DEPLOYMENT FOR UNIFIED NEURAL MOTION CONTROL IN ROBOTICS SYSTEMS AND APPLICATIONS

A unified neural motion control model using policy distillation and mode masks enables seamless adaptation to multiple control modes and interfaces, enhancing performance and reducing resource requirements.

DE102025140631A1Pending Publication Date: 2026-04-09NVIDIA CORP
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
DE · DE
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-10-06
Publication Date
2026-04-09

AI Technical Summary

Technical Problem

Conventional neural motion control models are adapted to specific control modes and interfaces, limiting their ability to efficiently transition to different types of control modes and interfaces, requiring additional time and resources for training and deployment of multiple models.

Method used

A machine learning model is trained using policy distillation to generate actions within a unified command space, incorporating mode masks to filter irrelevant attributes, enabling it to adapt to multiple control modes and interfaces, and is integrated into a controller for seamless task switching.

Benefits of technology

The model can perform a variety of motion control tasks effectively with fewer errors and reduced resource expenditure by operating across different control modes and interfaces, improving adaptability and performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

In various examples, a technique for neural motion control involves determining a multitude of masks that correspond to one or more control modes for an articulated object. The technique also involves applying the multitude of masks to a multitude of target state attributes for the articulated object to generate one or more masked target state attributes. Furthermore, the technique involves generating one or more actions based on at least the one or more masked target state attributes and performing a task using the articulated object based on at least the one or more actions by executing a trained machine learning model.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL AREA

[0001] Embodiments of the present disclosure relate generally to machine learning and motion control, and in particular to its use for unified neural motion control in robot systems and applications. BACKGROUND

[0002] Neural motion control refers to the use of a neural network (or another type of machine learning model) to move and / or animate a physical and / or virtual entity. For example, a deep learning model can be trained to perform motion control—or at least to serve as part of a machine's controller—by generating a sequence of poses (e.g., positions and orientations) corresponding to the movement of a human, animal, robot, vehicle, and / or other type of articulated object or machine. The output poses can then be integrated into a robot or other autonomous machine, a game, animation, simulation, and / or other application involving the articulated object.

[0003] Conventional techniques for implementing neural motion control involve developing separate neural motion control models to handle specific control modes (e.g., mechanisms for controlling movements) and / or control interfaces (e.g., mechanisms for receiving control signals related to movements and / or motion sequences). For example, three different machine learning models can be trained to perform navigation by tracking the root velocity of a virtual character, to generate expressive movements by tracking joint angles, or to perform teleoperation by kinematically tracking selected key body points.Furthermore, each machine learning model can be adapted to process inputs and / or commands associated with (but not limited to) a particular type of control interface such as a joystick, keyboard, motion capture system, exoskeleton and / or virtual reality, augmented reality and / or mixed reality (VR / AR / MR) headset.

[0004] However, adapting a particular neural motion control model to a specific set of control modes and / or interfaces can prevent that model from efficiently adapting to other types of control modes and / or interfaces. For example, a machine learning model that uses center-of-body velocity tracking to perform bipedal locomotion for a robot on uneven terrain may struggle to transition to a task involving precise bimanual manipulation. Instead, additional time and resources may be required to train and deploy a different machine learning model that uses joint angle and / or end-effector tracking to perform bimanual manipulation on the robot.

[0005] As the above illustrates, more effective techniques for carrying out neural movement control are needed in engineering. SUMMARY

[0006] The invention is defined by the claims. To illustrate the invention, aspects and embodiments that may or may not fall within the scope of the claims are described here.

[0007] A technique for neural motion control is disclosed, which involves determining a multitude of masks associated with one or more control modes for an articulated object. The technique also includes applying the multitude of masks to a multitude of target state attributes for the articulated object to generate one or more masked target state attributes. Furthermore, the technique includes generating one or more actions based on at least the one or more masked target state attributes and performing a task using the articulated object based on at least the one or more actions by executing a trained machine learning model.

[0008] The revelation extends to all novel aspects or features described and / or illustrated here.

[0009] Further features of the disclosure are characterized by the independent and dependent claims.

[0010] Any feature in one aspect of the disclosure can be applied to other aspects of the disclosure in any suitable combination. In particular, procedural aspects can be applied to device or system aspects and vice versa.

[0011] Furthermore, features implemented in hardware can also be implemented in software, and vice versa. Any reference to software and hardware features in this description should be interpreted accordingly.

[0012] Each system or device feature as described herein can also be provided as a process feature, and vice versa. Functionally described system and / or device aspects (including means and functional features) can alternatively be expressed in terms of their corresponding structure, such as a suitably programmed processor and allocated working memory.

[0013] It should also be understandable that certain combinations of the various features described and defined in any aspect of the revelation can be implemented and / or provided and / or used independently of one another.

[0014] The disclosure also provides computer programs and computer program products comprising software code designed to perform one of the methods described herein when executed on a data processing device and / or to embody one of the device and system features described herein, including one or all component steps of a method.

[0015] The disclosure also provides a computer or computing system (including networked or distributed systems) with an operating system that supports a computer program for carrying out one of the methods described herein and / or for embodying one of the device or system features described herein.

[0016] The revelation also provides a computer-readable medium on which one or more of the aforementioned computer programs are stored.

[0017] The revelation also provides a signal that carries one or more of the aforementioned computer programs.

[0018] The disclosure extends to processes and / or devices and / or systems as described herein with reference to the accompanying drawings.

[0019] Aspects and embodiments of the disclosure will now be described purely by way of example with reference to the attached drawings. BRIEF DESCRIPTION OF THE DRAWINGS Fig. Figure 1 illustrates a block diagram of a computer system configured to implement one or more aspects of at least one embodiment; Fig. Figure 2 is a more detailed illustration of the training unit and the execution unit from Fig. 1 according to at least one embodiment; Fig. Figure 3 illustrates an example of a distillation process, starting from the training unit. Fig. 2 is used to train a Student Policy (learning strategy) according to at least one embodiment; Fig. 4A illustrates the exemplary operation of the execution unit from Fig. 1 in the generation of a movement for an articulated object according to at least one embodiment; Fig. Figure 5 illustrates a flowchart of a procedure for training a machine learning model within the framework of a unified neural motion control according to at least one embodiment; Fig. 6A is an example of sensor positions with corresponding fields of view or sensor fields for exemplary autonomous or semi-autonomous machines according to at least some embodiments of the present disclosure; Fig. Figure 6B is an illustration of an example of component and sensor positions on an autonomous or semi-autonomous vehicle according to at least some embodiments of the present disclosure; Fig. 6C is a block diagram of an exemplary system architecture for an autonomous or semi-autonomous vehicle, an autonomous or semi-autonomous robot and / or other type of machine according to at least some embodiments of the present disclosure; Fig. Figure 6D is a block diagram of an exemplary architecture of a computer system - such as a system-on-a-chip (SoC) - according to at least some embodiments of the present disclosure; Fig. 6E is a system diagram for communication between one or more cloud-based servers and an exemplary autonomous or semi-autonomous vehicle, robot and / or other type of machine according to at least some embodiments of the present disclosure; Fig. Figure 7 is a system diagram illustrating an ecosystem with three computers, including a computer system for generating or creating artificial intelligence (AI) - such as AI training and validation data -, a computer system for training artificial intelligence, and a computer system that deploys the AI ​​at the system edge, according to at least some embodiments of the present disclosure; Fig. Figure 8 is a block diagram of an exemplary computer system for generative artificial intelligence (AI) according to at least some embodiments of the present disclosure; and Fig. Figure 9 is a block diagram of an exemplary computing device according to at least some embodiments of the present disclosure. DETAILED DESCRIPTION

[0020] Conventional techniques for implementing neural motion control involve adapting individual neural motion control models to different control modes and / or interfaces. Consequently, a conventional neural motion control model configured for a specific set of control modes and / or interfaces may struggle to adapt to other types of control modes and / or interfaces. Instead, additional time and resources can be spent training and deploying additional machine learning models to handle other types of control modes and / or interfaces.

[0021] To overcome the aforementioned limitations, the disclosed techniques train and execute a machine learning model to perform neural motion control within a unified command space. For example, the machine learning model can be used to output actions that control the movement of the entire body in a humanoid robot, robotic arm or end effector, forklift truck, construction machine, warehouse machine, and / or other type of articulated object or machine.The actions can be generated within various control modes and associated command spaces, such as (but not limited to) kinematic position tracking, where desired positions of important rigid body points in the articulated object are tracked; joint angle tracking, where desired angles for joints and / or motors in the joints of the articulated object are tracked; and / or center-of-body tracking, where a desired velocity, height, orientation and / or other attribute of a center-of-body position in the articulated object is tracked.

[0022] The machine learning model can correspond to a student policy (learning strategy; also: learning guideline) that is trained using a policy distillation technique to output actions consistent with those of an oracle policy (reference strategy; also: reference guideline). The oracle policy can comprise a different machine learning model trained to mimic extensive human motion data (e.g., from a motion capture dataset). Motor skills related to balance, coordination, and / or motion control can thus be learned from the oracle policy and transferred to the student policy, enabling it to adapt to different or multiple control modes and generalize across various types of tasks involving these motor skills.

[0023] The input to the student policy includes the proprioception associated with the articulated object and a representation of a target state. The proprioception can include (but is not limited to) a history of joint positions, joint velocities, base angular velocities, actions, and / or other states over a specified number of recent time steps. The target state can include (but is not limited to) changes to be induced in positions, orientations, linear velocities, angular velocities, and / or other attributes during the current time step. During student policy training, mode masks corresponding to one or more command modes for one or more parts of the articulated object (e.g., the upper and lower body of a humanoid robot) are used to remove attributes not associated with the one or more command modes from the target state.A sparsity mask is also used to selectively omit a subset of the remaining attributes from the target state, resulting in a masked target state that is used as input to the student policy. The student policy is executed to generate an action for the current time step, and it is trained with a loss calculated using the generated action and a reference action generated by the Oracle policy for the same time step.

[0024] After the student policy has been trained, it is integrated into a controller that generates movements for the articulated object based on control inputs assigned to different control modes. For example, the trained student policy can be used to generate actions that allow the articulated object to smoothly switch between control modes and / or tasks, use different control modes with different parts of the articulated object, and / or otherwise generate movements according to multiple control modes.

[0025] One advantage of the revealed techniques over previous approaches is the ability to generate actions in the articulated object using a single machine learning model, supporting multiple control modes and / or control interfaces. Consequently, the machine learning model can perform a variety of motion control tasks more effectively than conventional techniques that adapt individual neural motion control models to different control modes and / or interfaces. Furthermore, because the machine learning model learns motor skills shared across all control modes, the movements generated by the machine learning model may exhibit fewer errors and / or improved performance compared to machine learning models trained to process specific control modes.Furthermore, the ability of the machine learning model to operate within different combinations of control modes and / or control interfaces can reduce the time and resource expenditure compared to conventional approaches where multiple machine learning models are trained and deployed to handle different types of tasks and / or inputs.

[0026] The examples mentioned above are in no way intended to be restrictive. As experts in this field know, techniques for automatically generating dialogue flows from unlabeled conversation data can generally be implemented in any suitable application.

[0027] The systems and methods described herein can be used for a number of purposes, including but not limited to use in systems related to machine control, machine locomotion, machine propulsion, synthetic data generation, model training, perception, augmented reality, virtual reality, mixed reality, robotics, security and surveillance, simulation and digital twinning, autonomous or semi-autonomous machine applications, deep learning, environmental simulation, object or actor simulation and / or digital twinning, data center processing, conversational AI, light transport simulation (e.g., ray tracing, path tracing, etc.), collaborative content creation for 3D assets, cloud computing, and / or other suitable applications.

[0028] Disclosed embodiments may be included in a wide variety of different systems, such as automotive systems (e.g., an infotainment or plug-in gaming / streaming system of an autonomous or semi-autonomous vehicle), systems implemented with a robot, aviation systems, media systems, boat systems, intelligent area surveillance systems, systems for performing deep learning operations, systems for performing simulation operations, systems for performing digital twin operations, systems implemented using an edge device, systems incorporating one or more virtual machines (VMs), systems for performing operations to generate synthetic data, systems implemented at least partially in a data center, systems for performing conversational AI operations, systems,that implement one or more language models such as Large Language Models (LLMs), Small Language Models (SLMs), Vision Language Models (VLMs), and / or multimodal language models capable of processing text, audio, and / or image data; systems for performing light transport simulation; systems for performing collaborative content creation for 3D assets (e.g., systems or platforms using Universal Scene Descriptor (USD) data such as OpenUSD); systems implemented at least partially using cloud computing resources; systems for performing operations with generative AI; and / or other types of systems.

[0029] Approaches according to various embodiments can be used to generate one or more parameters for a content generation environment. In at least one embodiment, a trained machine learning (ML) and / or artificial intelligence (AI) system, such as a large language model (LLM) or a vision language model (VLM), can be used to generate parameters for the content generation environment, such as camera settings, scene lighting, video parameters, and / or the like, but not limited to those used to display objects within a scene. The parameters can be based on input provided by a user or a proxy instance to a trained language model (e.g., LLM, VLM, etc.), which can then generate one or more environments according to the input.Various embodiments can be used to generate settings in two-dimensional (2D) or three-dimensional (3D) environments. In embodiments that include one or more language models—that is, one or more LLMs, one or more VLMs, or a combination of LLMs and VLMs—the one or more language models can receive input (e.g., a prompt, a request, a query, etc.) that is parsed or otherwise formatted to generate deterministic output. For example, the input provided to the language model might include a specific format for the output results, an example of desired output results, a specific list of parameters and their respective formatting, and the like. An input generator (e.g.,A prompt generator, which can be controlled or otherwise directed by one or more AI and / or ML systems, can be used to generate this input based on an initial input received from a user, device, proxy, and / or the like. A modified input generated by the prompt generator can then be provided to the language model, which generates an output parameter set. This output can be further evaluated by a validator or other system to ensure its appropriateness. Subsequently, a configuration file can be generated, and / or the parameters can be directly provided to an environment to configure various components (e.g., camera settings, lighting, etc.) based on the parameters generated by the language model.

[0030] In some examples, the machine learning models described herein (e.g., deep neural networks, language models, LLMs, VLMs, multimodal language models, perceptual models, tracking models, fusion models, transformer models, diffusion models, pure encoder models, pure decoder models, encoder-decoder models, neural rendering field (NERF) models, etc.) can be packaged as a microservice—such as an inference microservice (e.g., NVIDIA NIMs)—which may contain a container (e.g., an operating system-level virtualization package) that may contain an application programming interface (API) layer, a server layer, a runtime layer, and / or at least one model "engine." The inference microservice may, for example, contain the container itself and one or more models (e.g., weights and biases). In some cases, where one or more machine learning models are small enough (e.g.,(Having a sufficiently small number of parameters), the one or more models can be contained within the container itself. In other examples—where, for instance, the one or more models are large—the one or more models can be hosted / stored in the cloud (e.g., in a data center) and / or hosted on-premises and / or at the network edge (e.g., on a local server or local computing device, but outside the container). In such embodiments, access to the one or more models can be provided via one or more APIs, such as REST APIs. This, and in some embodiments, allows the machine learning models described herein to be deployed as an inference microservice to accelerate the deployment of a model or models to any cloud, data center, or edge computing system while ensuring data security.For example, the inference microservice may include one or more APIs, a pre-configured container for simplified deployment, an optimized inference engine (e.g., implemented using standardized software for deploying and running AI models such as NVIDIA's Triton Inference Server and / or one or more APIs for high-performance deep learning inference, which may include an inference runtime environment as well as model optimizations to provide low latency and high throughput for production applications - such as NVIDIA's TensorRT), and / or enterprise-related management data for telemetry (e.g., including identity, metrics, system checks, and / or monitoring).

[0031] The one or more machine learning models described herein can be included as part of the microservice, along with an accelerated infrastructure capable of single-command deployment and / or orchestration and auto-scaling using a container orchestration system on an accelerated infrastructure (e.g., from a single device to data center scaling). Thus, the inference microservice can include the one or more machine learning models (e.g., optimized for high-performance inference), inference runtime software for executing the one or more machine learning models and providing outputs / responses to inputs (e.g., user queries, prompts, etc.), and enterprise-grade management software for providing system checks, identity verification, and / or other monitoring.In some embodiments, the inference microservice may include software that enables the replacement and / or updating of one or more machine learning models in place. During the replacement or update process, the software performing the replacement / update can retain the user configurations of the inference runtime software and the enterprise management software.

[0032] In some embodiments, the system and methods described herein can be used in a robotics application. For example, a robot or robotic system may contain one or more onboard processors (e.g., CPUs, GPUs, hardware-based deep learning accelerators (DLAs), hardware-based programmable vision accelerators (PVAs) which may include one or more vector processing units (VPUs), direct memory access systems (DMAs) and / or pixel processing units (PPEs), hardware-based optical flow accelerators (OFAs), SoCs, etc.) and working memory and / or storage (e.g., for storing control algorithms, sensor data, and one or more machine learning models). The robotic system can use these processors to execute one or more machine learning models (e.g.,Language models, vision language models (VLMs), large language models (LLMs), vision language action (VLA) models, multimodal language models (MMLMs), etc., enable it to perform complex tasks autonomously or semi-autonomously, such as interacting with and / or manipulating static and / or dynamic objects or navigating environments using sensors like cameras, LiDAR, radar, ultrasonic sensors, etc. The system can use sensor fusion techniques to combine data from multiple sensors (e.g., cameras, infrared, LiDAR, radar, accelerometers) and create a comprehensive model of the robot's environment. This data can be processed locally on the robot or sent to remote servers for more computationally intensive tasks such as 3D mapping or SLAM (Simultaneous Localization and Mapping). In one or more embodiments, data from individual robots (e.g.,Sensor data, task status, or environmental conditions are uploaded to the cloud, where centralized AI models can analyze optimized commands and distribute them to an entire fleet. In some embodiments, the machine learning models described herein (e.g., language models, VLMs, VLAs, LLMs, SLMs, MMLMs, diffusion models, NeRF models, DNNs, etc.) can be used to enable the robot to perceive its environment, draw inferences, and / or communicate with one or more other robots and / or people in the environment. In some embodiments, the robot can communicate (e.g., using one or more network interface cards (NICs) and / or data processing units (DPUs)) with one or more locally hosted servers / computing devices and / or with one or more remote servers / computing devices (e.g., in one or more data centers).

[0033] Although examples relating to the use of machine learning models such as neural networks may be described herein, this is not intended as a limitation. For example, and without limitation, any of the various machine learning models and / or neural networks described herein may include any type of machine learning model, such as one or more machine learning models, linear regression, logistic regression, decision trees, support vector machines (SVMs), Naive Bayes, k-nearest neighbor (Knn), k-means clustering, random forest, dimensionality reduction algorithms, gradient boosting algorithms, neural networks (e.g., neural autoencoder networks, artificial neural networks (ANNs), convolutional neural networks (CNNs), recurrent neural networks (RNNs), perceptrons, long / short-term memory (LSTM) networks,Multilayer perceptron (MLP) networks, deep-stacking networks (DSNs), generative pre-training (GPT) models or networks, feedforward networks, radial basis-function ANNs, self-organizing maps (SOMs), Kohonen maps, Hopfield networks, Boltzmann machines, deep-belief neural networks, deconvolutional neural networks, generative adversarial networks (GANs), liquid-state machines, modular neural networks, liquid-state machines, sequence-to-sequence models, networks with transformer architectures, state-space models (SSMs) (e.g., networks with Mamba architectures (e.g., Mamba-1, Mamba 2, etc.)). Networks with selective state-space models, networks with structured state-space sequence models, etc.), diffusion models (e.g., probabilistic diffusion models,Score-based generative models, etc.), neural radiance field models (NeRF), Gaussian splat models, Kolmogorov-Arnold networks (KANs), models with pure encoder architectures, models with pure decoder architectures, models with encoder-decoder architectures, generative machine learning models, language models, large language models (LLMs), vision-language models (VLMs), multimodal language models (MMLMs), large action models (LAMs), vision-language-action (VLA) models, etc.) and / or other types of machine learning models. SYSTEM OVERVIEW

[0034] Fig. Figure 1 is a block diagram illustrating a computer system 100 configured to implement one or more aspects of at least one embodiment. In at least one embodiment, the computer system 100 may comprise any type of computing device, including, without limitation, a server computer, a server platform, a desktop computer, a laptop computer, a handheld / mobile device, a digital kiosk, an in-vehicle infotainment system, a smart speaker or display, a television, and / or a portable device. In at least one embodiment, the computer system 100 is a server computer operating in a data center or cloud computing environment that provides scalable computing resources as a service over a network.

[0035] In various embodiments, the computer system 100 comprises, without limitation, one or more processors 102 and one or more memories 104, which are coupled to a parallel processing subsystem 112 via a memory bridge 105, and a communication path 113. The memory bridge 105 is further coupled to an input / output (I / O) bridge 107 via a communication path 106, and the I / O bridge 107 is in turn coupled to a switch 116.

[0036] In one embodiment, the I / O bridge 107 is configured to receive user input information from optional input devices 108 (but not limited to), such as a keyboard, mouse, touchscreen, sensor data analysis (e.g., for evaluating gestures, speech, or other information about one or more applications in a field of view or sensor field of one or more sensors), a VR / MR / AR headset, a gesture recognition system, a steering wheel, mechanical, digital, or touch-sensitive buttons or input components, and / or a microphone, and to forward the input information to the one or more processors 102 for processing. In at least one embodiment, the computer system 100 can be a server computer in a cloud computing environment. In such embodiments, the computer system 100 can omit input devices 108 and receive equivalent input information as commands (e.g., commands, commands, etc.).in response to one or more inputs from a remote computing device) and / or receive messages transmitted over a network and received via the network adapter 118. In at least one embodiment, the switch 116 is configured to provide connections between the I / O bridge 107 and other components of the computer system 100, such as a network adapter 118 and various expansion cards 120 and 121.

[0037] In at least one embodiment, the I / O bridge 107 is coupled to a system hard disk 114, which can be configured to store content and applications as well as data for use by the one or more processors 102 and the parallel processing subsystem 112. In one embodiment, the system hard disk 114 provides non-volatile memory for applications and data and can include fixed or removable hard disk drives, flash memory devices, and CD-ROM (Compact Disc Read-Only Memory), DVD-ROM (Digital Versatile Disc-ROM), Blu-ray, HD-DVD (High-Definition DVD), or other magnetic, optical, or solid-state storage devices. In various embodiments, other components such as universal serial bus or other port connectors, compact disc drives, digital versatile disc drives, movie recording devices, and the like can also be connected to the I / O bridge 107.

[0038] In various embodiments, the memory bridge 105 can be a northbridge chip and the I / O bridge 107 a southbridge chip. Furthermore, the communication paths 106 and 113, as well as other communication paths within the computer system 100, can be implemented using any technically suitable protocols, including, without limitation, AGP (Accelerated Graphics Port), HyperTransport, or any other bus or point-to-point communication protocol known in the art.

[0039] In at least one embodiment, the parallel processing subsystem 112 comprises a graphics subsystem that transmits pixels to an optional display device 110, which may be a conventional cathode ray tube, a liquid crystal display, a light-emitting diode display, and / or the like. In such embodiments, the parallel processing subsystem 112 may include circuits optimized for graphics and video processing, including, for example, video output circuits. Such circuits may be integrated via one or more parallel processing units (PPUs), also referred to herein as parallel processors, which are included in the parallel processing subsystem 112.

[0040] In at least one embodiment, the parallel processing subsystem 112 comprises circuits optimized for general-purpose and / or computational processing (e.g., undergoing optimization). Such circuits may also be included in one or more power processing units (PPUs) contained within the parallel processing subsystem 112 and configured to perform such general-purpose and / or computational operations. In further embodiments, the one or more PPUs contained within the parallel processing subsystem 112 may be configured to perform graphics processing, general-purpose processing, and / or computational operations. The one or more memory locations 104 contain at least one device driver configured to manage the processing operations of the one or more PPUs within the parallel processing subsystem 112.Furthermore, the one or more memory units contain 104 instructions for implementing a training unit 122 and an execution unit 124, which can be executed by one or more processors and / or the parallel processing subsystem 112.

[0041] In various embodiments, the parallel processing subsystem 112 can be combined with one or more of the other elements from Fig. 1. It can be integrated to form a single system. For example, the parallel processing subsystem 112 can be integrated with one or more processors 102 and other interconnect circuitry on a single chip to form a system-on-a-chip (SoC).

[0042] The one or more processors 102 can include any suitable processor, which can be a central processing unit (CPU), graphics processing unit (GPU), application-specific integrated circuit (ASIC), field-programmable gate array (FPGA), artificial intelligence (AI) accelerator, deep learning accelerator (DLA), parallel processing unit (PPU), data processing unit (DPU), vector / vision processing unit (VPU), or programmable vision accelerator (PVA) (which may include one or more VPUs, pixel processing engines (PPEs), and / or direct memory access (DMA) systems).any other type of processing unit or a combination of different processing units, such as one or more CPUs configured to operate in conjunction with one or more GPUs, is implemented. In general, the one or more processors 102 can comprise any technically feasible hardware unit capable of processing data and / or executing software applications. Furthermore, in the context of this disclosure, the computing elements shown in the computer system 100 can correspond to a physical computer system (e.g., a system in a data center or machine) and / or to a virtual computing instance running within a computing cloud.

[0043] In at least one embodiment, the one or more processors 102 issue instructions that control the operation of the PPUs. In at least one embodiment, the communication path 113 is a link of a Peripheral Component Interconnect Express (PCIe) connection, in which dedicated channels are assigned to each PPU. Other communication paths may also be used. The PPU advantageously implements a highly parallel processing architecture, and the PPU may be equipped with any amount of local parallel processing memory (PP memory).

[0044] It is understood that the system shown here serves for illustrative purposes and that variations and modifications are possible. The connection topology, including the number and arrangement of the bridges, the number of processors 102, and the number of parallel processing subsystems 112, can be modified as desired. For example, in at least one embodiment, the one or more memories 104 can be directly connected to the one or more processors 102 instead of via a memory bridge 105, and other devices can communicate with the one or more memories 104 via the memory bridge 105 and the processors 102. In other embodiments, the parallel processing subsystem 112 can be connected to the I / O bridge 107 or directly to the one or more processors 102 instead of to the memory bridge 105.In further embodiments, the I / O bridge 107 and the memory bridge 105 can be integrated into a single chip instead of being provided as one or more discrete devices. In certain embodiments, one or more of the components described in... Fig. The components shown in Figure 1 may not be present. For example, the switch 116 may be omitted, and the network adapter 118 and the expansion cards 120, 121 would be directly connected to the I / O bridge 107. Furthermore, in certain embodiments, one or more of the components shown in Figure 1 may be omitted. Fig. The components shown in Figure 1 can be implemented as virtualized resources in a virtual computing environment, such as a cloud computing environment. In particular, the parallel processing subsystem 112 can be implemented as a virtualized parallel processing subsystem in at least one embodiment. For example, the parallel processing subsystem 112 can be implemented as one or more virtual graphics processing units (vGPUs) that render graphics on one or more virtual machines (VMs) running on one or more server computers, whose GPU(s) and other physical resources are shared by one or more VMs. Unified Neural Motion Control

[0045] Fig. Figure 2 is a more detailed illustration of training unit 122 and execution unit 124 from Fig. 1 according to at least one embodiment. As explained herein, the training unit 122 and the execution unit 124 are configured to train (update) and execute one or more machine learning models to perform unified neural motion control. Each of these components is described in more detail below.

[0046] In one or more embodiments, the neural motion control comprises the iterative execution of a controller 250, which includes one or more machine learning models, to generate a different action 252 for each frame or time step in a motion 240 of an articulated object. For example, the controller 250 can be used to generate a specific motion 240 for a human, an animal, a robot, and / or another type of articulated object. Each action 252 output by the controller 250 can be used to update a configuration of joints and / or other parts of the articulated object at a corresponding time step, resulting in a corresponding motion 240 of the articulated object.A sequence of actions generated by the Controller 250 for a corresponding sequence of time steps can be used to replicate one or more motor skills, such as (for example and without limitation): walking, racing, running, jumping, turning, pivoting, leaning, dancing, kicking, hitting, waving, climbing, dismounting, gesturing, squatting, hopping, skipping and / or manipulating objects.

[0047] As in Fig. As shown in Figure 2, each action 252 is generated by the controller 250 (which may implement or contain components, features, etc., of the controller 636) based on a target state 230 and a proprioception 236. The target state 230 comprises a state to be achieved by performing the action 252 during the current time step. The target state 230 may, for example, include desired positions, orientations, angles, linear velocities, angular velocities, and / or other attributes of joints and / or other parts of the articulated object. The target state 230 may also, or alternatively, include higher-level tasks such as (but not limited to) maintaining balance during locomotion, achieving a specific spatial position for object manipulation, and / or adhering to trajectory constraints derived from human motion data.

[0048] In some embodiments, the target state 230 is generated using a control input 232 and a group of mode masks 238 that are associated with one or more control modes 234. The control input 232 includes commands, sensor data, data received via one or more input devices, and / or other types of external control signals related to movements and / or motion sequences of the articulated object. These external control signals can be provided via various control interfaces such as (but not limited to) joysticks, keyboards, game controllers, motion capture systems, virtual reality (VR) headsets and / or controllers, exoskeletons, and / or robotic arms.For example, the control input can include (but is not limited to) motion detection data, positions and / or states of a joystick and / or other type of input device, gestures, haptic feedback and / or electroencephalography (EEG) signals.

[0049] The mode masks 238, which are assigned to the various control modes 234 for the articulated object, are applied to attributes in the target state 230 to generate inputs in the controller 250. The control modes 234 correspond to different mechanisms and / or strategies for controlling the movement of the articulated object. Each control mode is associated with a specific type of control input 232 and / or one or more attributes contained in the target state 230. For example, a control mode corresponding to tracking the body center for locomotion can convert a control input 232, which specifies a direction, velocity, trajectory, and / or other attribute associated with the overall movement of the articulated object, into a desired velocity, height, and / or orientation for a position of the body center within the articulated object.In another example, a control mode corresponding to joint angle tracking can convert a control input 232, which specifies a desired position of an end effector in the articulated object, into a group of desired joint angles for joints in the articulated object. In a third example, a control mode corresponding to kinematic position tracking can be used to convert a control input 232 from a VR / AR / MR controller and / or other control interface that tracks a user's pose into desired positions of key points (e.g., head, torso, shoulders, elbows, hands, hips, knees, ankles, etc.) on the articulated object.

[0050] In some embodiments, individual control modes 234 are selected and / or set for different parts of the articulated object. For example, different control modes 234 can be used to control the upper and lower body of a humanoid robot, enabling the humanoid robot to perform complex, multifunctional movements (e.g., walking in a particular direction while carrying an object, reaching for a door handle while maintaining balance, etc.). These control modes 234 can also be dynamically and / or dynamically updated as needed to adapt the movement 240 to different types of tasks, allowing the articulated object to seamlessly switch between different behaviors and / or motor skills.

[0051] After determining the control modes 234 for one or more parts of the articulated objects, the mode masks 238 assigned to these control modes 234 are applied to attributes in the target state 230. More precisely, mode masks 238 are used to filter attributes of the target state 230 based on the selected and / or specified control modes 234, so that attributes of the target state 230 that are not relevant to the control modes 234 and / or the parts of the articulated object to which the control modes 234 refer are omitted from the input to the controller 250. For example, mode masks 238 assigned to a control mode for tracking the body center can be used to retain attributes of the body center in the target state 230 and / or to filter out attributes relating to joint angles and / or kinematic positions from the target state 230.In another example, mode masks 238, which are assigned to a control mode for joint angle tracking, can be used to filter out attributes relating to the body center and / or joint angles in the articulated object from the target state 230 and / or to retain attributes relating to desired joint angles in the target state 230. In a third example, mode masks 238, which are assigned to a kinematic control mode for position tracking, can be used to retain desired positions of key points in the target state 230 and / or to omit joint angles and / or attributes of the body center from the target state 230.

[0052] In one or more embodiments, the proprioception 236 comprises data representing the current and / or past state of the articulated object. For example, the proprioception 236 may include (but is not limited to) a history of the k last joint positions, joint velocities, angular velocities, linear accelerations, contact forces, actions, and / or other physical attributes of the articulated system. This data may be derived from inertial data, torque data, force data, and / or other sensory inputs available to the articulated object in a real and / or simulated environment.

[0053] Based on the input, which includes the target state 230 and the proprioception 236, the controller 250 generates a corresponding action 252. For example, the controller 250 can use a machine learning model to map the target state 230 and the proprioception 236 to an action 252 that includes joint positions, joint orientations, key point positions, motor commands, body center velocity, and / or other types of output related to the motion 240 of the articulated object. The action 252 can then be used to actuate the degrees of freedom in the articulated object (e.g., via a proportional-derivative (PD) controller), resulting in an updated proprioception 236 for the next time step. The process can be repeated until a desired overall goal (e.g., a location to be reached by the articulated object, the completion of a manipulation task, the demonstration of a motor skill, etc.) is achieved.) is reached, no additional target state 230 is defined for the next time step and / or another predefined termination condition is met.

[0054] The training unit 122 generates the controller 250 by training machine learning models that include an Oracle Policy (reference strategy) 202 and a Student Policy (learning strategy) 204. In some embodiments, the Oracle Policy 202 acts as an expert and / or teacher model trained to mimic extensive motion data for humans (or other types of articulated objects), and the Student Policy 204 is trained to learn motor skills from the Oracle Policy 202 within a unified command space that includes multiple control modes 234 and / or types of control inputs 232.

[0055] In some embodiments, the training unit 122 trains the Oracle Policy 202 with a set of training data 212 and a set of training targets 210. The training data 212 comprises a set of reference movements 214(1)-214(X) (each referred to individually here as a reference movement 214) that correspond to the ground truth movements to be learned by the Oracle Policy 202. For example, each reference movement 214 may comprise a sequence of ground truth states representing the movement of the articulated object. Each state in the sequence may be associated with a different frame, or time step, in the corresponding reference movement 214.Each state in the sequence may also or instead include a center-of-body velocity, center-of-body height, joint positions, joint orientations, joint velocities, key point positions and / or other features that describe the configuration of the articulated object at a corresponding time step.

[0056] In one or more embodiments, the training unit 122 generates reference movements 214 in the training data 212 via a motion transfer process that converts human motion data (e.g., from one or more motion capture datasets) reflecting human body structures, shapes, and / or dynamics into humanoid motion data that reflects body structures, shapes, and / or dynamics in a humanoid articulated object. In the motion transfer technique, the training unit 122 can use forward kinematics to compute key point positions in the humanoid (e.g., in a Cartesian coordinate space).Training Unit 122 can also adapt a skinned multi-person linear (SMPL) model to the humanoid's structure by optimizing the SMPL parameters to match the calculated keypoint positions, thereby establishing a correspondence between human and humanoid motion that accounts for differences in joint structures, proportions, and / or joint connections between humans and humanoids. Training Unit 122 can then transfer the human motion data by matching keypoints between the adapted SMPL model and the humanoid. For example, Training Unit 122 can use gradient descent and / or another optimization technique to minimize the discrepancy between keypoints in the human motion data and corresponding keypoints in the humanoid.The motion transfer process can thus ensure that human-like movements can be performed using the kinematic structure of the humanoid.

[0057] The training data 212 also include several groups of training proprioceptions 216(1)-216(Y) (each referred to individually here as training proprioceptions 216) and training target states 218(1)-218(Z) (each referred to individually here as training target states 218), each group of training proprioceptions 216 and each group of training target states 218 being derived from a corresponding reference movement 214. Like the proprioception 236, each training proprioception includes a representation of current and / or past states associated with a particular time step in a corresponding reference movement 214. For example, a particular training proprioception may include rigid body positions, orientations, linear velocities, angular velocities, prior actions, and / or other state representations for each time step in a corresponding reference movement 214.

[0058] Like the target state 230, each training target state comprises a state that is to be achieved at a specific time step within a corresponding reference movement 214. For example, each group of training target states 218 may include (but is not limited to) desired joint positions, key point positions, center-of-body velocities, and / or other attributes of the articulated object at various time steps within a corresponding reference movement 214.

[0059] In some embodiments, the training unit 122 uses reinforcement learning (RL) to train an agent that includes Oracle Policy (Reference Strategy 202) and / or Student Policy (Learning Strategy) 204 to track and / or mimic human movements (or movements associated with another type of articulated object) in real time. For example, a particular policy (which may include Oracle Policy 202 and / or Student Policy 204) can be defined by π(a t | s t ) are represented, where s t a state at time step t that affects proprioception 236 stp and the target state 230 stg includes, and Action 252 a t ∈ ℝ 19 The desired joint positions are represented, which are converted into a corresponding movement of 240. A reward rt=R(stp,stg) Policy optimization is defined using proprioception 236 and the target state 230, and proximal policy optimization (PPO) can be used to determine the cumulative discounted reward. E[∑t=tT−1γt−1rt] to maximize. This RL framework can be viewed as a command-following task, where the articulated object learns to follow the training target states 218 and to track the reference motion 214 at each time step.

[0060] The training unit 122 also trains the Oracle Policy 202 using representations of reference movements 214, training proprioceptions 216, and training target states 218 in the training data 212. More precisely, the Oracle Policy 202 can be described as πoracle(at|stp−oracle,stg−oracle) be represented, whereby stp−oracle the training proprioceptions entered into Oracle Policy 202 are referred to as 216, stg−oracle the training target states 218 entered into Oracle Policy 202 and a t The training actions 222 issued by Oracle Policy 202 based on the entered training proprioceptions 216 and training target states 218 are referred to as the training actions 222. Each training proprioception can be defined by stp−oracle≜[pt,θt,p˙t,ωt,at−1] can be arranged and includes a rigid body position p t , an orientation θ t , a linear velocity ṗ t and an angular velocity ω t as well as an action from the previous time step a t-1 Values ​​in a specific training proprioception can be determined by simulating the physics of the articulated object. Each training target state can be described as stg−oracle≜[θ^t+1⊖θt,p^t+1−pt,v^t+1−vt,ω^t+1−ωt,θ^t,p^t] are represented and include a reference pose (θ̂ t , p̂ t ) for the time step t and a one-frame difference of the reference position p̂ t+1 , the orientation θ̂t +1 , the linear velocity ν̂ t+1 and the angular velocity ω̂ t+1 for the next time step t + 1 and the position p t , the orientation θ t , the linear velocity ν t and the angular velocity ω t for the current time step t.

[0061] In some embodiments, Oracle Policy 202 includes a multi-layer perceptron (MLP) with three layers and corresponding dimensions of [512, 256, 128]. During each time step of a training episode, the MLP processes an input training proprioception and a target state and outputs a corresponding training action. The training unit 122 computes one or more training targets 210 using the output training action and updates the parameters of Oracle Policy 202 (e.g., using PPO and / or another optimization technique) in a way that optimizes the training targets 210. For example, the training targets 210 might be a reward r tThis includes terms calculated as the sum (or another combination) of three components: 1) penalty, 2) regularization, and 3) a set of task rewards. The penalty component may include terms related to torque limits that reduce excessive use of joint torque; joint position and / or joint velocity limits that ensure realistic joint movements; termination components that penalize unstable movements (e.g., falls); and / or other terms that counteract undesirable behaviors and / or movements.The regularization component can include terms related to joint acceleration and / or velocity, upper and / or lower body movement rates, differences between actions for successive time steps, torques, the time the articulated object's feet are in the air, the maximum foot height at a given step, contact force between the feet and the ground, stumbling, slipping, foot orientation, body center orientation, and / or other attributes that can be used to refine the articulated object's motion. Task rewards can include differences between reference and actual positions, velocities, rotations, and / or other attributes of joints in the articulated object, the articulated object's body, and / or other parts of the articulated object. The task rewards can thus be used to optimize whole-body tracking in real time.

[0062] During Oracle Policy 202 training, Training Unit 122 can randomize the parameters of the simulated environment (e.g., NVIDIA Isaac Sim™, NVIDIA Isaac Gym™, NVIDIA Isaac Lab™, and / or NVIDIA Drive Sim™, which are registered trademarks of NVIDIA Corporation) used to determine the training proprioceptions 216. This allows the movements learned by Oracle Policy 202 to be adapted to different simulated and / or real-world environments and / or conditions. For example, Training Unit 122 can randomize the terrain in a given simulated environment by adding obstacles and / or varying the topography and / or roughness of the terrain. Training Unit 122 can also, or alternatively, generate external disturbances at random intervals to simulate (but is not limited to) real-world disturbances such as uneven terrain, wind, and / or other environmental factors.The training unit 122 can also or instead randomize parameters associated with the dynamics (e.g., coefficients of friction, base mass offset, link mass, proportional gain, differential gain, amplification, random torque force input, control delay, motion reference offset, etc.) of the simulated environment by extracting them from corresponding distributions.

[0063] After training Oracle Policy 202, Training Unit 122 uses a policy distillation technique and the reference actions 226 output by the trained Oracle Policy 202 to train Student Policy 204. More specifically, Training Unit 122 inputs representations of training proprioceptions 216 and training target states 218, mapped to different time steps in reference movements 214, into both Oracle Policy 202 and Student Policy 204. Training Unit 122 executes Oracle Policy 202 to generate reference actions 226 for the same time steps. Training Unit 122 also executes Student Policy 204 to generate training actions 224 for the same time steps. The training unit 122 calculates one or more training objectives 220 using the generated training actions 224 and the corresponding reference actions 226 and updates the parameters of the student policy 204 (e.g.using PPO and / or another optimization technique) in a way that optimizes the training objectives 220.

[0064] In some versions, Student Policy 204 is considered πstudent(at|stp−student,stg−student) depicted, whereby stp-student the training proprioceptions 2 216 are referred to as being entered into Student Policy 204, stg−oracle The training target states 218 are designated as entered into Student Policy 204 and a t The training actions 224 issued by the Student Policy 204 based on the input training proprioceptions 216 and training target states 218 are designated. Each training proprioception can be stp−student≜[q,q,ωbase,g]t−25:t∪[at−25:t−1] It is represented and includes the joint positions q, the joint velocities q̇, the base angular velocities ω. base, the gravity vector g and the actions a for the previous 25 time steps.

[0065] In one or more embodiments, the training unit 122 applies a group of mode masks 206 and a group of sparsity masks 208 to training target states 218 associated with a specific reference motion 214 in training data 212 to generate a corresponding group of masked target states 228. The training unit 122 uses masked target states 228 as representations of training target states 218 that are input into the student policy 204.

[0066] Like mode masks 238, mode masks 206 are also assigned control modes for one or more parts of the articulated object and can be used to filter training target states 218 based on their relevance to the control modes. For example, a particular training episode might be assigned a control mode for kinematic position tracking for the upper body of the articulated object and control modes for tracking joint angles and body center positions for the lower body of the articulated object. Mode masks 206 can thus be used to omit attributes in training target states 218 that are not assigned to upper body kinematic position tracking, lower body joint angle tracking, and / or body center tracking of the articulated object from the masked target states 228.

[0067] After applying the mode masks 206 to a specific training target state, sparsity masks 208 are applied to the remaining unmasked training target states 218 to selectively omit a subset of the attributes associated with the selected control modes from the masked target states 228. Continuing the example above, sparsity masks 208 can be applied to track only the kinematic positions of the hands in the upper body and / or only the joint angles of the trunk in the lower body.

[0068] The generation of masked target states 228 can thus be achieved by stg−student≜Msparsity⊙[Mmode⊙stg−upper,Mmode⊙stg−lower] are represented, where M sparsity which represents the sparsity masks 208, M mode which represents mode masks 206, stg−upper represents a subset of a specific training target state that is assigned to the upper body of the articulated object, and stg−lower represents a subset of a specific training target state associated with the lower body of the articulated object. The mode masks 206 can be used to selectively activate different command spaces (e.g., different groups and / or types of commands used to control the articulated object) to accommodate different types of tasks and / or to allow the student policy 204 to adapt to different control modes 234 and / or types of control inputs 232. The sparsity masks 208 can improve the robustness of the student policy 204 to missing inputs and / or further refine the tasks learned by the student policy 204.

[0069] In one or more embodiments, the mode masks 206 and the sparsity masks 208 are randomized at the beginning of each training episode. For example, each bit in mode masks 206 and sparsity masks 208 can be selected from a Bernoulli distribution. B(0.5) can be drawn. Once selected for a specific training episode, the same mode masks 206 and sparsity masks 208 can be applied to the training target states 218 throughout the entire training episode.

[0070] In some embodiments, the Student Policy 204 includes a multi-layer perceptron (MLP) with the same architecture as the Oracle Policy 202 and / or a different architecture than the Oracle Policy 202. During a given training episode, the training unit 122 uses the training actions 224a generated by the MLP. t , the masked target states 228 and the training proprioceptions 216, to establish a trajectory of (stp−student,stg−student) to obtain. Training unit 122 also calculates a corresponding trajectory of (stp−oracle,stg−oracle) for Oracle Policy 202 and executes Oracle Policy 202 to perform reference actions. t to obtain. Training unit 122 calculates one or more training goals 210 as a loss. L=‖a^t−at‖22 and updates the parameters of Oracle Policy 202 (e.g., using PPO and / or another optimization technique) in a way that minimizes the loss.

[0071] Fig. Figure 3 illustrates an example of a distillation process, which is based on training unit 122. Fig. 1 is used to train Student Policy (learning strategy) 204 according to at least one embodiment. As in Fig. As shown in Figure 3, reference movements 214 from a data set (e.g., a data set of human movements transferred to humanoids) are used to determine target states 218, which include key point target states 218(A), joint target states 218(B), and body center target states 218(C).

[0072] Key point target states 218(A) can include attributes in target states 218 that are used to perform key point position tracking. For example, key point target states 218(A) can include desired 3D positions of key points on the articulated object.

[0073] Joint target states 218(B) can include attributes in target states 218 that are used to perform joint angle tracking. For example, joint target states 218(B) can contain target joint angles for different joints and / or joint motors in the articulated object.

[0074] Target states of the body center 218(C) can include desired attributes of a body center in the articulated object. For example, target states of the body center 218(C) can include a desired velocity, height, and / or orientation of the body center for the articulated object.

[0075] Mode masks 206 are used to remove a subset of attributes not associated with one or more command modes from keypoint target states 218(A), joint target states 218(B), and / or body center target states 218(C). For example, mode masks 206 can be used to omit lower body positions and upper body joint angles from keypoint target states 218(A), joint target states 218(B), and / or body center target states 218(C). Mode masks 206 can thus correspond to command modes for tracking upper body keypoints, tracking lower body joint angles, and tracking the body center.

[0076] Sparsity masks 208 are used to select one or more remaining attributes 302(1)-302(4) for inclusion in masked target states 228. For example, sparsity masks 208 can be used to select a first attribute 302(1) corresponding to the torso position in the keypoint target states 218(A), a second attribute 302(2) corresponding to a specific joint angle of the lower body in the joint target states 218(B), a third attribute 302(3) corresponding to a velocity along the x-axis in the body center target states 218(C), and a fourth attribute 302(4) corresponding to a velocity along the z-axis in the body center target states 218(C) for inclusion in the masked target states 228.

[0077] The masked target states 228 and a group of training proprioceptions 216(A) are input into the student policy 204, and the student policy 204 is executed to perform a corresponding group of training actions 224a. t to generate. The unmasked training target states 218 and another group of training proprioceptions 216(B) are inputted into Oracle Policy 202, and Oracle Policy 202 is executed to generate a corresponding group of reference actions 226. t to generate. The Student Policy 204 is then updated by optimizing the training goals 220 calculated from the training actions 224 and the reference actions 226.

[0078] In one or more embodiments, the Student Policy 204 is trained using the DAgger policy distillation technique based on data set aggregation. During the DAgger distillation process, the training unit 122 first trains the Student Policy 204 with the training proprioceptions 216 and training target states 218 derived from the reference movements 214 in the training data 212 and corresponding reference actions 226 generated by the Oracle Policy 202. After the Student Policy 204 has been trained to output training actions 224 that mimic reference actions 226 associated with reference movements 214, the training unit 122 uses a weighted combination of the Oracle Policy 202 and the Student Policy 204 to capture new trajectories of training proprioceptions 216 and training target states 218.Training Unit 122 additionally uses Oracle Policy 202 to generate reference actions 226 for these new trajectories and trains a new version of Student Policy 204 on an aggregation of training proprioceptions 216, training target states 218, and reference actions 226 associated with the reference actions 226 and the new trajectories. Training Unit 122 can repeat the process using additional trajectories taken from the new version of Student Policy 204 until a certain number of versions of Student Policy 204 have been generated and / or another condition is met. This iterative data set aggregation process thus ensures that Student Policy 204 is trained on a distribution of inputs likely to occur during execution by Student Policy 204, rather than on the distribution of inputs generated by Oracle Policy 202.

[0079] Returning to the discussion of Fig. 2: After the training of the Student Policy 204 is complete, the Execution Unit 124 integrates the trained Student Policy 204 into the Controller 250 to generate a new motion 240 for the articulated object based on one or more selected control modes 234 and a control input 232. For example, the Execution Unit 124 can receive control modes 234 in the form of user inputs, configuration parameters associated with the motion 240, and / or other types of data. The Execution Unit 124 can also, or alternatively, determine and / or dynamically update the control modes 234 based on the control input 232, the environment of the articulated object, and / or other factors. The Execution Unit 124 can receive control inputs 232 from an input device and / or another source of control signals for the articulated object.The execution unit 124 can apply mode masks 238, which are assigned to the control modes 234, to the target state 230 to generate a masked target state 230. The execution unit 124 can also determine the proprioception 236 using sensor inputs at the articulated object, a simulation that includes the articulated object, and / or another information source relating to the state of the articulated object. The execution unit 124 can input the masked target state 230 and the proprioception 236 into the controller 250 and use the trained student policy 204 within the controller 250 to generate a corresponding action 252. The execution unit 124 can convert the action 252 into a corresponding movement 240 and use the action 252 to generate an updated proprioception 236 for the next time step.The execution unit 124 can repeat the process for a specified number of time steps until no further control input 232 is provided, until the control input 232 indicates that the generated movement 240 is complete, and / or until another condition is met. While the execution unit 124 generates additional actions and / or movements, it can vary the types of control modes 234 and / or control input 232 used to generate the actions and / or movements to reflect different types of tasks (e.g., bipedal locomotion, bimanual manipulation, teleoperation, etc.) performed via the actions and / or movements.

[0080] The execution unit 124 can also integrate any generated movement 240 into various applications. For example, the execution unit 124 can simulate the articulated object performing a movement 240 within a game, video, virtual world, visualization, and / or other environment. The execution unit 124 can also, or alternatively, generate commands that cause a robot corresponding to the articulated object to perform a movement 240.

[0081] Fig. Figure 4 illustrates the exemplary operation of the execution unit 124. Fig. 1 in the generation of a movement for an articulated object according to at least one embodiment. As in Fig. As shown in Figure 4, the control input 232 can include a variety of signals specified via a variety of control interfaces. For example, the control input 232 can include (but is not limited to): head and hand poses specified by a VR controller; a full-body pose determined by an image, video, and / or other visual representation of a person (or other type of articulated object); full-body joint positions derived from an exoskeleton; upper-body joint positions derived from a robotic arm; a full-body posture specified by a motion-sensing system; and / or center-of-body commands received from a joystick, keyboard, game controller, and / or other type of input device.

[0082] The control input 232 is used to determine different representations of the target state 230(A), 230(B), and / or 230(C) for the articulated object. For example, the control input 232 can be combined with proprioceptive 236 data for the articulated object to determine desired key point positions in target state 230(A), desired joint angles in target state 230(B), and / or desired attributes of the body center in target state 230(C).

[0083] Mode masks 238(A), 238(B), and / or 238(C) can also be applied to each target state 230(A), 230(B), and / or 230(C) to generate a masked target state that is input into the controller 250. For example, mode mask 238(A) can be used to include a torso position and a left ankle position from target state 230(A) in the masked target state and / or to filter out other targeted key point positions in target state 230(A) from the masked target state. Mode mask 238(B) can be used to select a subset of joint angles from target state 230(B) for inclusion in the masked target state.The mode masks 238(C) can be used to select a desired velocity of the body center from the target state 230(C) for inclusion in the masked target state and / or to exclude other desired attributes of the body center in the target state 230(C) from the masked target state.

[0084] Based on the entered masked target state and proprioception 236 (not in Fig. (As shown in Figure 4), the controller 250 generates a corresponding action 252 and / or converts the action 252 into a movement 240. For example, the controller 250 can use the trained student policy 204 to predict the action 252 as desired joint positions and / or other representations of a state to be achieved. The controller 250 can also use a PD controller and / or another type of control system to generate a corresponding actuation signal that causes the movement 240.

[0085] Now, with reference to Fig. Each block of the Procedure 500 described herein comprises a computational process that can be performed using any combination of hardware, firmware, and / or software. Various functions can be performed, for example, by a processor executing instructions stored in memory. The procedures can also be embodied as computer-usable instructions stored on computer storage media. The procedures can be provided by a standalone application, a service, a hosted service (standalone or in combination with another hosted service), or a plug-in for another product, to name just a few. Furthermore, Procedure 500 is illustrated by a simulated example with respect to the systems of Fig. 1-2 described. However, these procedures can additionally or alternatively be performed by any system or combination of systems, including, but not limited to, the systems described herein. Furthermore, the operations in Procedure 500 can be omitted, repeated, and / or performed in any order without deviating from the scope of this disclosure.

[0086] Fig. Figure 5 illustrates a flowchart of a method 500 for training a machine learning model within the framework of a unified neural motion control according to at least one embodiment. As shown in Fig. As shown in Figure 5, Procedure 500 begins with Operation 502, in which Training Unit 122 determines training data that includes one or more reference movements, training proprioceptions, and / or target training states. For example, Training Unit 122 may generate the reference movements using a motion capture technique, a pose estimation technique, a motion transfer technique, and / or another technique applied to "real-world" movements of humans and / or other types of articulated objects. Training Unit 122 may use simulation techniques to generate a training proprioception for a specific time step in each reference movement, representing a current state of the articulated object.Training unit 122 can additionally generate a training target state for the same time step, which includes a unified representation of key point positions, joint angles, body center attributes and / or other attributes to be achieved.

[0087] In Operation 504, Training Unit 122 trains an Oracle Policy using one or more reference movements, one or more training proprioceptions, and / or one or more training target states. The Oracle Policy can, for example, include a machine learning program (MLP), a recurrent neural network (RNN), a convolutional neural network (CNN), a graph neural network (GNN), and / or another type of machine learning model. For each time step of a training episode associated with a specific reference movement, Training Unit 122 can input a corresponding training proprioception and training target state into the Oracle Policy. Training Unit 122 can then execute the Oracle Policy to generate a training action based on the input data.Training Unit 122 can compute one or more rewards from the training action, one or more training proprioceptions associated with the generated training action, and / or one or more training target states associated with the generated training action. Training Unit 122 can additionally use PPO and / or another optimization technique to update Oracle Policy parameters based on a cumulative discounted reward calculated from multiple rewards associated with the training episode.

[0088] In Operation 506, Training Unit 122 determines whether to continue training the Oracle Policy. For example, Training Unit 122 may determine that training should continue until one or more conditions are met. These conditions include (but are not limited to) convergence of the Oracle Policy parameters, an increase in the calculated reward(s) above one or more thresholds, and / or a certain number of training steps, iterations, episodes, and / or epochs. As long as training is to continue, Training Unit 122 repeats one or more iterations of Operations 502, 504, and 506. For example, Training Unit 122 may use Operations 502 and 504 to train the Oracle Policy to generate actions that replicate and / or mimic additional reference movements in the training data.Training Unit 122 can also perform Operation 506 after a certain number of training steps, episodes, batches and / or epochs to determine whether or not to continue training the Oracle Policy.

[0089] Once Training Unit 122 determines that Oracle Policy training should not continue, it executes Operation 508, in which it determines additional training data. This data may include one or more reference actions from the Oracle Policy, additional training proprioceptions, and / or additional training target states. For example, Training Unit 122 can use the Oracle Policy to generate the reference action(s) for one or more time steps of a training episode associated with a specific reference movement. Training Unit 122 can also generate any additional training proprioception as a history of the recent states and / or actions associated with a Student Policy up to a specific time step.Training Unit 122 can also generate any additional training target state as a unified representation of key point positions, joint angles, body center attributes and / or other attributes to be achieved in a corresponding time step.

[0090] In Operation 510, the training unit 122 applies mode masks and sparsity masks to the one or more additional training target states to generate one or more corresponding masked target states. For example, the training unit 122 can determine one or more control modes associated with one or more parts of the articulated object by sampling from one or more distributions and / or by using another technique. The training unit 122 can use mode masks associated with the one or more control modes to remove attributes not associated with the one or more control modes from the one or more additional training target states.Training Unit 122 can also determine a set of sparsity masks for remaining attributes associated with one or more control modes by sampling from one or more distributions and / or by using another technique. Training Unit 122 can then use the sparsity masks to filter out a subset of the remaining attributes from the one or more masked target states.

[0091] In Operation 512, Training Unit 122 trains a Student Policy using one or more reference actions, one or more additional training proprioceptions, and / or one or more masked target states. The Student Policy can, for example, be an MLP, a recurrent neural network (RNN), a convolutional neural network (CNN), a graph neural network (GNN), and / or another type of machine learning model. For a time step of a training episode associated with a specific reference movement, Training Unit 122 can input a corresponding additional training proprioception and a corresponding additional training target state into the Student Policy. Training Unit 122 can then execute the Student Policy to generate a training action based on the input data.Training Unit 122 can calculate one or more losses using the generated training action and a corresponding reference action. Training Unit 122 can additionally use PPO and / or another optimization technique to update parameters of the student policy based on the one or more losses.

[0092] In Operation 514, Training Unit 122 determines whether to continue training the Student Policy. For example, Training Unit 122 may determine that training should continue until one or more conditions are met. These conditions include (but are not limited to) the convergence of the Student Policy parameters, an increase in the calculated reward(s) above one or more thresholds, and / or a certain number of training steps, iterations, episodes, and / or epochs. As long as training is to continue, Training Unit 122 repeats one or more iterations of Operations 508, 510, 512, and 514. For example, Training Unit 122 may use Operations 508, 510, and 512 to train the Student Policy to generate actions that replicate and / or mimic the reference actions.Training Unit 122 can also perform Operation 514 after a certain number of training steps, episodes, batches and / or epochs to determine whether or not to continue training the Student Policy.

[0093] Once the training unit 122 determines that the training of the student policy should not continue, the execution unit 124 performs operation 516, in which the execution unit 124 generates a new movement for the virtual character based on a control input and / or one or more control modes by executing the trained student policy. For example, the execution unit 124 can receive the control input and / or the one or more control modes from a user and / or via a control interface. The execution unit 124 can use the control input and / or the mode masks associated with the one or more control modes to generate proprioception and a set of target state attributes for a start-time step.Execution Unit 124 can input proprioception and target state attributes into the trained Student Policy and use the trained Student Policy to generate a corresponding action. Execution Unit 124 can also use a PD controller and / or another type of control system to generate a movement from the action. Execution Unit 124 can also use the action, the movement, the control input, and / or one or more control modes to generate updated proprioception and a set of target state attributes for the next time step. Execution Unit 124 can repeat the process for a specified number of time steps until no further control input is received, until control input and / or another type of signal indicating that the generated movement is complete is received, and / or until another condition is met.

[0094] In summary, the disclosed techniques train and execute a machine learning model to perform neural motion control within a unified command space. For example, the machine learning model can be used to output actions that control the whole-body movement of a humanoid robot and / or other type of articulated object.The actions can be generated within various control modes and associated command spaces, such as kinematic position tracking, where desired positions of important rigid body points in the articulated object are tracked; joint angle tracking, where desired angles for joints and / or motors in the joints of the articulated object are tracked; and / or center of mass tracking, where a desired velocity, height, orientation and / or other attribute of a center of mass position in the articulated object is tracked (but are not limited to these).

[0095] The machine learning model can correspond to a student policy (learning strategy; also: learning guideline) that is trained using a policy distillation technique to output actions consistent with those of an oracle policy (reference strategy; also: reference guideline). The oracle policy can comprise a different machine learning model trained to mimic extensive human motion data (e.g., from a motion capture dataset). Motor skills related to balance, coordination, and / or motion control can thus be learned from the oracle policy and transferred to the student policy, enabling it to adapt to different or multiple control modes and generalize across various types of tasks involving these motor skills.

[0096] The input to the student policy includes the proprioception associated with the articulated object and a representation of a target state. The proprioception can include (but is not limited to) a history of joint positions, joint velocities, base angular velocities, actions, and / or other states over a specified number of recent time steps. The target state can include (but is not limited to) changes to be induced in positions, orientations, linear velocities, angular velocities, and / or other attributes during the current time step. During training of the control strategy to be learned, mode masks are used, which assign one or more command modes to one or more parts of the articulated object (e.g.,(corresponding to the upper and lower body of a humanoid robot) to remove attributes not associated with one or more command modes from the target state. A sparsity mask is also used to selectively omit a subset of the remaining attributes from the target state, resulting in a masked target state that serves as input to the control strategy being learned. The student policy is executed to generate an action for the current time step, and the student policy is trained with a loss calculated using the generated action and a reference action generated by the oracle policy for the same time step.

[0097] After the student policy has been trained, it is integrated into a controller that generates movements for the articulated object based on control inputs assigned to different control modes. For example, the trained student policy can be used to generate actions that allow the articulated object to smoothly switch between control modes and / or tasks, use different control modes with different parts of the articulated object, and / or otherwise generate movements according to multiple control modes.

[0098] One advantage of the revealed techniques over previous approaches is the ability to generate actions in the articulated object using a single machine learning model, supporting multiple control modes and / or control interfaces. Consequently, the machine learning model can perform a variety of motion control tasks more effectively than conventional techniques that adapt individual neural motion control models to different control modes and / or interfaces. Furthermore, because the machine learning model learns motor skills shared across all control modes, the movements generated by the machine learning model may exhibit fewer errors and / or improved performance compared to machine learning models trained to process specific control modes.Furthermore, the ability of the machine learning model to operate within different combinations of control modes and / or control interfaces can reduce the time and resource expenditure compared to conventional approaches where multiple machine learning models are trained and deployed to handle different types of tasks and / or inputs.

[0099] The systems and procedures described herein can be used without restriction by non-autonomous vehicles or machines, semi-autonomous vehicles or machines (e.g., in one or more adaptive driver assistance systems (ADAS)), autonomous vehicles or machines, steered and unsteered robots or robotic platforms, warehouse vehicles, all-terrain vehicles, vehicles coupled with one or more trailers, hydrofoils, watercraft, shuttles (e.g., robotaxis), emergency vehicles, motorcycles, electric or motorized bicycles, aircraft, construction vehicles, underwater vehicles (e.g., steered or unsteered submarines), drones, and / or other types of vehicles.Furthermore, the systems and methods described herein can be used for a number of purposes, including but not limited to machine control, machine locomotion, machine propulsion, synthetic data generation, model training, perception, augmented reality, virtual reality, mixed reality, robotics, security and surveillance, simulation and digital twinning, autonomous or semi-autonomous machine applications, deep learning, environmental simulation, object or actor simulation and / or digital twinning, data center processing, conversational AI, light transport simulation (e.g., ray tracing, path tracing, etc.), collaborative content creation for 3D assets (e.g., NVIDIA's Omniverse), cloud computing, and / or other suitable applications.

[0100] Disclosed embodiments may be included in a variety of different systems, such as automotive systems (e.g., a control system for an autonomous or semi-autonomous machine, a perception system for an autonomous or semi-autonomous machine, etc.).), systems implemented with a robot, aviation systems, media systems, boat systems, intelligent area surveillance systems, systems for performing deep learning operations, systems for performing simulation operations, systems for performing digital twin operations, systems implemented using an edge device, systems implementing language models such as large language models (LLMs), vision language models (VLMs) and / or multimodal language models, systems that use or employ one or more inference microservices, systems that implement one or more machine learning models in a service or microservice together with an operating system-level virtualization package (e.g.a container) or use systems that include or utilize one or more virtual machines (VMs), systems for performing operations to generate synthetic data, systems that are at least partially implemented in a data center, systems for performing conversational AI operations, systems for performing light transport simulations, systems for performing collaborative content creation for 3D assets, systems for performing generative AI operations, systems that are at least partially implemented using cloud computing resources, and / or other types of systems. EXEMPLARY AUTONOMOUS OR SEMI-AUTOMATIC MACHINE

[0101] Fig. Figure 6A is an example of sensor positions with corresponding fields of view or sensor fields for an autonomous or semi-autonomous vehicle 600a, an autonomous mobile robot (AMR) 600b, and a humanoid robot 600c according to at least some embodiments of the present disclosure. Although three types of machines 600 are illustrated, this is not intended to be a limitation, and the one or more machines 600 described herein may include: a passenger vehicle, a car, a truck, a bus, an emergency service vehicle, a shuttle, an electric or motorized bicycle, a motorcycle, a fire engine, a police or emergency vehicle, an ambulance, a watercraft, a construction vehicle, an underwater vehicle, a robot (e.g., AMR, humanoid, robotic arm, end effector, forklift, etc.), a drone, an aircraft, a vehicle coupled with a trailer (e.g., a trailer ...a semi-trailer truck used for transporting cargo), and / or another type of vehicle (e.g., one that is unmanned and / or carries one or more passengers). The vehicle 600a, the AMR 600b, the humanoid robot 600c, and / or other machine types may, in some cases, be collectively referred to here as Machine 600.

[0102] With regard to the 600A vehicles, autonomous and semi-autonomous vehicles are generally described in terms of levels of automation defined by the National Highway Traffic Safety Administration (NHTSA), a division of the U.S. Department of Transportation, and the Society of Automotive Engineers (SAE) standard "Taxonomy and Definitions for Terms Related to Driving Automation Systems for On-Road Motor Vehicles" (Standard No. J3016-201806, published June 15, 2018, Standard No. J3016-201609, published September 30, 2016, and earlier and future versions of this standard). The 600A vehicle may exhibit functionality corresponding to one or more of the Levels 3 through 5 of autonomous driving levels.The Machine 600 can exhibit functionality corresponding to one or more of the Levels 1 to 5 of autonomous driving. For example, depending on its embodiment, the Machine 600 may be capable of driver assistance (Level 1), partial automation (Level 2, Level 2+, Level 2++), conditional automation (Level 3), high automation (Level 4), and / or full automation (Level 5). The term "autonomous," as used herein, may encompass any and / or all types of autonomy for the Machine 600 or any other machine, such as fully autonomous, highly autonomous, conditionally autonomous, partially autonomous, assistive autonomy, semi-autonomous, primary autonomous, or any other designation.

[0103] In relation to Fig. 6A The sensors and their respective fields of view (not illustrated for clarity) or sensor fields (not illustrated for clarity) are an exemplary embodiment and are not intended to be limiting. Although not illustrated, each sensor may have a corresponding field of view (e.g., a 360-degree field of view of an ambient camera 668D, a 180-degree field of view of a wide-angle camera 668B, a 360-degree sensor field of a LiDAR sensor 664, etc.). For example, only a subset of the illustrated sensors may be included, additional sensors may be included, alternative sensors may be included, the number of each sensor modality may vary, and the sensor modalities may differ (e.g., they may not include LiDAR or radar, they may include SONAR, thermal sensors, etc.).The sensor positions (included in the diagrams) may differ from those shown on the Vehicle 600a, AMR 600b, and / or Humanoid Robot 600c, etc. For example, with respect to the Vehicle 600a, depending on the type (e.g., SUV, truck, sedan, robot, motorcycle, etc.), size (e.g., semi-trailer truck, moving van, small sedan, etc.), and associated functionality (e.g., L2 vs. L5), the positions, numbers, modalities, and / or other sensor information may vary. Similarly, the shape, size, purpose, implementation, model, etc., may dictate the number and type of sensors used for the AMR 600b and / or the Humanoid Robot 600c.

[0104] As in Fig. Figure 6A illustrates that the autonomous or semi-autonomous vehicle 600A, the AMR 600B and the humanoid robot 600C can have different sensor types, numbers and positions. As a non-restrictive example, the vehicle 600A can contain twelve cameras 668, such as a front wide-angle camera (e.g., with a field of view (FOV) of 120 degrees), a front telephoto camera (e.g., FOV of 30 degrees), a rear left side camera (e.g., FOV of 70 degrees), a rear right side camera (e.g., FOV of 70 degrees), a front fisheye camera (e.g., FOV of 200 degrees), a rear fisheye camera (e.g., FOV of 200 degrees), a left fisheye camera (e.g., FOV of 200 degrees), a right fisheye camera (e.g., FOV of 200 degrees), a front telephoto satellite camera (e.g., FOV of 30 degrees), a rear telephoto camera (e.g., FOV of 30 degrees), and a left side camera (e.g., FOV of 120 degrees). and a horizontal camera on the right (e.g., FOV of 120 degrees).In certain embodiments, one or more cameras 668 can use a Gigabit Multimedia Serial Link (GMSL) interface - such as GMSL2 - as input / output (I / O).

[0105] In some embodiments, which, however, are not in Fig. As illustrated in Figure 6A, the vehicle 600A may include a system for monitoring the occupants and / or the driver inside the vehicle, which may contain various sensors. For example, the interior sensors may include various cameras 668, such as a driver monitoring camera (e.g., with a 55-degree field of view, positioned in front of and aimed at the driver's seat), a front-mounted occupant monitoring camera (e.g., with a 190-degree field of view, positioned in front of and aimed at the front seats), and a rear-mounted occupant monitoring camera (e.g., with a 190-degree field of view, positioned in front of and aimed at the rear seats). Similar to the one or more outward-facing cameras 668, the one or more interior cameras 668 may, in embodiments, use a GMSL interface (such as GMSL2) for I / O.

[0106] As a further, non-restrictive example, the vehicle 600A can also contain nine RADAR sensors 660. For example, the vehicle 600A can contain a front center imaging RADAR sensor (e.g., FOV or sensor field of 120 degrees), a front left corner RADAR sensor (e.g., FOV or sensor field of 160 degrees), a front right corner RADAR sensor (e.g., FOV or sensor field of 160 degrees), a rear right corner RADAR sensor (e.g., FOV or sensor field of 160 degrees), a left side RADAR sensor (e.g., FOV or sensor field of 160 degrees), a right side RADAR sensor (e.g., FOV or sensor field of 160 degrees), a rear left RADAR sensor (e.g., FOV or sensor field of 50 degrees), and a rear right RADAR sensor (e.g., FOV or sensor field of 50 degrees). In some embodiments, one or more RADAR sensors 660 can use an Ethernet interface as I / O.

[0107] The one or more vehicles 600A can furthermore, as a non-limiting example, contain twelve ultrasonic sensors 662. As in Fig. As illustrated in Figure 6A, the ultrasonic sensors can be positioned along the front and rear bumpers of the vehicle 600A and along the side of the vehicle 600A and can be used to detect (static and dynamic) objects in the immediate vicinity of the vehicle 600A. In some embodiments, one or more ultrasonic sensors 662 can use a DS13 interface as I / O.

[0108] The one or more vehicles 600A may further, as a non-limiting example, include a LiDAR sensor 664, such as a front center LiDAR sensor (e.g., with a horizontal FOV or sensor field of 120 degrees and a vertical FOV or sensor field of 30 degrees). In some embodiments, for example, when additional or alternative LiDAR sensors are used, the LiDAR sensor may have different horizontal and vertical fields of view or sensor fields. For example, a LiDAR sensor 664 may have a horizontal FOV or sensor field of 360 degrees (as with a rotating LiDAR sensor) and a vertical FOV or sensor field of 90 degrees. In some embodiments, the one or more LiDAR sensors 664 may use an Ethernet interface as I / O.

[0109] The autonomous mobile robot (AMR) 600B can, as a non-restrictive example, incorporate three LiDAR sensors 664. For instance, the uppermost LiDAR sensor 664 shown can be a beam or 3D LiDAR sensor (e.g., with a horizontal or vertical FOV or sensor field of 360 degrees or 90 degrees, respectively), and the front and rear LiDAR sensors can be planar or 2D LiDAR sensors (e.g., with a horizontal FOV or sensor field of 180 degrees).

[0110] The AMR 600B can further, in a non-restrictive embodiment, include eight cameras 668, such as a front stereo camera (e.g., FOV of 120 degrees), a rear stereo camera (e.g., FOV of 120 degrees), a left stereo camera (e.g., FOV of 120 degrees), a right stereo camera (e.g., FOV of 120 degrees), a front fisheye camera (e.g., FOV of 202 degrees ± 3 degrees), a rear fisheye camera (e.g., FOV of 202 degrees ± 3 degrees), a left fisheye camera (e.g., FOV of 202 degrees ± 3 degrees), and a right fisheye camera (e.g., FOV of 202 degrees ± 3 degrees).

[0111] The AMR 600B can further include a charging port, charging port contacts, a status indicator light, one or more (e.g., four) RGB LEDs, one or more IMU sensors 666, a magnetometer, and a barometer. The AMR 600B is capable of highly accurate time synchronization between sensors using hardware timestamps and PTP over Ethernet with a sensor acquisition time of less than 10 microseconds. In certain embodiments, the AMR 600B enables simultaneous camera recording from all cameras 668 within 100 microseconds of a single hardware trigger and can write to disk at 4 GB / second (e.g., to ROSbags for the Robot Operating System (ROS)).Thus, the AMR 600B is able to run the ROS (such as NVIDIA's Isaac ROS), can be teleoperated (as described here), can map an environment, and can navigate within an environment using visual cameras 668, LiDAR sensors 664, and / or other sensor types or modalities.

[0112] The humanoid robot 600C can, as a non-restrictive example, include a LiDAR sensor 664. For example, the LiDAR sensor 664 can include a beam or 3D LiDAR sensor (e.g., with a horizontal or vertical FOV or sensor field of 360 degrees or 90 degrees, respectively), or it can include a planar or 2D LiDAR sensor (e.g., with a horizontal FOV or sensor field of 180 degrees).

[0113] The humanoid robot 600C can further, as a non-restrictive embodiment, include four cameras 668, such as a front stereo camera (e.g., FOV of 120 degrees), a rear stereo camera (e.g., FOV of 120 degrees), a front fisheye camera (e.g., FOV of 202 degrees ± 3 degrees), and a rear fisheye camera (e.g., FOV of 202 degrees ± 3 degrees).

[0114] The humanoid robot 600C can further, as a non-restrictive embodiment, include four ultrasonic sensors 662, such as an ultrasonic sensor for the left arm, an ultrasonic sensor for the right arm, an ultrasonic sensor for the left leg and an ultrasonic sensor for the right leg.

[0115] The 600C humanoid robot can also include any number of actuators, for example, to enable the control and maneuverability of joints. For instance, the 600C humanoid robot can include actuators that allow different degrees of freedom (DoF), depending on the design. In a non-restrictive embodiment, the 600C humanoid robot can have a total of 40 degrees of freedom (DoF) (e.g., 6 DoF x 2 for the arms, 6 DoF x 2 for the hands, 6 DoF x 2 for the legs, 2 DoF for the torso, and 2 DoF for the neck). The actuators can convert energy into physical motion, thus enabling actions such as joint movements, locomotion, and grasping / manipulation. For example, joint movements can be performed using motors and servos to control the rotation of joints in an arm or manipulator, enabling the reaching, grasping, and manipulation of objects.Locomotion can be achieved using wheels, tracks, or other propulsion devices (robot legs) to move within the environment. Grasping and manipulation can be accomplished using end effectors or hands / fingers, which may be equipped with actuators to grasp objects, exert force, and perform specific tasks. In some examples, the 600C humanoid robot may include position and orientation sensors such as encoders, gyroscopes, and the like to determine the robot's position in space, enabling location tracking and motion tracking. The 600C humanoid robot may also include force and pressure sensors to detect interactions with the environment, allowing the robot to grasp objects with the appropriate force and avoid obstacles in its path. Perception sensors (e.g., cameras, LiDARs, radars, ultrasound, sonar, etc.) are also included.These sensors can be used in conjunction with tactile sensors so that the 600C robot can perceive objects, shapes, and textures, and understand when a touch begins and ends (along with force sensors that regulate the force applied during the touch). As a non-limiting example, the 600C humanoid robot could be approximately 1 to 2 meters tall (e.g., 1.7 meters or 5' 6") and weigh 50 to 70 kg, be capable of moving at speeds of 8 km / h or more, and, depending on the design and system requirements, be able to carry payloads of 20 to 100 kg.

[0116] The 600C humanoid robot can, in certain configurations, incorporate a dialogue system—for example, a dialogue system based on language models (such as LLMs, VLMs, MMLMs, VLAs, etc.)—to understand its environment, reason, and communicate with humans, animals, devices, and / or other robots, and / or to make planning, control, and navigation decisions. Thus, in addition to performing various tasks, the 600C humanoid robot can utilize integrated sensors, microphones, and speakers to understand speech, acoustic and visual signals, and so on, while simultaneously communicating with its environment.

[0117] With reference to the cameras 668 of the one or more machines 600, the camera types for the cameras 668 can include, but are not limited to, digital cameras designed for use with the components and / or systems of the machine 600. In an implementation of the vehicle 600a, the one or more cameras 668 can operate at Automotive Safety Integrity Level (ASIL) B and / or at another ASIL. Depending on the embodiment, the camera types can be capable of any frame rate, such as 30 frames per second (fps), 60 fps, 120 fps, 240 fps, etc. The cameras can use roller shutters, global shutters, another type of shutter, or a combination thereof.In some examples, the color filter array may include a red-clear-clear-clear (RCCC) color filter array, a red-clear-clear-blue (RCCB) color filter array, a red-blue-green (RBGC) color filter array, a Foveon X3 color filter array, a Bayer sensor color filter array (RGGB), a monochrome sensor color filter array, and / or another type of color filter array. In some embodiments, cameras with clear pixels, such as cameras with an RCCC, RCCB, and / or RBGC color filter array, may be used to increase light sensitivity.

[0118] Cameras with a field of view that includes portions of the environment in front of the Machine 600 (e.g., forward-facing cameras) can be used for environmental viewing to help identify forward paths and obstacles, and to provide, with the help of one or more Controller 636 units and / or control SoCs, information critical for creating an occupancy grid and / or determining preferred machine movements, trajectories, and / or paths. Forward-facing cameras can be used to perform many of the same ADAS functions as LiDAR, including emergency braking, pedestrian detection, and collision avoidance.Forward-facing cameras can also be used for ADAS functions and systems, including lane departure warnings (LDW), autonomous cruise control (ACC) and / or other functions such as traffic sign recognition.

[0119] A variety of cameras can be used in a forward-facing configuration, including, for example, a monocular camera platform containing a complementary metal oxide semiconductor (CMOS) color imager. Another example is one or more 668B wide-angle cameras, which can be used to detect objects moving into the field of view from the periphery (e.g., pedestrians, warehouse vehicles, other robots, crossing vehicles, or bicycles). In addition, any number of 668E long-range cameras (e.g., a pair of long-range stereo cameras) can be used for depth-based object detection, particularly for objects for which a neural network has not yet been trained. The one or more 668E long-range cameras can also be used for object detection and classification, as well as basic object tracking.

[0120] Any number of stereo cameras 668A can also be included in a forward-facing and / or other (e.g., rear-facing) configuration. In at least one embodiment, one or more of the stereo cameras 668A can include an integrated control unit comprising a scalable processing unit that can provide programmable logic (“FPGA”) and a multicore microprocessor with an integrated controller area network (“CAN”) or Ethernet interface on a single chip. Such a unit can be used to create a 3D map of the environment of the machine 600 that includes a distance estimate for points in the image (e.g., a disparity image or depth image).One or more alternative stereo cameras 668A can include one or more compact stereo vision sensors, which may contain two camera lenses (one left and one right) and an image processing chip that can measure the distance between the vehicle and the target object and use the generated information (e.g., metadata) to activate the autonomous emergency braking and lane departure warning functions. Other types of stereo cameras 668A can be used in addition to or as alternatives to those described here. For example, in some embodiments, stereo depth estimation can be performed using cameras other than stereo cameras, such as two monocular cameras with at least partially overlapping fields of view.

[0121] Cameras with a field of view that includes parts of the environment to the sides of the machine 600 (e.g., side cameras) can be used, for example, for all-around vision and provide information that can be used to create and update the occupancy grid, as well as to generate side-impact collision warnings and / or, for example, to indicate to an AMR 600B or humanoid robot 600C that objects, features, and / or people are located to the side. For example, one or more environment cameras 668D can be positioned on the machine 600. The one or more environment cameras 668D can include one or more wide-angle cameras 668B, one or more fisheye cameras, one or more 360-degree cameras, and / or the like. For example, four fisheye cameras can be mounted on the front, rear, and sides of the machine 600. In an alternative arrangement, the machine 600 can have three environment cameras 668D (e.g.,use (left, right and rear) and use one or more other cameras (e.g. a front-facing camera) as a fourth surround-view camera.

[0122] Cameras 668 with a field of view that includes parts of the environment behind the machine 600 (e.g., reversing cameras) can be used to capture objects, features, people, and / or other information behind the machine 600, for example, for parking assistance, surround view, rear-impact warnings, planning, control, and navigation decisions, and / or for creating and updating an occupancy grid, a BEV image of the environment, a height map, etc. A variety of cameras 668 can be used, including, but not limited to, cameras 668 that are also suitable as one or more front cameras (e.g., one or more long-range and / or medium-range cameras 668E, one or more stereo cameras 668A), one or more infrared cameras 668C, etc.), one or more rear-facing cameras, one or more side-facing cameras, one or more downward-facing cameras, one or more upward-facing cameras and / or the like, as described herein.

[0123] Similarly, for LiDAR sensors 664, RADAR sensors 660, ultrasonic sensors 662 and / or other sensor modalities or types, the position and arrangement of the sensors and their corresponding fields of view or sensor fields can be determined depending on the application, implementation or design of the respective machine 600.

[0124] The one or more machines 600 contain, for example, one or more RADAR sensors 660, which are used by the machine 600 for long-range object detection, even in darkness and / or adverse weather conditions. The functional safety level of the RADAR can be ASIL B in some embodiments. The one or more RADAR sensors 660 can use CAN and / or bus 602 (e.g., for transmitting the data generated by the one or more RADAR sensors 660) for control and access to object tracking data, with access to the raw data in some examples occurring via Ethernet. A variety of RADAR sensor types can be used. The one or more RADAR sensors 660 can be suitable, for example, for front, rear, and side RADAR applications without restriction. In some examples, one or more pulse-Doppler RADAR sensors are used.

[0125] The single or multiple RADAR 660 sensors can incorporate various configurations, such as long-range with a narrow field of view, short-range with a wide field of view, side coverage with short-range, and so on. In some examples, long-range RADAR can be used for adaptive cruise control (ACC). Long-range RADAR systems can provide a wide field of view, achieved through two or more independent scans, for example, at a range of 250 m. The single or multiple RADAR 660 sensors can assist in distinguishing between static and moving objects and can be used by ADAS systems for emergency braking assistance and forward collision warning, and by robots for detecting dynamic objects in various environments—such as those with low or no lighting.Long-range radar sensors can incorporate a monostatic multimodal radar with multiple (e.g., six or more) fixed radar antennas and a high-speed CAN and FlexRay interface. In a six-antenna example, the four central antennas can generate a focused beam pattern designed to cover the machine's surroundings at higher speeds with minimal peripheral interference (e.g., from traffic in adjacent lanes). The other two antennas can extend the field of view, enabling the rapid detection of objects entering or leaving the machine's immediate path (e.g., the lane).

[0126] Medium-range radar systems can, for example, have a range of up to 660 m (front) or 80 m (rear) and a field of view of up to 42 degrees (front) or 150 degrees (rear). Short-range radar systems can include, among other things, radar sensors designed to be installed at both ends of a lateral surface (e.g., a rear bumper), allowing two beams to be used to continuously monitor the blind spot behind and beside the vehicle (e.g., vehicle, robot, etc.). Short-range radar systems can therefore be used in an ADAS system for blind spot detection and / or as a lane change assistant.

[0127] The machine 600 can also include one or more ultrasonic sensors 662. The one or more ultrasonic sensors 662, which can be positioned at the front, rear, and / or sides of the machine 600, can be used to support close-range sensing, for example, for parking assistance, collision avoidance (e.g., for robot parts), and / or for creating and updating an occupancy grid, an evidence grid map (EGM), a height map, a BEV image, and / or other representation of objects and features in the machine 600's environment. A variety of ultrasonic sensors 662 can be used, and different ultrasonic sensors 662 can be used for different detection ranges (e.g., 2.5 m, 4 m). The one or more ultrasonic sensors 662 can, for example, operate at functional safety levels of ASIL B.

[0128] The machine 600 can include one or more LiDAR sensors 664. The LiDAR sensor(s) 664 can be used for object and feature detection, pedestrian and other robot detection, emergency braking, collision avoidance, simultaneous localization and mapping (SLAM), free space detection, and / or other functions. The one or more LiDAR sensors 664 can operate in embodiments with functional safety levels of ASIL B. In some examples, the machine 600 can include multiple LiDAR sensors 664 (e.g., two, four, six, etc.) that can use Ethernet (e.g., to deliver data to a Gigabit Ethernet switch).

[0129] In some examples, one or more LiDAR sensors 664 may be able to provide a list of objects and their distances for a 360-degree field of view. Commercially available LiDAR sensors 664 may, for example, have a specified range of approximately 600 m, with an accuracy of 2 cm to 3 cm and support for a 600 Mbit / s Ethernet connection. In some examples, one or more non-protruding LiDAR sensors 664 may be used. In such examples, the one or more LiDAR sensors 664 may be implemented as a small device that can be embedded in the front, rear, sides, and / or corners of the machine 600. In such examples, one or more LiDAR 664 sensors can provide a horizontal field of view of up to 120 degrees and a vertical field of view of up to 35 degrees, with a range of 200 m, even with objects of low reflectivity.The one or more front-mounted LiDAR 664 sensors can be configured for a horizontal field of view between 45 degrees and 135 degrees.

[0130] In some examples, LiDAR technologies, such as 3D flash LiDAR, can also be used. 3D flash LiDAR uses a laser pulse as a transmission source to illuminate the vehicle's surroundings up to approximately 200 m. A flash LiDAR unit contains a sensor that records the travel time of the laser pulse and the reflected light at each pixel, which in turn corresponds to the distance between the vehicle and the objects. Flash LiDAR can enable the generation of highly accurate and distortion-free images of the surroundings with each laser pulse. In some examples, four flash LiDAR sensors can be used, one on each side of the vehicle. Available 3D flash LiDAR systems include a solid-state 3D focal plane array LiDAR camera that contains no moving parts other than a fan (e.g., a non-scanning LiDAR device).The flash LiDAR device can use a 5-nanosecond pulse of a Class I (eye-safe) laser per frame and capture the reflected laser light in the form of 3D distance point clouds and co-registered intensity data. By using flash LiDAR, and because flash LiDAR is a solid-state device with no moving parts, the single or multiple LiDAR sensors can be less susceptible to motion blur, vibration, and / or shock.

[0131] Fig. Figure 6B illustrates the sensor and component positions of an exemplary autonomous or semi-autonomous vehicle 600A (hereafter referred to alternatively as "Vehicle 600", "Ego-Vehicle 600", "Ego-Machine 600", or "Machine 600") according to some embodiments of the present disclosure. Although Vehicle 600A is illustrated, this is not intended to be a limitation, and similar components and / or sensors may be included in any other type of machine without departing from the scope of the present disclosure. Similar sensors and / or components may be used, for example, in: a passenger vehicle, a car, a truck, a bus, an emergency service vehicle, a shuttle, an electric or motorized bicycle, a motorcycle, a fire engine, a police vehicle, an ambulance, a watercraft, a construction vehicle, an underwater vehicle, a robot (e.g., AMR, humanoid, robotic arm, end effector, forklift, etc.).), a drone, an aircraft, a vehicle coupled to a trailer (e.g. a semi-trailer truck used for transporting cargo), and / or another type of vehicle or machine (e.g. one that is unmanned and / or carries one or more passengers).

[0132] Fig. Figure 6C is a block diagram of an exemplary system architecture for a machine 600, such as an autonomous or semi-autonomous vehicle 600A, an autonomous mobile robot (AMR) 600B, a humanoid robot 600C, and / or other types of machines according to some embodiments of the present disclosure. It is understood that these and other arrangements described herein are presented only as examples. Other arrangements, components, features, and elements (e.g., machines, interfaces, functions, sequences, groupings of functions, etc.) may be used in addition to or instead of those shown, and some elements may be omitted entirely. Furthermore, many of the arrangements, components, features, elements, etc., described herein are functional units that can be implemented as separate or distributed components or in conjunction with other components and in any suitable combination and location (e.g.,on a local device, in a vehicle or machine at the system edge, on-site—for example, on locally hosted servers, remote servers—for example, in one or more computing or server devices in one or more data centers, in the cloud, and / or at other locations. Various functions described herein, performed by entities, can be executed by hardware, firmware, and / or software. For example, various functions can be performed using one or more processors (e.g., central processing units (CPUs), graphics processing units (GPUs), microprocessors, microcontrollers, embedded processors, digital signal processors (DSPs), image signal processors (ISPs), physics accelerators (PPUs), field-programmable gate arrays (FPGAs), accelerators (e.g.,Deep learning accelerators (DLAs), deep learning accelerator clusters (XNNs), neural network accelerators (NNAs) and / or neural processing units (NPUs), programmable vision accelerators (PVAs), optical flow accelerators (OFAs), etc.), application-specific integrated circuits (ASICs), data processing units (DPUs), quantum processors, etc.) are executed, carrying out instructions stored in memory. In some embodiments, the systems, methods, and processes described herein can be executed using components, features, and / or functionalities similar to those of the exemplary machine 600 from [reference missing]. Fig. 6A-6E, the exemplary 700-series computer ecosystem Fig. 7, the exemplary generative language model system 800 from Fig. 8 and / or the exemplary calculating device 900 from Fig. 9 are the same.

[0133] Each of the components, features and each of the systems of the Machine 600 in Fig. 6C is illustrated as being connected via a 602 bus (alternatively referred to as the "602 machine communication network" or simply the "602 communication network"). The 602 bus may contain a Controller Area Network (CAN) data interface (alternatively referred to here as the "CAN bus"). A CAN bus may be a network within the 600 machine that serves to support the control of various features and functions of the 600 machine, such as the actuation of brakes, acceleration, braking, steering, windshield wipers, etc. A CAN bus may be configured to have dozens or even hundreds of nodes, each with its own unique identifier (e.g., a CAN ID). The CAN bus can be read to determine the steering wheel angle, vehicle speed, engine speed (rpm), button positions, and / or other vehicle status indicators. The CAN bus may be ASIL B compliant.In some embodiments, the 602 bus can include, in addition to or as an alternative to, a CAN bus, FlexRay, an embedded bus (e.g., SPI, I2C), a local interconnect network (LIN), NVIDIA NVLink, USB (2.0, 3.0, higher), radio frequency (RF), Ethernet (e.g., 10BASE / 100BASE, 1000BASE, 10G, etc.), and / or another communication protocol or functionality. Furthermore, while a single line is used to represent the 602 bus, this is not intended as a limitation. For example, there can be any number of 602 buses, which may contain one or more CAN buses, one or more FlexRay buses, one or more Ethernet buses, and / or one or more other types of buses using a different protocol. In some examples, two or more buses 602 can be used to perform different functions, and / or they can be used for redundancy.For example, a first bus 602 can be used for collision avoidance functionality and a second bus 602 for actuation control. In each example, each bus 602 can communicate with one of the components of the machine 600, and two or more buses 602 can communicate with the same components. In some examples, each SoC 604, each controller 636, and / or each computer or computer engine within the machine 600 can have access to the same input data (e.g., inputs from sensors of the machine 600) and be connected via a common bus, such as a CAN bus.

[0134] The machine 600 can include components such as a chassis, a vehicle body, wheels (e.g., 2, 4, 6, 8, 18, etc.), tires, axles, batteries, side mirrors, and / or other components of a vehicle or machine. The machine 600 can include a drive system 650, such as an internal combustion engine, a hybrid electric power plant, a pure electric motor, a hydrogen-powered engine, and / or another type of drive system. The drive system 650 can be connected to a drivetrain of the machine 600, which may include a transmission to enable the propulsion of the machine 600. The drive system 650 can be controlled in response to signals received from the throttle or accelerator device 652.

[0135] A steering system 654, which may include a steering wheel and / or other steering device (e.g., remote steering and / or local steering), can be used to steer the machine 600 (e.g., along a desired path or route) when the drive system 650 is in operation (e.g., when the vehicle is in motion). The steering system 654 can receive signals from a steering actuator 656. In some embodiments, a steering wheel or other steering mechanism may not be included, such as in a machine 600 that enables full automation functionality (e.g., Level 5).

[0136] The brake sensor system 646 can be used to actuate the vehicle brakes in response to receiving signals from the brake actuators 648 and / or the brake sensors.

[0137] The machine 600 can contain one or more controllers 636, as described here in relation to Fig. 6A. The one or more Controllers 636 can be used for a variety of functions and can be coupled with any other components and systems of the Machine 600. For example, the Controllers 636 can be used to control the Machine 600, to run artificial intelligence on the Machine 600, for infotainment for the Machine 600, and / or the like. For example, one Controller 636 can be used for some or all functionalities, or different Controllers 636 can be used for different functionalities—for example, to ensure availability and a secure separation between different controllers for different tasks. For example, the one or more Controllers 636 can use plans calculated by the system—for example,Paths or trajectories for the vehicles 600A or AMRs 600B, or movements, component trajectories, motion positions or displacements, etc., for joints or components (e.g., of manipulators, end effectors, limbs, hands, fingers, legs, feet, etc.) of a humanoid robot 600C, to control the one or more machines 600 in the environment. In some cases, the one or more controllers 636 may include a proportional-integral-differential (PID) controller, a fuzzy logic controller, a neural controller (e.g., a controller implemented as one or more neural networks), a force control controller, a programmable logic controller (PLC), and / or another type of controller.In a 600C humanoid robot, for example, one or more 636 controllers can function as the brain, responsible for analyzing sensor data, making decisions, and sending commands to the actuators. The one or more 636 controllers can include a lower-level controller that handles basic motor control and ensures accurate and precise movements of individual joints and actuators. The one or more 636 controllers can also include an upper-level controller that coordinates multiple actuators and sensors, plans complex movements, and adapts to changing environments.

[0138] The one or more Controller 636 can, in various embodiments, include an artificial intelligence controller that can use AI algorithms (e.g., DNNs, MLMs, etc.) to learn, make decisions, and autonomously perform tasks for the Machine 600. In some embodiments, the one or more Controller 636 can use an open control algorithm that is fixed and does not adapt actions to the environment. In other embodiments, a closed control with feedback mechanisms can be used to monitor the robot's performance and make necessary adjustments. The one or more Controller 636 can implement reactive control to respond directly to sensor inputs, enabling rapid reflexes and real-time changes.Furthermore, in some examples deliberative control can be implemented, using internal models and planning algorithms to generate higher-level actions suitable for complex tasks requiring logical reasoning, decision-making, and long-term planning.

[0139] The one or more controllers 636, which are one or more systems-on-chips (SoCs) 604 ( Fig. 6C and Fig. 6D), CPUs, GPUs, accelerators, etc., can provide signals (e.g., representative of instructions or messages) to one or more components and / or systems of the machine 600. Although the one or more controllers 636 are listed separately from the one or more SoCs 604, this is not intended to be a limitation, and in some embodiments, one or more components of the one or more SoCs 604 can perform the operations of the controllers 636. For example, the one or more controllers can send signals to actuate the machine brakes via one or more brake actuators 648, to actuate the steering system 654 via one or more steering actuators 656, to actuate the propulsion system 650 via one or more throttle / accelerator 652, etc. The one or more controllers 636 can include one or more integrated (e.g., built-in) computing devices (e.g.,Supercomputers) that process sensor signals and issue operating commands (e.g., command-representing signals) to enable autonomous or semi-autonomous navigation and movement and / or to assist a human operator in using the machine 600. The one or more controllers 636 may include a first controller 636 for autonomous driving and navigation functions, a second controller 636 for functional safety functions, a third controller 636 for artificial intelligence functions (e.g., computer vision), a fourth controller 636 for infotainment functions, a fifth controller 636 for emergency redundancy, and / or other controllers.For example, the hardware used for safety monitoring and other safety functions (such as a functional safety island) can be discrete or partitioned (physically or by separation of processing) with respect to the hardware used for processing sensor data for perception and making decisions for vehicle control. Similarly, hardware (e.g., a controller, a SOC, etc.) for controlling the vehicle's infotainment system and / or interior monitoring can be discrete or separate from the hardware used for vehicle perception and control. In some examples, a single Controller 636 can perform two or more of the functionalities mentioned above, two or more Controller 636s can perform a single functionality, and / or any combination thereof.

[0140] The one or more controllers 636 can provide the signals for controlling one or more components and / or systems of the machine 600 in response to sensor data received from one or more sensors (e.g., sensor inputs). The sensor data can be received, for example, without limitation, from one or more of the following: one or more global navigation satellite systems (GNSS) sensors 658 (e.g., one or more global positioning system sensors), one or more radar sensors 660, one or more ultrasonic sensors 662, one or more LiDAR sensors 664, one or more inertial measurement units (IMUs) 666 (e.g., accelerometer(s), gyroscope(s), magnetic compass(s), magnetometer(s), etc.), one or more microphones 696, one or more cameras 668 (e.g.,one or more stereo cameras 668A, one or more wide-angle cameras 668B (e.g., fisheye cameras), one or more infrared cameras 668C, one or more surround-view cameras 668D (e.g., 360-degree cameras), one or more long-range and / or medium-range cameras 668E and / or other camera types), one or more speed sensors 644 (e.g., for measuring the speed of the machine 600), one or more vibration sensors 642, one or more steering sensors 640, one or more brake sensors (e.g., as part of the brake sensor system 646), actuators and / or other sensor types.

[0141] One or more of the controllers 636 can receive inputs (e.g., in the form of input data) from an instrument cluster 632 of the machine 600 and provide outputs (e.g., in the form of output data, display data, etc.) via a human-machine interface (HMI) display 634 (e.g., screen, field-of-view display, mirror display, face display, robot display, etc.), an acoustic signal, a loudspeaker, a transducer, and / or via other components of the machine 600. The outputs can include information such as machine speed, rotational speed, time, map data from one or more cards 622 of Fig. 6C) correspond (e.g., from a navigation chart, a standard-resolution map, a high-resolution (High Definition, "HD") map, etc.), location data (e.g., the location of machine 600, for example, on a map 622), direction, location of other vehicles (e.g., an occupancy map, a height map, a bird's-eye view (BEV) image, a grid, etc.), information about objects and the status of objects as perceived by the system, information about the system status, etc. For example, one or more HMI displays 634 can show information about the presence of one or more objects (e.g., a road sign, a warning sign, a changing traffic light, etc.) and / or information about driving maneuvers that the vehicle has performed, is currently performing, or will perform (e.g., changing lanes now, taking exit 34B in two miles, etc.).

[0142] The Machine 600 can contain one or more systems on a chip (SoCs) 604 (described in more detail in Fig. 6D). The one or more SoCs 604 can contain one or more CPUs 606, one or more GPUs 608, one or more processors 610, one or more caches 612, one or more accelerators 614, one or more data storage devices 616, and / or other components and features. The one or more SoCs 604 can be used to process and provide data for various operations, such as navigation, planning, reasoning, inference, perception, control, and / or actuation operations of the machine 600 in a variety of platforms and systems. For example, in addition to map data corresponding to one or more maps 622 (e.g., HD map, SD card, navigation map, occupancy map, etc.), the one or more SoCs 604 can also provide live perception data (e.g., from camera, LiDAR, radar, ultrasound, etc.).) process data to perform or assist various operations of the machine 600. When a map and / or AI is used, the map and / or AI (e.g., model parameter updates, fine-tuning, etc.) are accessed via a network interface 624 from one or more servers (e.g., the servers 678). Fig. 6E) - for example, one or more servers in a cloud-based data center - updated and / or renewed.

[0143] Although one or more SoCs 604 are consistently in Fig. As illustrated in Figures 6A-6E, additional or alternative components and / or architectures may be used—such as multi-chip modules (MCMs), application-specific integrated circuits (ASICs), system-in-packages (SiPs), field-programmable gate arrays (FPGAs), heterogeneous integration (HI), and single-board computers (SBCs)—without deviating from the scope of this disclosure. For example, depending on the type of Machine 600, the use of the Machine 600, the model of the Machine 600, and the required capabilities of the Machine 600, one or more SoCs 604 and / or alternative architectures and / or components may be used to fulfill the respective implementation.

[0144] The Machine 600 can contain one or more CPUs 618 (e.g., one or more discrete CPUs or one or more dCPUs) that can be coupled to the one or more SoCs 604 via a high-speed connection (e.g., PCIe). The one or more CPUs 618 can, for example, contain an x86 processor. The one or more CPUs 618 can be used, for example, to perform a variety of functions, including reconciling potentially inconsistent results between ADAS sensors and the one or more SoCs 604 and / or monitoring the status and health of the one or more Controllers 636 and / or the Infotainment SoC 630.

[0145] The Machine 600 can contain one or more GPUs 620 (e.g., one or more discrete GPUs or one or more dGPUs) that can be coupled to the one or more SoCs 604 via a high-speed connection (e.g., NVIDIA's NVLink). The one or more GPUs 620 can provide additional artificial intelligence capabilities, such as running redundant and / or distinct neural networks, and can be used to train and / or update neural networks based on input (e.g., sensor data) from sensors in the Machine 600.

[0146] The machine 600 can also include the network interface 624, which can contain one or more wireless antennas 626 and / or modems (e.g., one or more wireless antennas for different communication protocols, such as a cellular antenna, a Bluetooth antenna, etc.). The network interface 624 can be used to enable a wireless connection via the internet to the cloud (e.g., to one or more servers 678 and / or other network devices), to other vehicles, and / or to computing devices (e.g., passenger client devices). To communicate with other vehicles, a direct connection between the two vehicles and / or an indirect connection (e.g., via networks and the internet) can be established. Direct connections can be established via vehicle-to-vehicle communication.Vehicle-to-vehicle communication can provide the Machine 600 with information about vehicles in its vicinity (e.g., vehicles in front of, beside, and / or behind the Machine 600). This functionality can be part of a cooperative adaptive speed control function of the Machine 600.

[0147] The 624 network interface can include a system-on-a-chip (SoC) that provides modulation and demodulation functions, enabling one or more 636 controllers to communicate over wireless networks. The 624 network interface can include a high-frequency (RF) front end for up-conversion from baseband to RF and down-conversion from RF to baseband. The frequency conversions can be performed using known methods and / or superheterodyne techniques. In some examples, the RF front-end functionality can be provided by a separate chip.The network interface 624 can, for example, be capable of communication via Long-Term Evolution (LTE), Wideband Code Division Multiple Access (WCDMA), Universal Mobile Telecommunications System (UMTS), Global System for Mobile Communication (GSM), IMT-CDMA Multi-Carrier (CDMA2000), fifth generation mobile communication technology (5G), sixth generation mobile communication technology (6G), and / or other mobile communication and / or wireless communication standards. The one or more wireless antennas 626 can also enable communication between objects in the environment (e.g., vehicles, mobile devices, etc.) using local area networks such as Bluetooth, Bluetooth Low Energy (LE), Z-Wave, ZigBee, etc.and / or enable low power wide area networks (LPWANs), such as LoRaWAN, SigFox, etc.

[0148] The machine 600 may further include one or more data memories 628, which may be located outside the chip (e.g., outside the SoCs 604). The one or more data memories 628 may contain one or more memory elements, including RAM, SRAM, DRAM, VRAM, flash, hard disks, and / or other components and / or devices capable of storing at least one bit of data.

[0149] The Machine 600 can also include one or more GNSS sensors 658. The one or more GNSS sensors 658 (e.g., GPS, supported GPS sensors, differential GPS (DGPS) sensors, etc.) assist with mapping, perception, occupancy grid creation, and / or path planning. Any number of GNSS sensors 658 can be used, including, for example, and without limitation, a GPS unit that uses a USB connection with an Ethernet-to-serial (RS-232) bridge.

[0150] The machine 600 may further include one or more IMU sensors 666. In some examples, the one or more IMU sensors 666 may be located in the center of the rear axis of the machine 600. The one or more IMU sensors 666 may, for example, and without limitation, include one or more accelerometers, one or more magnetometers, one or more gyroscopes, one or more magnetic compasses, and / or other sensor types. In some examples, such as six-axis applications, the one or more IMU sensors 666 may include accelerometers and gyroscopes, while in nine-axis applications, the one or more IMU sensors 666 may include accelerometers, gyroscopes, and magnetometers.

[0151] In some embodiments, the one or more IMU sensors 666 can be implemented as a miniaturized, high-performance GPS-based inertial navigation system (GPS / INS) that combines inertial sensors of a microelectromechanical system (MEMS), a highly sensitive GPS receiver, and advanced Kalman filter algorithms to provide estimates of position, velocity, and orientation. Thus, in some examples, the one or more IMU sensors 666 can enable the machine 600 to estimate its course without requiring input from a magnetic sensor by directly observing and correlating velocity changes from the GPS with the one or more IMU sensors 666. In some examples, the one or more IMU sensors 666 and the one or more GNSS sensors 658 can be combined in a single integrated unit.

[0152] The vehicle may contain one or more microphones 696, which are mounted in and / or around the machine 600. The one or more microphones 696 may be used, among other things, for the detection and identification of emergency vehicles.

[0153] The machine 600 can further include one or more vibration sensors 642. The one or more vibration sensors 642 can measure vibrations of the machine's components, such as the arms or legs of a humanoid robot 600C or the one or more axles of a vehicle 600A or an AMR 600B. For example, changes in vibrations can indicate a change in the road, sidewalk, or drivable surface. In another example, if two or more vibration sensors 642 are used, the differences between the vibrations can be used to determine the friction or slippage on the road surface (e.g., if the difference in vibration is between a driven axle and a freely rotating axle).

[0154] Machine 600 can contain an ADAS system 638, for example, if Machine 600 is a vehicle 600A. In some examples, the ADAS system 638 can contain one or more dedicated SoCs.The ADAS system 638 may include autonomous / adaptive / automatic cruise control (ACC), cooperative adaptive cruise control (CACC), forward crash or collision warning (FCW), automatic emergency braking (AEB), lane departure warning (LDW), lane keep assist (LKA), blind spot warning (BSW), blind spot monitoring (BSM), rear cross-traffic warning (RCTW), pedestrian detection, driver monitoring, collision warning systems (CWS), traffic sign recognition, speed limit recognition, automatic parking, lane centering (LC), high beam safety system, and / or other features and functions.

[0155] The Machine 600 may also include the Infotainment SoC 630 (e.g., an in-vehicle infotainment system (IVI)). Although illustrated and described as an SoC, the infotainment system may not actually be an SoC and may include one or more discrete components, such as multi-chip modules (MCMs), application-specific integrated circuits (ASICs), system-in-packages (SiPs), heterogeneous integration (HI), single-board computers (SBCs), etc. The Infotainment SoC 630 may include a combination of hardware and software that can be used to provide the Machine 600 with audio (e.g., music, a personal digital assistant, navigation directions, news, radio, etc.), video (e.g., TV, movies, streaming, etc.), telephone (e.g., hands-free calling), network connectivity (e.g., LTE, Wi-Fi, etc.), and / or information services (e.g.,Navigation systems, rear parking assistance, a radio data system, vehicle-related information such as fuel level, total distance traveled, brake fluid level, oil level, door open / close status, air filter information, etc.). The Infotainment SoC 630 can, for example, include radios, turntables, navigation systems, video players, USB and Bluetooth connectivity, car computers, in-car entertainment, Wi-Fi, steering wheel audio controls, hands-free calling, a head-up display (HUD), an HMI display 634, a telematics device, a control panel (e.g., for controlling and / or interacting with various components, functions, and / or systems), and / or other components. The Infotainment SoC 630 can also be used to provide information (e.g.,to provide visual and / or acoustic information such as information from the ADAS system 638, autonomous driving information such as planned vehicle maneuvers, road layouts, environmental information (e.g. intersection information, vehicle information, road information, etc.), and / or other information.

[0156] The Infotainment SoC 630 can include GPU functionality. The Infotainment SoC 630 can communicate with other devices, systems, and / or components of the Machine 600 via the 602 bus (e.g., CAN bus, Ethernet, etc.). In some examples, the Infotainment SoC 630 can be coupled with a monitoring MCU so that the Infotainment System's GPU can perform some self-driving functions if one or more of the primary controllers 636 (e.g., the primary and / or backup computers of the Machine 600) fail. In such an example, the Infotainment SoC 630 can put the Machine 600 into a chauffeur-to-safe-stop mode, as described here.

[0157] In some implementations, the infotainment system can provide a digital or virtual assistant, which may be voice-controlled only or may include a visual component (e.g., in the form of a digital person or a digital avatar). The assistant can provide basic functions such as sending text messages, adjusting vehicle settings, controlling music or video, navigation, etc., and / or provide advanced functions as supported by one or more language models—such as large language models (LLMs), vision language models (VLMs), multimodal language models (MMLMs), etc. For example, the driver and / or passengers can interact with the assistant much like a user interacts with a language model, e.g., by...to ask general or specific questions, to request recommendations and / or locations for restaurants, gas stations, and / or other amenities, to inquire about vehicle functions, or to troubleshoot problems (e.g., to request information about tire pressure, oil changes, battery replacements, etc.). Thus, the Machine 600—be it a Vehicle 600A, AMR 600B, Humanoid Robot 600C, and / or another type of machine—can contain one or more locally stored language models and / or communicate with a remotely hosted language model (e.g., via one or more APIs) to provide users of the one or more Machine 600s with more detailed and comprehensive communication features.

[0158] In some examples, an infotainment SoC 630, one or more SoCs 604, and / or another SoC or computer / processing system can perform monitoring of the driver and / or occupants inside the vehicle. For example, the computer system can perform facial recognition, and the vehicle owner identification system can use data from cameras and / or other sensors to identify the presence of an authorized driver and / or owner of the Machine 600. The always-on sensor processing unit can be used to unlock the vehicle when the owner approaches the driver's door and turn on the lights, and to disable the vehicle in security mode when the owner leaves the vehicle. In this way, the one or more SoCs 604 provide security against theft and / or carjacking.

[0159] In some embodiments, an interior surveillance camera sensor can be monitored by one or more neural networks running on a separate or dedicated SoC—for example, an SoC for the vehicle's infotainment or surveillance systems—configured to identify and respond to events within the vehicle. An in-cabin system can perform lip-reading to activate cellular service and make a call, dictate emails, change the destination, activate or modify the vehicle's infotainment system and settings, or enable voice-activated web browsing. The in-cabin system can also include one or more AI agents or assistants that can use one or more APIs or plug-ins to interact with one or more LLMs, VLMs, MMLMs, etc., in the cloud.For example, the AI ​​agents or assistants in the interior can provide directions, vehicle or machine feedback information, answer general questions, handle music / video and / or other requests, activate windows, doors and / or other vehicle components, etc. Therefore, one or more dedicated SoCs and / or processor sets can be used to perform the infotainment and / or interior monitoring (e.g., as an occupant monitoring system (OMS)) for the Machine 600.

[0160] The Machine 600 may further include an Instrument Cluster 632 (e.g., a digital instrument cluster, an electronic instrument cluster, a digital instrument panel, etc.). The Instrument Cluster 632 may include a controller and / or supercomputer (e.g., a discrete controller or supercomputer). The Instrument Cluster 632 may contain a number of instruments, such as a speedometer, fuel gauge, oil pressure gauge, tachometer, odometer, turn signals, shift indicator, seatbelt warning light(s), parking brake warning light(s), engine malfunction light(s), airbag system (SRS) information, lighting controls, safety system controls, navigation information, etc. In some examples, information from the Infotainment SoC 630 and the Instrument Cluster 632 may be displayed and / or shared. In other words, the Instrument Cluster 632 may be included as part of the Infotainment SoC 630, or vice versa.

[0161] Fig. 6D is a block diagram of an exemplary architecture of a computer system (a subset of the architecture of 6D in relation to 6D). Fig. 6C system) according to at least some embodiments of the present disclosure. Although illustrated as one or more SoCs 604, this is not intended to be a limitation and the computer system may additionally or instead include multi-chip modules (MCMs), application-specific integrated circuits (ASICs), system-in-packages (SiPs), heterogeneous integration (HI), single-board computers (SBCs) and / or other components and / or architectures without deviating from the scope of the present disclosure.

[0162] The single or multiple 604 SoCs can form an end-to-end platform with a flexible architecture covering automation levels 2-5, or they can be specifically designed for a particular automation level (e.g., a first 604 SoC for level 2 to level 2++, a second 604 SoC for level 3, a third 604 SoC for level 4, etc.). This provides a comprehensive functional safety architecture that leverages and efficiently utilizes computer vision, neural network inference, robot planning, control and navigation, ADAS techniques, and similar technologies with diversity and redundancy to provide a platform for a flexible, reliable software stack for vehicle or robot control, along with deep learning tools. The single or multiple 604 SoCs can be faster, more reliable, and even more energy-efficient and compact than conventional systems.For example, the one or more accelerators 614 in combination with the one or more CPUs 606, the one or more GPUs 608 and the one or more data storage devices 616 can form a fast, efficient platform for autonomous vehicles of levels 2-5 as well as for the safe planning, navigation and control of the AMRs 600B, humanoid robots 600C and / or other robot or machine types.

[0163] In some embodiments, such as when one or more SoCs 604 include a GPU 608 with 2000 or more cores (e.g., 2048 cores), 60 or more tensor cores (e.g., 64 tensor cores), and a maximum GPU frequency above 1 GHz (e.g., 1.3 GHz), a CPU 606 with 10 or more cores (e.g., 12 cores), with 64-bit memory, 3 MB L2 and 6 MB L3 cache, and a maximum frequency of 2 GHz or more (e.g., 2.2 GHz), one or more deep learning accelerators (DLAs), deep learning accelerator clusters (XNNs), neural network accelerators (NNAs), or neural processing units (NPUs) 609 (e.g., 2 DLAs / XNNs / NNAs / NPUs (609) and a vision accelerator—such as a programmable vision accelerator (PVA) (607), a single SoC (604)—can achieve an AI performance of 275 teraoperations per second (TOPS). For example, NVIDIA's Jetson AGX Orin 64 GB SoC meets these criteria and achieves this performance.

[0164] Likewise, in embodiments where one or more SoCs 604 include a GPU 608 with 1700 or more cores (e.g., 1792 cores), 50 or more tensor cores (e.g., 64 tensor cores), and a maximum GPU frequency exceeding 900 MHz (e.g., 930 MHz), a CPU 606 with 8 or more cores (e.g., 8 cores), with 64-bit architecture, 2 MB L2 and 4 MB L3 cache memory, and a maximum frequency of 2 GHz or more (e.g., 2.2 GHz), one or more deep learning accelerators (DLAs), deep learning accelerator clusters (XNNs), neural network accelerators (NNAs), or neural processing units (NPUs) 609 (e.g., 2 DLAs / XNNs / NNAs / NPUs (609) and a vision accelerator—such as a programmable vision accelerator (PVA) (607), a single SoC (604)—can achieve an AI performance of 200 teraoperations per second (TOPS). For example, NVIDIA's Jetson AGX Orin 32 GB SoC meets these criteria and achieves this performance.

[0165] In some embodiments, such as when one or more SoCs 604 include a GPU 608 with 1000 or more cores (e.g., 1024 cores), 28 or more tensor cores (e.g., 64 tensor cores), and a maximum GPU frequency exceeding 900 MHz (e.g., 1173 MHz), a CPU 606 with 8 or more cores (e.g., 8 cores), with 64-bit architecture, 2 MB L2 and 4 MB L3 cache memory, and a maximum frequency of 2 GHz or more (e.g., 2 GHz), one or more deep learning accelerators (DLAs), deep learning accelerator clusters (XNNs), neural network accelerators (NNAs), or neural processing units (NPUs) 609 (e.g., 1 DLA / XNN / NNA / NPU) may be included. 609) and a vision accelerator - such as a programmable vision accelerator (PVA) 607, a single SoC 604) can achieve an AI performance of 157 teraoperations per second (TOPS). For example, NVIDIA's Jetson AGX Orin NX 16 GB SoC meets these criteria and achieves this performance.

[0166] In various configurations, such as when one or more SoCs 604 incorporate a GPU 608 with 1000 or more cores (e.g., 1024 cores), 28 or more tensor cores (e.g., 64 tensor cores), and a maximum GPU frequency exceeding 900 MHz (e.g., 1020 MHz), a CPU 606 with 6 or more cores (e.g., 6 cores), 64-bit architecture, 1.5 MB L2 and 4 MB L3 cache, and a maximum frequency of 1.5 GHz or higher (e.g., 1.7 GHz), a single SoC 604 can achieve an AI performance of 67 teraoperations per second (TOPS). For example, NVIDIA's Jetson Orin Nano 8GB SoC meets these criteria and achieves this performance.

[0167] The one or more SoCs 604 can contain one or more CPUs 606. In some embodiments, the one or more CPUs 606 can contain a CPU cluster or CPU complex (here alternatively referred to as "CCPLEX"). The one or more CPUs 606 can contain multiple cores and / or caches (e.g., L2, L3). In some embodiments, the one or more CPUs 606 can, for example, contain twelve cores in a coherent multiprocessor configuration. In some embodiments, the one or more CPUs 606 can contain four dual-core clusters, each cluster having a dedicated L2 cache (e.g., a 3 MB L2 cache). The one or more CPUs 606 (e.g., the CCPLEX) can be configured to support the concurrent operation of clusters, so that any combination of clusters of the one or more CPUs 606 can be active at any given time.

[0168] The one or more SoCs 604 can contain any type and number of GPUs 608. For example, in some embodiments, one or more integrated GPUs (alternatively referred to herein as one or more "iGPUs") can be used. The one or more GPUs 608 can be programmable and can be efficient for parallel workloads. The one or more GPUs 608 can use an extended Tensor instruction set in some examples. The one or more GPUs 608 can contain one or more streaming microprocessors, each streaming microprocessor being able to contain an L1 cache (for example, an L1 cache with a minimum of 96 KB of memory), and two or more of the streaming microprocessors being able to share an L2 cache (for example, an L2 cache with a minimum of 512 KB of memory). In some embodiments, the one or more GPUs 608 can contain at least eight streaming microprocessors.The one or more GPUs 608 can use one or more application programming interfaces (APIs) for computations. Furthermore, the one or more GPUs 608 can use one or more parallel computing platforms and / or programming models (e.g., NVIDIA's CUDA).

[0169] The one or more GPUs 608 can be power-optimized for best performance in automotive, robotics, and / or other embedded applications. The one or more GPUs 608 can be manufactured, for example, on a FinFET field-effect transistor. However, this is not a limitation, and the one or more GPUs 608 can also be manufactured using other semiconductor or fabrication processes. Each streaming microprocessor can contain an array of mixed-precision processing cores, divided into multiple blocks. For example, and without limitation, 64 PF32 cores and 32 PF64 cores can be divided into four processing blocks. In such an example, each processing block could have 16 FP32 cores, 8 FP64 cores, 16 INT32 cores, two mixed-precision NVIDIA TENSOR COREs for deep learning matrix arithmetic, an instruction cache (e.g., L0), a warp scheduler, a dispatch unit, and / or a (e.g., a 12-bit) 12-bit ...B. a 64 KB register file. Furthermore, the streaming microprocessors can include independent parallel integer and floating-point data paths to enable efficient execution of workloads with a mix of computations and addressing calculations. The streaming microprocessors can include an independent thread scheduling function to enable fine-grained synchronization and cooperation between parallel threads. The streaming microprocessors can include a combined L1 data cache and a shared memory unit to improve performance while simplifying programming.

[0170] The one or more GPUs 608 can include high-bandwidth memory (HBM) and / or a 16 GB HBM2 subsystem to provide a peak memory bandwidth of approximately 900 GB / second in some examples. In some examples, synchronous graphics random access memory (SGRAM), such as double-data-rate type five synchronous graphics random access memory (GDDR5), can be used in addition to or as an alternative to HBM memory.

[0171] The one or more GPUs 608 can incorporate a unified memory technology that includes access counters to enable more accurate migration of memory pages to the processor that accesses them most frequently, thereby improving the efficiency of memory areas shared by processors. In some examples, support for Address Translation Services (ATS) can be used so that the one or more GPUs 608 can directly access the page tables of the one or more CPUs 606. In such examples, if the Memory Management Unit (MMU) of the one or more GPUs 608 fails, an address translation request can be sent to the one or more CPUs 606.In response, the one or more CPUs 606 can search their page tables for the virtual-physical mapping for the address and send the translation back to the one or more GPUs 608. In this way, the unified memory technology enables a single, unified virtual address space for the memory of both the one or more CPUs 606 and the one or more GPUs 608, thereby simplifying the programming of the one or more GPUs 608 and the porting of applications to the one or more GPUs 608.

[0172] The one or more SoCs 604 can contain any number of cache(s) 612, including those described here. For example, the one or more caches 612 can include L0 caches, L1 caches, L2 caches, L3 caches (e.g., those available to both the one or more CPUs 606 and the one or more GPUs 608 (e.g., those associated with both the one or more CPUs 606 and the one or more GPUs 608)), etc. The one or more caches 612 can include a write-back cache capable of tracking the states of the rows, for example, by using a cache coherence protocol (e.g., MEI, MESI, MSI, etc.). The (e.g., L3) cache can be 4 MB or larger, depending on the embodiment, although smaller cache sizes can also be used.

[0173] The one or more SoCs 604 can contain one or more arithmetic logic units (ALUs) 665, which can be used in performing processing related to one of the many tasks or operations of the machine 600—such as computer vision, machine learning or deep learning processing, world model management, etc. Furthermore, the one or more SoCs 604 can contain one or more floating-point units (FPUs) 667—or other types of mathematical or numeric coprocessors—for performing mathematical operations within the system. For example, the one or more SoCs 604 can contain one or more FPUs 667, which are integrated as execution units into one or more CPUs 606 and / or one or more GPUs 608.

[0174] The one or more SoCs 604 can contain one or more accelerators 614 (e.g., hardware accelerators, software accelerators, or a combination thereof). For example, the one or more SoCs 604 can contain a hardware acceleration cluster, which may include optimized hardware accelerators and / or a large amount of on-chip memory. The large on-chip memory 615 (e.g., 4 MB SRAM, 32 GB and / or 64 GB of 256-bit LPDDR5 at 204.8 GB / s, 8 GB and / or 16 GB of 128-bit LPDDR5 at 102.4 GB / s, and / or other memory types and sizes) can enable the hardware acceleration cluster to accelerate neural network processing, transformer processing, optical flow processing, vision processing, and / or other computations or processing operations.The hardware acceleration cluster can be used to complement one or more GPUs 608 and offload some of the tasks from the one or more GPUs 608 (e.g., to free up more cycles of the one or more GPUs 608 for other tasks). For example, the one or more accelerators 614 can be used for specific workloads (e.g., perception, convolutional neural networks (CNNs), deep neural networks (DNNs), language models (LLMs, VLMs, MMLMs, VLAs, etc.), transformer models, diffusion models, pure encoder models, encoder-decoder models, etc.) that are stable enough to be suitable for acceleration.

[0175] The one or more Accelerators 614 (e.g., the Hardware Acceleration Cluster) can contain a Deep Learning Accelerator (DLA) 609 (alternatively referred to here as "Deep Learning Accelerator Cluster (XNN) 609," "Neural Network Accelerator (NNA) 609," or "Neural Processing Unit (NPU) 609"). The one or more DLAs 609 can contain one or more Tensor Processing Units (TPUs) 641 configured to provide an additional ten trillion operations per second for deep learning applications and inference. The TPUs 641 may be accelerators configured and optimized for performing image processing functions (e.g., for CNNs, RCNNs, etc.).The single or multiple DLAs 609 can also be optimized for a specific set of neural network types and floating-point operations, as well as for inference. The design of the single or multiple DLAs can deliver more performance per millimeter than a general-purpose GPU and far surpasses the performance of a CPU. The single or multiple TPUs 641 can perform multiple functions, including a single-instance convolution function that supports, for example, INT8, INT16, and FP16 data types for both features and weights, as well as post-processing functions.Although the one or more TPUs 641 are described as part of the one or more DLAs 609, this is not to be understood as a limitation, and the one or more TPUs 641 may be contained in one or more additional or alternative accelerators 614 and / or other components and / or be contained as one or more discrete processing components.

[0176] One or more DLAs 609 can quickly and efficiently run neural networks on processed or unprocessed data for a variety of functions, including, but not limited to, the following: object identification and recognition (e.g., vehicles, pedestrians, other robots, road markings, road boundary lines, debris, potholes, boxes, stock items, etc.).) using data from one or more sensor modalities; for distance estimation using data from one or more sensor modalities; for the detection and identification of emergency vehicles using data from microphones and / or image-based sensors; for facial recognition; for pick-and-place operations; for handling operations; for occupant monitoring; for vehicle owner identification; and / or for other interior operations using data from interior camera sensors and / or other sensor types; and / or for safety and / or security-related events, to name just a few.

[0177] The one or more DLAs 609 can perform any function of the one or more GPUs 608, and by using an inference accelerator, a developer can, for example, allocate either the one or more DLAs 609 or the one or more GPUs 608 to each function. For example, the developer can concentrate the processing of CNNs and floating-point operations on the one or more DLAs 609 and leave other functions to the one or more GPUs 608 and / or other accelerators 614. The one or more DLAs 609 can be used to run any type of network to improve control and security; this includes, for example, a neural network that outputs a confidence measure for each object detection.

[0178] The one or more Accelerators 614 (e.g., the Hardware Acceleration Cluster) can include a Programmable Vision Accelerator (PVA) 607, which may alternatively be referred to here as a Computer Vision Accelerator or more generally as a Vision Accelerator. The one or more PVAs 607 can be designed and configured to accelerate computer vision algorithms for Advanced Driver Assistance Systems (ADAS), semi-autonomous driving, autonomous driving, robotics applications, security and surveillance applications, augmented reality (AR), virtual reality (VR), and / or mixed reality (MR) applications, etc. The one or more PVAs 607 can offer a balance between performance and flexibility.Each PVA 607 can, for example, without limitation, contain any number of cores from Reduced Instruction Set Computers (RISCs), Direct Memory Access (DMA) systems, Pixel Processing Engines (PPEs), Vector Processors or Vector Processing Units (VPUs), and / or other components. The PVA engine can include an advanced Very Long Instruction Word (VLIW) digital signal processor and Single Instruction, Multiple Data (SIMD) capabilities. The one or more PVA 607s can be optimized for image processing tasks and accelerating computer vision algorithms.For example, one or more PVAs 607 offer excellent performance with extremely low power consumption and can be used asynchronously and simultaneously with one or more CPUs 606, one or more GPUs 608 and / or other accelerators in the system (e.g. vehicle, robot, etc.) as part of a heterogeneous computing pipeline.

[0179] The one or more PVAs 607 can contain one or more (e.g., two) Vector Processing Subsystems (VPS), each VPS being able to contain one or more Vector Processing Unit cores (VPU cores), one or more Decoupled Look-up Units (DLUTs), one or more Shared Memory or Vector Memory Elements (VMEMs), and one or more Instruction Caches (I-caches). The one or more VPU cores can be the main processing unit and include a computer vision-optimized vector SIMD VLIW DSP 643. The one or more VPU cores can retrieve instructions via the one or more I-caches and access data via the one or more VMEMs. The one or more DLUTs can include a special hardware component that improves the efficiency of parallel look-up operations.For example, the one or more DLUTs enable parallel lookup operations using a single copy of the lookup table by running these lookups in a decoupled pipeline independent of the primary processor pipeline. In this way, the one or more DLUTs minimize or reduce memory usage and improve throughput while avoiding data-dependent memory bank conflicts—ultimately leading to improved overall system performance. The one or more VPU VMEMs can provide local data storage for the VPU, enabling efficient implementation of various image processing and computer vision algorithms. The one or more VPU VMEMs can support access from hosts outside the VPS, such as Direct Memory Access (DMA), and the one or more CPUs 606 (e.g.,ARM Cortex-R5 processor), which facilitates data exchange with the one or more CPUs 606 and other system-level components. The VPU-I cache can supply instruction data to the one or more VPUs as needed, request missing instruction data from system memory, and / or maintain a temporary instruction memory for the VPU. For each VPU task, the one or more CPUs 606 can configure the DMA system, optionally preload the VPU program into the VPU-I cache, and / or start each VPU-DMA pair to process a task. The one or more PVAs 607 can also include L2 SRAM memory shared by one or more (e.g., two) sets of VPS and DMA. In some embodiments, one or more (e.g., two) DMA devices are used to transfer data between external memory, the PVA L2 memory, the VMEMs (e.g.,to move data to the DRAM (in one VPS), the tightly coupled memory (TCM) of the one or more CPUs, the DMA descriptor memory, and / or the configuration registers at the PVA level. In a lightly loaded system, two parallel DMA accesses to the DRAM can achieve a read / write bandwidth of up to 15 GB / s each, and in a heavily loaded system, this bandwidth can reach up to 10 GB / s each. In terms of compute capacity, the INT8 gigamultiplication-accumulation operations per second (GMACs) can be 2048 or more, excluding the DLUT. The FP32 GMACs can include 32 per PVA instance.

[0180] The RISC cores can interact with image sensors (e.g., the image sensors of one of the cameras described here), image signal processors, and / or the like. Each RISC core can contain any amount of memory. Depending on the implementation, the RISC cores can use any number of protocols. In some examples, the RISC cores can run a real-time operating system (RTOS). The RISC cores can be implemented with one or more integrated circuits, application-specific integrated circuits (ASICs), and / or memory devices. For example, the RISC cores can include an instruction cache and / or tightly coupled RAM.

[0181] The DMA system can enable components of the PVA(s) 607 to access the system's main memory independently of the one or more CPUs 606. The DMA can support any number of features that optimize the one or more PVAs 607, including, but not limited to, support for multidimensional addressing and / or circular addressing. In some examples, the DMA can support up to six or more dimensions of addressing, which may include block width, block height, block depth, horizontal block stepping, vertical block stepping, and / or depth stepping.

[0182] The vector processors, or VPUs, can be programmable processors designed to efficiently and flexibly execute programming for computer vision algorithms and provide signal processing capabilities. In some examples, the one or more PVAs 607 can contain a PVA core and two vector processing subsystem partitions. The PVA core can include a processor subsystem, one or more DMA machines (e.g., two DMA machines), and / or other peripheral devices. The vector processing subsystem can act as the primary processing unit of the one or more PVAs 607 and can include one or more vector processing units (VPUs), one or more pixel processing engines (PPEs) – which process a 2D layout of interconnected (e.g.,A VPU core may contain processing elements (for north-south, east-west, and west-west intercommunication), one or more instruction caches, and / or one or more shared or vector memories (e.g., VMEMs). A VPU core may also contain a digital signal processor, such as a single instruction, multiple data (SIMD) and a very long instruction word (VLIW). The combination of SIMD and VLIW can increase throughput and speed.

[0183] In some embodiments, each vector processor can contain an instruction cache and be coupled to dedicated memory. Therefore, in some examples, each vector processor can be configured to operate independently of the others. In other examples, the vector processors contained in one or more specific PVAs 607 can be configured to use data parallelism. For example, in some embodiments, the multiple vector processors contained in one or more single PVAs 607 can execute the same computer vision algorithm, but on different regions of an image.In other examples, the vector processors contained in one or more specific PVAs 607 can simultaneously execute different computer vision algorithms on the same image, or even different algorithms on successive images or sections of an image. Among other things, any number of PVAs 607 can be included in the hardware acceleration cluster, and any number of vector processors can be contained in each of the PVAs. Furthermore, the one or more PVAs 607 can contain additional memory for error-correcting code (ECC) to enhance the overall system security.

[0184] The single or multiple 614 accelerators (e.g., the hardware accelerator cluster) have a wide range of applications for controlling autonomous and semi-autonomous machines. The single or multiple 607 PVAs can be programmable vision accelerators used for critical processing steps in perception, robotics understanding and logic, ADAS, semi-autonomous and autonomous vehicles, and more. The capabilities of the PVA 607 are well-suited for algorithmic domains requiring predictable processing with low power consumption and low latency. In other words, the single or multiple 607 PVAs are well-suited for semi-dense or dense regular computations, even with small datasets, that demand predictable runtimes with low latency and low power consumption.In the context of platforms for autonomous vehicles and robotics, the PVAs 607 are therefore designed to execute classic computer vision algorithms, as they are efficient at object recognition and work with integer mathematics.

[0185] According to one embodiment of the technology, the PVA 607 is used, for example, to perform computer stereo vision. In some examples, a semi-global matching-based algorithm can be used, although this is not intended as a limitation. Many Level 3-5 autonomous driving applications require spontaneous motion estimation or spontaneous stereo matching (e.g., structure of motion, pedestrian detection, lane detection, etc.). One or more PVA 607s can perform computer stereo vision on input from two monocular cameras.

[0186] In some examples, one or more PVAs 607 can be used to perform dense optical flow processing. This involves processing raw radar data (e.g., using a 4D Fast Fourier Transform) to provide processed radar data. In other examples, one or more PVAs 607 are used for time-of-flight depth processing, for example, by processing raw time-of-flight data to deliver processed time-of-flight data.

[0187] Although the VPU(s), DMA(s), RISC core(s), VMEM(s), and decoupled coprocessors (e.g., the DLUT(s)) are described as being contained in the one or more PVAs 607, this is not intended to be a limitation. In some embodiments, these components may be contained in alternative or additional processing components and / or the one or more accelerators 614 and / or as discrete components of the one or more SoCs 604 and / or other computer system architectures.

[0188] In some examples, one or more SoCs 604 can include a real-time ray tracing hardware accelerator (RTA) 651, which can be used to quickly and efficiently determine the positions and extents of objects (e.g., within a world model) for generating real-time visualization simulations, for radar signal interpretation, for sound propagation synthesis and / or analysis, for simulating SONAR, radar, LiDAR, camera, and / or other sensor modalities within a simulation, for general wave propagation simulation, for comparison with LiDAR data for localization purposes, for generating realistic training data for neural network training, and / or for other functions and purposes. In some embodiments, one or more tree traversal units (TTUs) can be used to perform one or more operations related to ray tracing.For example, the machine 600 (or any other machine or device) can be simulated within a simulation environment, and the simulation environment can be generated using one or more light transport simulation algorithms (e.g., ray tracing, path tracing, etc.). These ray tracing algorithms can then be accelerated using a ray tracing accelerator 651 and / or a ray tracing-optimized GPU 608—such as NVIDIA's RTX GPU.

[0189] The one or more Accelerators 614 (e.g., in a hardware acceleration cluster) can contain one or more Optical Flow Accelerators (OFAs) 611. For example, the one or more OFAS 611 can be used to calculate the optical flow and stereo disparity between individual frames of sensor data (e.g., images). The optical flow can be accelerated on the one or more OFAS 611 for applications such as object detection and tracking and / or for stereo depth estimation, where it is used to calculate the stereo disparity between stereo images (e.g., two or more images acquired with two or more image sensors with at least partially overlapping fields of view).

[0190] The one or more 604 SoCs can include one or more 623 Camera Serial Interfaces (CSIs). For example, the one or more 623 CSIs can also include a Mobile Industry Processor Interface (MIPI) for receiving video and camera input, a high-speed interface, and / or a video input block that can be used for camera and related pixel input functions. The one or more 604 SoCs can also include one or more software-controlled input / output controllers that can be used to receive I / O signals not assigned to a specific role. For example, the 623 CSI can include a MIPI-CSI-2 port—for example, a 16-lane MIPI-CSI-2 port, D-PHY 2.1 (up to 40 Gbps), and C-PHY 2.0 (up to 164 Gbit / s) to support 16 virtual channels and six or more cameras, an 8-lane MIPI CSI-2 connector, D-PHY 2.1 (up to 20 Gbit / s to support 8 virtual channels and 4 or more cameras and / or a 2x MIPI CSI-2, 22-pin camera connector, depending on the design and implementation.

[0191] The one or more Accelerators 614 (e.g., the Hardware Acceleration Cluster) can include a Computer Vision Network on Chip (CVNOC) 663 and SRAM to provide high-bandwidth, low-latency SRAM for the one or more Accelerators 614. In some examples, the on-chip memory can include at least 4 MB of SRAM, consisting, for example, and without limitation, of eight field-configurable memory blocks accessible to the PVA 607, the OFA 611, the DLA 609, and / or other Accelerators 614. Each pair of memory blocks can include an Advanced Peripheral Bus (APB) interface, configuration circuitry, a controller, and a multiplexer. Any type of Memory 615 can be used.The PVA 607, the OFA 611, the DLA 609, and / or one or more other 614 accelerators can access memory via a backbone that enables high-speed memory access for the accelerator(s). The backbone can include an on-chip computer vision network that connects the accelerator(s) to the memory (e.g., using the APB).

[0192] The CVNOC 663 can include an interface that determines, prior to the transmission of control signals / addresses / data, that one or more Accelerators 614 are providing ready and valid signals. Such an interface can provide separate phases and channels for the transmission of control signals / addresses / data, as well as burst communication for continuous data transmission. This type of interface can conform to ISO 26262 or IEC 61508 standards, although other standards and protocols can also be used.

[0193] The one or more SoCs 604 can contain the one or more data stores 616 and / or the memory 615. The one or more data stores 616 can be on-chip memory 615 on the one or more SoCs 604, in which neural networks and / or other algorithms can be stored to run on the one or more CPUs 606, the one or more GPUs 608, and / or one or more accelerators 614. In some examples, the one or more data stores 616 can be large enough to store multiple instances of neural networks for redundancy and security. For example, the one or more data stores 616 can include one or more L2 and / or L3 caches 612. The one or more memory 615 can include SRAM, LPDD5, and / or other memory types.For example, the one or more memory locations 615 may include 4 MB SRAM, 32 GB and / or 64 GB of 256-bit LPDDR5 at 204.8 GB / s, 8 GB and / or 16 GB of 128-bit LPDDR5 at 102.4 GB / s, and / or other memory types and sizes. The reference to the one or more data stores 616 may include a reference to the memory allocated to the PVA 607, the OFA 611, the DLA 609, and / or one or more other accelerators 614, as described herein.

[0194] The one or more data storage devices 616 can comprise various storage types, such as eMMC, NVMe, etc. For example, the one or more SoCs 604 can include storage in the form of an embedded multimedia card (eMMC) (e.g., 64 GB eMMC 5.1) and / or an SD card slot with external NVM Express capability (NVMe), e.g., via M.2 Key M. For example, the one or more data storage devices 616 and / or other storage can be accessed, for example, via NVMe using PCI Express (PCIe), RDMA, TCP, and / or other protocols.

[0195] The one or more SoCs 604 can contain one or more processors 610 (e.g., embedded processors). The one or more processors 610 can contain a boot and power management processor 653, which can be a dedicated processor and subsystem to handle boot power and management functions and the associated security enforcement. The BPMP 653 can be part of the boot sequence of the one or more SoCs 604 and can provide runtime power management services. The BPMP 653 can provide clock and voltage programming, support for system transitions to a low-power state, management of the thermals and temperature sensors of the one or more SoCs 604, and / or management of the power states of the one or more SoCs 604.Each temperature sensor can be implemented as a ring oscillator whose output frequency is proportional to the temperature, and the one or more SoCs 604 can use the ring oscillators to detect the temperatures of the one or more CPUs 606, the one or more GPUs 608, the one or more accelerators 614, and / or other components. If it is determined that the temperatures exceed a threshold, the BPMP 653 can enter a temperature fault routine and put the one or more SoCs 604 into a reduced-power state and / or put the machine 600 into a chauffeur-to-safe-stop mode (e.g., bring the machine 600 to a safe stop).

[0196] The one or more 610 processors can also contain a number of embedded processors that can serve as the 655 Audio Processing Engine (APE). The 655 APE can be an audio subsystem that provides full hardware support for multi-channel audio across multiple interfaces and a wide and flexible range of audio I / O interfaces. In some examples, the 655 APE is a dedicated processor core with a digital signal processor and dedicated RAM.

[0197] The one or more 610 processors can also include an Always-On Processor Unit (AOPE) 657, which can provide the necessary hardware functions to support low-power sensor management and wake-up of use cases. The AOPE 657 can include a processor core, tightly coupled RAM, supporting peripherals (such as timers and interrupt controllers), various I / O controller peripherals, and routing logic.

[0198] The one or more 610 processors can further comprise one or more 613 security processors (alternatively referred to as a "613 security island"), which may contain a security cluster unit. This unit includes a dedicated processor or processor subsystem for security management in automotive, robotics, and / or other applications. The one or more 613 security processors—and / or the security cluster unit—can contain two or more processor cores, tightly coupled RAM, supporting peripherals (such as timers, an interrupt controller, etc.), and / or routing logic. In a security mode, the two or more cores can operate in lockstep mode, functioning as a single core with comparison logic that detects any differences between their operations.In some embodiments, the one or more security processors 613 may comprise one or more discrete processors, so that failures of other system components cannot affect the performance and availability of the security processor 613.

[0199] The one or more processors 610 may further include a real-time or near-real-time sensor engine (SE) 659, which may include a dedicated processor subsystem for processing real-time or near-real-time cameras, LiDAR, RADAR and / or other sensor modalities.

[0200] The one or more processors 610 may further comprise one or more image signal processors (ISPs) 627, which may include a high dynamic range signal processor and / or a hardware unit that is part of one or more sensor processing pipelines.

[0201] The one or more 610 processors can include a 661 Video Image Compositor (VIC), which can be a processing block (e.g., implemented on a microprocessor) that implements video post-processing functions required by a video playback application to produce the final image for the player window. The VIC 661 can perform lens distortion correction on the one or more 668B wide-angle cameras, the one or more 668D ambient cameras, the sensors of the in-cabin surveillance camera, and / or other camera sensors with distorted fields of view.

[0202] A VIC 661 can incorporate enhanced temporal noise reduction for both spatial and temporal noise reduction. For example, if there is motion in a video, the noise reduction weights the spatial information accordingly and reduces the impact of information provided by adjacent frames. If a frame or portion of a frame contains no motion, the temporal noise reduction performed by the video image compositor can use information from the previous frame to reduce noise in the current frame.

[0203] A VIC 661 can also be configured to perform stereo equalization of the input stereo lens images. The video image compositor can also be used for user interface design when the operating system desktop is in use and the one or more GPUs 608 do not need to constantly render new surfaces. Even when the one or more GPUs 608 are powered on and actively performing 3D rendering, the video image compositor can be used to offload the workload from the GPUs 608, thus improving performance and responsiveness.

[0204] The one or more SoCs 604 can also include a wide range of peripheral input / output (I / O) interfaces 625 to enable communication with peripheral devices, audio codecs, power management, and / or other devices. The one or more SoCs 604 can be used to process data from cameras (e.g., via Gigabit Multimedia Serial Link and / or Ethernet), sensors (e.g., one or more LiDAR sensors 664, one or more RADAR sensors 660, etc., which may be connected via Ethernet), data from the bus 602 (e.g., machine speed 600, steering wheel position, etc.), and data from one or more GNSS sensors 658 (e.g., connected via Ethernet or CAN bus).The one or more SoCs 604 may also include dedicated high-performance mass storage controllers, which may contain their own DMA units and which can be used to offload routine data management tasks from the one or more CPUs 606. In some embodiments, the I / O ports 625 of the one or more SoCs 604 may include a header (e.g., a 40-pin header or a 40-pin expansion header) to support the following: Universal Asynchronous Receiver / Transmitter (UART), Serial Peripheral Interface (SPI), Inter-Integrated Circuit Sound (I.C.) bus, and a 40-pin expansion header. 2S), Inter-Integrated Circuit (I2C) bus, Controller Area Network (CAN), Pulse Width Modulation (PWM), Digital Microphone Interface (DMIC), Digital Speaker Station (DSPK), General Purpose I / O (GPIO), etc., an automation connector (e.g., a 12-pin automation connector), an audio panel connector (e.g., a 10-pin audio panel connector), a Joint Test Action Group (JTAG) connector (e.g., a 10-pin JTAG connector), a fan connector (e.g., a 4-pin fan connector), an RTC battery fuse connector (e.g.,a 2-pin battery fuse connector, a microSD slot, a DC power jack, power switch, force button(s), restore button(s) and reset button, one or more display connectors (e.g., DisplayPort (DP), such as DP 1.4A (+MST), eDP 1.41, HDMI 2.1 and / or a 4K30 multi-mode DP 1.2 (+MST) connector) and / or other I / O-625 elements, components or functions.

[0205] The one or more 604 SoCs can include in-machine networking capabilities that utilize, for example, Ethernet (e.g., automotive Ethernet), SERDES, Controller Area Network (CAN), FlexRay, Local Interconnect Network (LIN), Low Voltage Differential Signaling (LVDS), Media Oriented System Transport (MOST), another network type, and / or a combination thereof. For example, the one or more 604 SoCs can include an RJ45 port supporting up to 10 GbE, a 1 GbE port, and / or other types of network connections.

[0206] The one or more SoCs 604 can contain one or more digital signal processors (DSPs) 643. For example, the one or more DSPs 643 can contain a dedicated or specialized microprocessor chip optimized for digital signal processing—such as in audio signal processing, telecommunications, digital image processing, radar, sonar, LiDAR, and / or other sensor processing, speech recognition, and / or other applications.

[0207] The one or more SoCs 604 can contain one or more video encoders 619 and / or one or more video decoders 621. For example, the one or more video encoders 619 can include a hardware-based video encoder (e.g., as part of the one or more GPUs 608) that supports BH264, H.265, etc., and is HEVC-compliant, such as NVIDIA's NVENC, which can process image inputs (e.g., as YUV, RGB, etc.) to generate a video bitstream. The single or multiple Video Decoder 621 units can contain a video decoder unit capable of providing fully accelerated hardware video decoding capabilities (e.g., support for decoding bitstreams in various formats such as AV1, H.264, H.265, VP8, VP9, ​​MPEG-1, MPEG-2, MPEG-4, VC-1, etc., and HEVC compatibility, like NVIDIA's NVDEC). In some examples, the single or multiple Video Decoder 621 units can be hardware-based (e.g.,as part of one or more GPUs (608).

[0208] The one or more SoCs 604 can contain one or more General Compute Acceleration Clusters (GCACs) 629. For example, the one or more GCACs 629 can include various processor types that can be used to accelerate computation, such as one or more Vector Microcode Processors (VMPs) 633, one or more Multithreaded Processing Clusters (MPCs) 631, one or more Programmable Macro Arrays (PMAs) 635, and / or one or more other processor types. For example, the one or more GCACs 629 can include one PMA 635, two VMPs 633, and two MPCs 631.

[0209] The one or more SoCs 604 can contain one or more vector microcode processors (VMPs) 633. The one or more VMPs 633 can, in embodiments, comprise a wide-vector machine (very long instruction word, VLIW) and a single instruction multiple data machine (SIMD) that performs various operations, such as short integral-like operations commonly used in computer vision and deep learning algorithms.

[0210] The one or more SoCs 604 can contain one or more general-purpose multithreaded processing clusters (MPCs) 631. The one or more MPCs 631 can comprise a processing cluster that, in some implementations, is more versatile than a GPU and more efficient than a CPU. For example, the one or more MPCs 631 can include a multithreaded processor that allows multiple threads to share resources and execute instructions concurrently.

[0211] The one or more SoCs 604 can contain one or more Programmable Macro Arrays (PMAs) 635. The one or more PMAs 635 can comprise a Coarse-Grained Reconfigurable Architecture (CGRA) data flow machine, which has a unique architecture that delivers strong performance in dense computer vision and deep learning algorithms that may not be achievable in classic digital signal processing (DSP) architectures.

[0212] The one or more SoCs 604 can contain one or more Display Processing Units (DPUs) 645 for performing hardware-accelerated image processing. For example, the one or more DPUs 645 can retrieve pixel data from memory 615 and send it to a display peripheral device via standard interfaces. Thus, the one or more DPUs 645 can handle display processing and playback for displays in and / or on the machine.

[0213] The one or more 604 SoCs can contain one or more 639 Application Processing Units (APUs). For example, the one or more 639 APUs can contain a quad-core or dual-core processor with 48 KB / 32 KB L1 cache with parity and ECC, and a 1 MB L2 cache with ECC. The one or more 639 APUs can support NEON instructions as well as single- and double-precision floating-point operations.

[0214] The one or more 604 SoCs can contain one or more 669 Real-Time Processing Units (RTPUs). The one or more 669 RTPUs can contain a dual-core processor with 32 KB / 32 KB L1 cache and 256 KB TCM with ECC. The one or more 669 RTPUs can support single- and dual-precision floating-point operations.

[0215] The one or more SoCs 604 can contain one or more built-in self-test (BIST) components 637. For example, the one or more BIST components 637 can include a memory BIST (MBIST) for testing the system's memory and / or a logic BIST (LBIST) for testing the system's logic. The BIST components 637 can include embedded logic for directly testing the system's logic and / or memory.

[0216] The one or more SoCs 604 can contain one or more dynamically reconfigurable processors (DRPs) 671. The one or more DRPs 671 can be used, for example, to accelerate various computational operations. For instance, in embodiments, the one or more DRPs 671 can be combined with a MAC unit for use as an AI accelerator. In embodiments, the one or more DRPs 671 can execute applications while dynamically changing the configuration of the circuit interconnects of the on-chip arithmetic units (e.g., ALUs) at each operating clock cycle, according to the content being processed. Because only the necessary arithmetic circuits are used, the one or more DRPs 671 may consume less power than CPU processing and can achieve higher speeds.Furthermore, compared to CPUs where frequent external memory accesses can degrade performance due to cache errors and other causes, one or more DRPs 671 processors can pre-build the necessary data paths in hardware, resulting in less performance degradation and reduced jitter caused by memory accesses. One or more DRPs 671 processors can also incorporate a dynamic loading function that changes the circuit connection information with each algorithm change, enabling processing with limited hardware resources, even in robotics / automotive applications that require the processing of multiple algorithms.

[0217] In some embodiments, the one or more accelerators 614 may include an OpenCV accelerator to accelerate the processing of OpenCV, an industry-standard open-source library for image processing. In some embodiments, the combination of one or more DRPs 671, used as AI accelerators, together with one or more OpenCV accelerators, may enhance AI computational and image processing algorithms and enable complex and computationally intensive operations such as simultaneous visual localization and mapping (SLAM).

[0218] In contrast to conventional systems, the technology described herein, by providing a CPU complex, a GPU complex, and a hardware acceleration cluster, enables the simultaneous (e.g., at least partially parallel) and / or sequential execution of multiple neural networks and the combination of their results to enable Level 2-5 autonomous driving functionality and / or autonomous robot movements, control, planning, and / or navigation operations. Furthermore, since the one or more SoCs 604 can contain various computing units (e.g., Processors 610, CPUs 606, GPU(s) 608, Accelerators 614, etc.), tasks can be distributed among the computing units, in some cases without common-cause failures due to the discrete footprint of the computing units.Since the one or more SoCs 604 can also contain one or more dedicated safety processors 613 (or a safety island 613), critical safety or redundancy operations can be performed by the main processing components or computing units of the one or more SoCs 604 without common cause failures. Because of these features, the one or more SoCs 604 and / or the underlying systems of the machine 600 can be capable of meeting higher safety levels—for example, automotive safety integrity level (ASIL) D from the ISO 26262 standard.

[0219] Fig. 6E is a system diagram for communication between one or more cloud-based servers (e.g., a data center such as those described herein) and the exemplary autonomous or semi-autonomous vehicle or machine 600 from Fig. 6A according to some embodiments of the present disclosure. The system 676 may include one or more servers 678, one or more networks 690, and one or more machines 600. The one or more servers 678 may include a variety of GPUs 684(A)-684(H) (here collectively referred to as GPUs 684), switches 682(A)-682(D) (e.g., PCIe 4.0 / 5.0 switches, etc., M.2 slots, Thunderbolt, USB4, NVIDIA NVLink, NVIDIA NVSwitch, GPUDirect RDMA, GPUDirect Storage, etc.), CPUs 680(A)-680(B) (here collectively referred to as CPUs 680), accelerators, and / or other processor types. The GPUs 684, the CPUs 680 and the PCIe switches 682 can be interconnected using high-speed connections, such as, without limitation, the NVIDIA-developed NVLink interfaces 688 and / or PCIe connections 686.In some examples, the GPUs 684 are connected via NVLink and / or NVSwitch SoCs, and the GPUs 684 and PCIe switches 682 are connected via PCIe links. Although eight GPUs 684, two CPUs 680, and four PCIe switches are illustrated, this should not be considered a limitation. Depending on the configuration, each Server 678 can contain any number of GPUs 684, CPUs 680, and / or PCIe switches. For example, one or more Server 678s can each contain eight, sixteen, thirty-two, and / or more GPUs 684.

[0220] The one or more servers 678 can receive sensor data from the one or more machines 600 via the one or more networks 690, indicating information about new or previously unexplored locations, and / or sensor data indicating changes to previously seen / stored locations (e.g., unexpected or altered road conditions, such as recently started roadworks). The one or more servers 678 can transmit neural networks 692, updated neural networks 692, map information 694, etc., containing information about traffic and road conditions, to the one or more machines 600 via the one or more networks 690. The map information updates 694 can include updates for the HD map 622, the SD card, the navigation map, etc., such as information about construction sites, potholes, detours, floods, and / or other obstacles.In some examples, the neural networks 692, the updated neural networks 692, the map information 694 and / or the other information may result from new training and / or experience represented in the data received by any number of one or more machines 600 in the environment, and / or may be based on training performed in a data center (e.g. using one or more servers 678 and / or other servers).

[0221] One or more Server 678 machines can be used to train machine learning models (e.g., neural networks) based on training data. The training data can be generated by one or more Machine 600 machines and / or in a simulation (e.g., using a game engine). In some examples, the training data is tagged (e.g., if the neural network benefits from supervised learning) and / or subjected to other preprocessing, while in other examples, the training data is not tagged and / or preprocessed (e.g., if the neural network does not require supervised learning).Training can be performed using one or more classes of machine learning techniques, including, but not limited to, classes such as: supervised training, semi-supervised training, unsupervised training, self-learning, reinforcement learning, federated learning, transfer learning, feature learning (including principal component and cluster analysis), multilinear subspace learning, diverse learning, representational learning (including substitute dictionary learning), rule-based machine learning, anomaly detection, and all variants or combinations thereof. Once the machine learning models are trained, they can be used by one or more machines (e.g.,to the one or more machines 600 via the one or more networks 690 and / or the machine learning models can be used by the one or more servers 678 for remote monitoring and / or control of the one or more machines 600.

[0222] In some examples, one or more Server 678s can receive data from one or more Machines 600s and apply the data to current neural networks in real time for real-time intelligent inference. The one or more Server 678s can include deep learning supercomputers and / or dedicated AI computers powered by GPUs 684s, such as NVIDIA's DGX and DGX Station machines. However, in some examples, the one or more Server 678s can include a deep learning infrastructure that uses only CPU-powered data centers.

[0223] The deep learning infrastructure of one or more Server 678s can perform fast, real-time inference and can use this capability to assess and verify the state of the processors, software, and / or associated hardware in the Machine 600. For example, the deep learning infrastructure can receive periodic updates from the Machine 600, such as a sequence of images and / or objects that the Machine 600 has located within that sequence (e.g., via computer vision and / or other machine learning object classification techniques).The deep learning infrastructure can run its own neural network to identify the objects and compare them with the objects identified by the machine 600. If the results do not match and the infrastructure concludes that the AI ​​in the machine 600 is not functioning correctly, one or more servers 678 can send a signal to the machine 600, instructing a failsafe computer of the machine 600 to take control, notify the passengers, and perform a safety maneuver or operation—such as slowing down, handing control back to a driver, stopping, and / or pulling over to the side of the road / turning off the engine.

[0224] For inference, one or more Server 678s can include GPUs 684s and one or more programmable inference accelerators (e.g., NVIDIA's TensorRT). The combination of GPU-driven servers and inference accelerators can enable real-time responsiveness. In other scenarios, such as when performance is less critical, servers powered by CPUs, FPGAs, and other processors can be used for inference. COMPUTER ECOSYSTEM FOR THE GENERATION, TRAINING AND DEPLOYMENT OF AI

[0225] Fig. Figure 7 is a system diagram illustrating an ecosystem with three computers 700, including a first computer system 702 for generating or creating artificial intelligence (AI) – such as AI training and validation data – a second computer system 704 for training artificial intelligence, and a third computer system 706 (which contains one or more SoCs 604 of Fig. 6A-6E, which may contain or correspond to them), which deploys the AI ​​at the system edge, according to at least some embodiments of the present disclosure. For example, to develop and deploy embodied or physical AI, the ecosystem with three computers 700 can be used, including three accelerated computing systems for processing physical AI training, simulation, and runtime (e.g., deployment at the system edge). These systems can generate training data for multimodal base models (and / or other model types) and train them using scalable, physically based simulations of the one or more machines 600 and their worlds. In this way, the simulation of the one or more machines 600 can be performed at scale, enabling the refinement, testing, and optimization of capabilities (e.g., robotic capabilities) in a virtual world (e.g.,using NVIDIA's OMNIVERSE) which mimics the laws of physics - this helps to reduce the cost of real-world data acquisition and ensures that one or more 600 machines can operate safely in controlled environments.

[0226] The 704 computer system (e.g., NVIDIA's DGX platform) can be used to train and fine-tune high-performance basic and generative AI models. Models such as general-purpose basic models (e.g., NVIDIA's Project GROOT) can be used to enable robots and one or more other Machine 600s to understand natural language and mimic movements by observing human actions. The 704 computer system can include a platform that integrates software, infrastructure, and expertise into a modern, unified AI development and training solution. The 704 computer system can include individual 710 computer devices (e.g., NVIDIA's DGX B200, H200, etc.) and / or any number of 710 computer devices in a 712 data center infrastructure (e.g., NVIDIA's DGX SuperPOD).

[0227] For example, individual 710 compute units can include GPUs (e.g., 8 GPUs with a total of 1,440 GB of GPU memory) and CPUs (e.g., 2 CPUs with a total of 112 cores, 2.1 GHz or 4 GHz (with boost)), providing more than 72 petaFLOPS for training and 144 petaFLOPS for inference. The 710 compute units can include RAM (e.g., 4 TB of RAM) and storage (e.g., 2 x 1.9 TB NVMe M.2 OS storage and 8 x 3.84 TB NVMe U.2 internal storage). The Computing Devices 710 can include various networking and network management components, such as OSFP ports (e.g., 4 OSFP ports), which support intelligent single-port host channel adapters (e.g., 8 single-port ConnextX-7 virtual protocol interconnects, VPIs) and provide up to 400 GB / s of InfiniBand / Ethernet. The Computing Devices 710 can also include, for example, dual-port quad-small form-factor pluggable (QSFFP) data processing units (DPUs) (e.g.,The 710 compute devices contain two dual-port QSFP112 DPUs (such as NVIDIA's BlueField 3 DPUs) that provide up to 400 Gb / s InfiniBand / Ethernet. The one or more 710 compute devices can include an integrated network interface card (NIC) (e.g., an integrated 10 Gb / s NIC with RJ45), a dual-port Ethernet NIC (e.g., a 100 Gb / s dual-port Ethernet NIC), and / or a host baseboard management controller (MBC) (e.g., with RJ45). In some embodiments, the NICs used for the one or more 710 compute devices may include SuperNICs (e.g., NVIDIA's ConnectX-8 SuperNIC) to provide up to 800 Gb / s of data throughput for network-internal acceleration compute units and deliver the performance and robust feature set required to run AI models with trillions of parameters and scientific computational loads.In other embodiments, the 710 compute devices can include an intelligent host channel adapter (HCA) (e.g., NVIDIA's ConnectX-7) to provide extremely low latency and a throughput of 400 Gb / s for network-internal acceleration compute units.

[0228] The 712 data center infrastructure can include any number of 710 compute devices, along with an operating system (OS) (e.g., DGX OS extensions for Linux distributions) to maximize system availability, security, and reliability; network / storage acceleration libraries and management to accelerate end-to-end infrastructure performance; cluster management to scale and manage from one node (e.g., one 710 compute device) to thousands; job scheduling and orchestration to ensure the smooth execution of every developer task; AI workflow management and machine learning operations (MLOps) to move more models from prototype to production; and enterprise software to accelerate developer success.

[0229] The Computer System 702 (e.g., NVIDIA OVX Server) can provide a development and simulation platform for testing and optimizing physical AI with APIs and simulation frameworks (e.g., NVIDIA DriveSIM, ISAAC Sim, ISAAC Gym, ISAAC Lab, etc.). The Computer System 702 enables developers to use simulation frameworks to simulate and validate robot models and / or generate large amounts of physics-based synthetic data to initiate model training. The Computer System 702 can support learning frameworks that enable reinforcement learning and imitation learning for robots to accelerate the training and refinement of robot guidelines. For example, the Computer System 702 can be used to generate any number of simulations—for instance, within NVIDIA Omniverse.The Computer System 702 can be optimized to accelerate an entire software stack, from training, fine-tuning, and deploying generative AI to advancing industrial digitalization within a content collaboration platform. This platform offers APIs, Software Development Kits (SDKs), and services that enable the integration of OpenUSD, ray tracing rendering technologies (e.g., NVIDIA RTX), and generative physical AI into existing software tools and simulation workflows for industrial and robotic use cases (e.g., NVIDIA OMNIVERSE). For example, the Computer System 702 can host or support a native OpenUSD software platform, allowing companies to connect 3D pipelines and develop advanced, real-time 3D applications for industrial digitalization.With powerful, ray-tracing-accelerated AI and graphics capabilities, the Computing System 702 delivers high performance for workloads such as extended reality (XR), multi-user design collaboration, and digital twins. This enables the creation of physically accurate models with high-precision ray-tracing and path-tracing rendering of materials, the operation of large-scale AI-powered simulations, and the generation of photorealistic synthetic 3D data for training purposes. The Computing System 702 can include individual Computing Devices 714 (e.g., NVIDIA OVX L40S servers) and / or any number of Computing Devices 714 in a Data Center Infrastructure 716 (e.g., NVIDIA OVX systems).

[0230] The one or more computing devices 714 (which may comprise a server) may contain CPUs (e.g., 2 CPUs with 32 cores each) and GPUs (e.g., 4 or 8 GPUs, each with 48 GB of GDDR6 ECC memory, 864 GB / s memory bandwidth, PCIe Gen4 x 16: 64 GB / s bidirectional interconnect interface, 18,176 CUDA cores, 142 Raytracing (RT) cores, and 568 Tensor cores). The Computing Devices 714 can include various networking and network management components, such as intelligent host channel adapters (HCAs) (e.g., two or four single-port ConnextX-7s, each with 200 Gb / s, providing up to 800 Gb / s InfiniBand / Ethernet), one or more DPUs (e.g., dual-port QSFP112 DPUs—such as an NVIDIA BlueField 3 DPU) providing up to 400 Gb / s InfiniBand / Ethernet. In some embodiments, the NICs used for the one or more Computing Devices 714 can be SuperNICs (e.g.,The 714 compute devices can include an NVIDIA ConnectX-8 SuperNIC to provide up to 800 Gb / s of throughput for network-integrated acceleration compute units, delivering the performance and robust feature set needed to run trillions of parameters AI models and scientific workloads. In other embodiments, the one or more 714 compute devices can include an intelligent host channel adapter (HCA) (such as NVIDIA ConnectX-7) to provide ultra-low latency and 400 Gb / s throughput for network-integrated acceleration compute units. The one or more computing devices 714 can include host memory (e.g., 384 Gb DDR5 ECC for 4 GPUs or 768 Gb DDR5 ECC for 8 GPUs) and one or more dual in-line memory module (DIMM) slots, a host boot drive (e.g., 1 TB NVMe), and / or host storage (e.g., 2 4 TB NVMe drives).

[0231] Similar to the data center infrastructure 712, the data center infrastructure 716 can enable the combination of any number of computing devices 714 in a cluster configuration according to a reference architecture.

[0232] The 706 computer system can be used to deploy trained AI models on a runtime computer—such as one or more of the 604 SoCs described here. For example, these 706 computer systems can be designed for compact onboard computing needs, including a suite of control policy, vision, and speech models, etc., deployed on a power-efficient, onboard edge 706 computer system. Details regarding the components, features, and capabilities of the 706 computer system can be found herein with respect to Fig. 6A-6E will be described in more detail. EXEMPLARY GENERATIVE MODELS

[0233] In at least some embodiments, language models such as large language models (LLMs), vision language models (VLMs), multimodal language models (MMLMs), vision language action (VLA) models, and / or other types of generative artificial intelligence (AI) can be implemented. These models may be able to understand, summarize, translate, and / or otherwise generate text (e.g., natural language text, code, etc.), images, videos, computer-aided design (CAD) assets, OMNIVERSE and / or METAVERSE file information (e.g., in USD format, such as OpenUSD), and / or the like, based on the context provided in prompts or queries.These language models can be considered "large" in embodiments because they are trained on extensive datasets and feature architectures with a large number of learnable network parameters (weights and distortions)—approximately millions or billions of parameters. The LLMs / VLMs / MMLMs, etc., can be implemented to summarize text data, analyze data (e.g., text, images, videos, etc.) and derive insights from it, and generate new text / images / videos, etc., in user-defined styles, tones, and / or formats. The LLMs / VLMs / MMLMs, etc., of the present disclosure can, in embodiments, be used exclusively for text processing, while in other embodiments, multimodal LLMs can be implemented to process text and / or other types of content such as images, audio (tones, synthetic speech, etc.), 2D and / or 3D data (e.g.,in USD formats) and / or videos. For example, Vision Language Models (VLMs) or more generally Multi-Modal Language Models (MMLMs) can be implemented to accept image, video, sensor, audio, text, 3D design (e.g., CAD) and / or other input data types and / or to generate or output image, video, audio, text, 3D design and / or other output data types.

[0234] Different types of LLM / VLM / MMLM architectures, etc., can be implemented in various embodiments. For example, different architectures can be implemented that use different techniques for understanding and generating outputs—such as text, audio, video, images, 2D and / or 3D design, or asset data, etc. In some embodiments, LLM / VLM / MMLM architectures, etc., can use recurrent neural networks (RNNs) or long short-term memory networks (LSTMs), while in other embodiments, transformer architectures—such as those based on self-attention and / or cross-attention mechanisms (e.g., between contextual and textual data)—can be used to understand and recognize relationships between words or tokens and / or contextual data (e.g., other text, video, images, design data, USD, etc.).One or more generative processing pipelines, comprising LLMs / VLMs / MMLMs, etc., may also include one or more diffusion blocks (e.g., denoisers). The LLMs / VLMs / MMLMs, etc., of this disclosure may contain encoder and / or decoder blocks. For example, discriminative or pure encoder models such as BERT (Bidirectional Encoder Representations from Transformers) may be implemented for tasks involving language understanding, such as classification, sentiment analysis, question answering, and named entity recognition. As another example, generative or pure decoder models such as GPT (Generative Pretrained Transformer) may be implemented for tasks involving language and content generation, such as text completion, story creation, and dialogue generation.Architectures that contain both encoder and decoder components, such as T5 (Text-to-Text Transformer), can be implemented to understand and generate content, for example, for translations and summaries. These examples are not intended to be restrictive, and any architecture type—including, but not limited to, those described here—can be implemented depending on the specific implementation and the one or more tasks performed by the LLMs / VLMs / MMLMs, etc.

[0235] In various embodiments, LLMs / VLMs / MMLMs, etc., can be trained using unsupervised learning, where an LLM / VLM / MMLM, etc., learns patterns from large amounts of unlabeled text / audio / video / image / design / USD data, etc. Due to the extensive training, the models in some embodiments may not require task-specific or domain-specific training. LLMs / VLMs / MMLMs, etc., that have undergone extensive pre-training with large amounts of unlabeled data can be referred to as base models and may be proficient in a variety of tasks such as answering questions, summarizing, filling in missing information, translating, and generating images / videos / designs / USD / data. Some LLMs / VLMs / MMLMs, etc., can be further enhanced using techniques such as prompt tuning, fine-tuning, retrieval augmented generation (RAG), and the addition of adapters (e.g., a static, static, or static).B. adapted neural networks and / or neural network layers that adjust or adapt prompts or tokens to align the language model with a specific task or domain) and / or using other fine-tuning or adaptation techniques that optimize the models for use in specific tasks and / or within specific domains.

[0236] In some embodiments, the LLMs / VLMs / MMLMs, etc., of this disclosure can be implemented using various model alignment techniques. For example, in some embodiments, guardrails can be implemented to identify impermissible or unwanted inputs (e.g., prompts) and / or outputs of the models. The system can then use the guardrails and / or other model alignment techniques to either prevent a specific unwanted input from being processed by the LLMs / VLMs / MMLMs, etc., and / or to prevent the output or representation (e.g., display, audio output, etc.) of information generated by the LLMs / VLMs / MMLMs, etc. In some embodiments, one or more additional models—or layers thereof—can be implemented to identify problems with the inputs and / or outputs of the models.For example, these “security models” can be trained to identify inputs and / or outputs that are “safe” or otherwise acceptable or desirable, and / or that are “unsafe” or otherwise undesirable for the specific application / implementation. As a result, the LLMs / VLMs / MMLMs, etc., of this disclosure are less likely to output speech / text / audio / video / design data / USD data, etc., that is offensive, vulgar, inappropriate, unsafe, out-of-domain, and / or otherwise undesirable for the specific application / implementation.

[0237] In some implementations, the LLMs / VLMs, etc., can be configured or enabled to access or use one or more plug-ins, application programming interfaces (APIs), databases, data stores, directories, etc. For example, for certain tasks or operations for which it is not ideally suited, the model may include instructions (e.g., as a result of training and / or based on instructions in a specific prompt) to access one or more plug-ins (e.g., third-party plug-ins) that assist in processing the current input. In such an example, where at least part of a prompt is related to restaurants or the weather, the model may access one or more restaurant or weather plug-ins (e.g., via one or more APIs) to retrieve the relevant information.In another example, where at least part of an answer requires a mathematical calculation, the model can access one or more math plugins or APIs to help solve the one or more problems and then use the answer from the plugin and / or API in the model's output. This process can be repeated—for example, recursively—for any number of iterations and using any number of plugins and / or APIs until an answer to the input prompt can be generated that addresses each requirement / question / request / process / operation, etc. Thus, the model relies not only on its own knowledge gained from training with large datasets but also on the expertise or optimization of one or more external resources—such as APIs, plugins, and / or the like.

[0238] In some embodiments, multiple language models (e.g., LLMs / VLMs / MMLMs, etc., multiple instances of the same language model, and / or multiple prompts served to the same language model or instance of the same language model) can be implemented, executed, or accessed (e.g., using one or more plug-ins, user interfaces, APIs, databases, data stores, directories, etc.) to provide output that responds to the same query or to separate parts of a query. In at least one embodiment, multiple language models, e.g., language models with different architectures, language models trained on different (e.g., updated) data corpora, can be served with the same input query and the same prompt (e.g., a set of constraints, control variables, etc.).In one or more embodiments, the language models can be different versions of the same base model. In one or more embodiments, at least one language model can be instantiated as multiple agents—for example, more than one prompt can be provided to constrain, control, or otherwise influence a style, content, character, etc., of the output provided. In one or more non-constraint embodiments, the same language model can be instructed to provide output corresponding to a different role, perspective, character, knowledge base, etc., as defined by a provided prompt.

[0239] In each of these embodiments, the output of two or more (e.g., all) language models, two or more versions of at least one language model, two or more instantiated agents of at least one language model, and / or two or more prompts provided to at least one language model can be further processed, e.g., aggregated, compared, or filtered, or used to determine (and provide) a consensus response. In one or more embodiments, the output of a language model—or a version, instance, or agent—can be provided as input to another language model for further processing and / or validation. In one or more embodiments, a language model can be prompted to generate or otherwise obtain output with respect to input source material, the output being associated with the input source material.Such mapping can, for example, involve generating a caption or text snippet that is embedded (e.g., as metadata) in input source text or an input source image. In one or more embodiments, an output from a language model can be used to determine the validity of input source material for further processing or inclusion in a dataset. For example, a language model can be used to evaluate the presence (or absence) of a target word in a text snippet or an object in an image, annotating the text or image to indicate this presence (or absence). Alternatively, the determination from the language model can be used to decide whether the source material should be included in a compiled dataset, for example, without restriction.

[0240] Fig. Figure 8 is a block diagram of one or more exemplary generative language model systems 800 suitable for use in the implementation of at least some embodiments of the present disclosure. In the Fig. In the 8 illustrated example, the generative language model system 800 contains a retrieval augmented generation (RAG) component 892, an input processor 805, a tokenizer 810, an embedding component 820, the plug-ins / APIs 895, and a generative language model (LM) 830 (which may include an LLM, a VLM, an MMLM, a VLA model, etc.).

[0241] At a higher level, the input processor 805 can receive an input 801 that includes text and / or other types of input data (e.g., audio data, video data, image data, sensor data (e.g., LiDAR, RADAR, ultrasound, etc.), 3D design data, CAD data, Universal Scene Descriptor (USD) data such as OpenUSD, etc.), depending on the architecture of the generative LM 830 (e.g., LLM / VLM / MMLM, etc.). In some embodiments, the input 801 contains plain text in the form of one or more sentences, paragraphs, and / or documents. Additionally or alternatively, the input 801 can include numeric sequences, pre-computed embeddings (e.g., word or sentence embeddings), and / or structured data (e.g., in tabular formats, JSON, or XML).In some implementations where the generative LM 830 is capable of processing multimodal input, the input 801 can combine text with image data, audio data, video data, design data, USD data, and / or other types of input data, such as, but not limited to, those described herein (or it can omit text). Using raw input text as an example, the input 805 can prepare raw input text in various ways. For instance, the input 805 can perform various types of text filtering to remove noise (such as special characters, punctuation, HTML tags, stop words, parts of one or more images, parts of audio, etc.) from relevant text content. In an example involving stop words (frequent words that typically have little semantic meaning), the input 805 can remove stop words to reduce noise and allow the generative LM 830 to focus on more meaningful content.The 805 input processor can apply text normalization (TN), for example, by converting all characters to lowercase, removing accents, and / or handling special cases such as contractions or abbreviations to ensure consistency (e.g., converting ¼ to a quarter). Likewise, the 805 input processor and / or a post-processor can perform inverse text normalization (ITN) to convert plain language back into canonical or other forms (e.g., converting a quarter to ¼). These are just a few examples, and other types of input and / or output processing can also be applied.

[0242] In some embodiments, a RAG component 892 (which can contain one or more RAG models and / or be executed using the generative LM 830 itself) can be used to retrieve additional information to be used as part of the input 801 or prompt. RAG can be used to enhance the input to the LLM / VLM / MMLM, etc., with external knowledge, making answers to specific questions or queries more relevant, such as in cases where specialized knowledge is required. The RAG component 892 can retrieve this additional information (e.g., background information such as background text / image / video / audio / USD / CAD, etc.) from one or more external sources, which can then be passed along with the prompt to the LLM / VLM / MMLM, etc., to improve the accuracy of the model's answers or outputs.

[0243] In some embodiments, the input 801 can be generated, for example, using the query or input into the model (e.g., a question, a query, etc.) in addition to the data retrieved by the RAG component 892. In some embodiments, the input processor 805 can analyze the input 801 and communicate with the RAG component 892 (or the RAG component 892 can be part of the input processor 805 in some embodiments) to identify relevant text and / or other data to be provided to the generative LM 830 as additional context or as information sources from which the answer, solution, or output 890 can be identified in general.For example, if the input indicates that the user is interested in a desired tire pressure for a specific vehicle make and model, the RAG component 892—for example, using a RAG model that performs a vector search in an embedding space—can retrieve the tire pressure information or the relevant text from a digital (embedded) version of the owner's manual for that specific vehicle make and model. Similarly, if a user revisits a chatbot in connection with a specific product offering or service, the RAG component 892 can retrieve a previously stored conversation history—or at least a summary of it—and include the previous conversation history along with the current query / request as part of the 801 input to the generative LM 830.

[0244] The RAG component 892 can employ various RAG techniques. For example, naive RAG can be used when documents are indexed, divided into sections, and applied to an embedding model to generate embeddings that correspond to the sections. A user query can also be applied to the embedding model and / or another embedding model of the RAG component 892, and the section embeddings can be compared with the query embeddings to identify the embeddings most similar to the query that can be provided to the generative LM 830 to generate output.

[0245] In some embodiments, more advanced RAG techniques can be used. For example, sections can undergo retrieval preparation processes (e.g., forwarding, rewriting, metadata analysis, extension, etc.) before being passed to the embedding model. Furthermore, retrieval post-processing processes (e.g., re-evaluation, prompt compression, etc.) can be performed on the outputs of the embedding model before the final embeddings are generated and used for comparison with an input query.

[0246] As another example, modular RAG techniques can be used, which are similar to those of naive and / or advanced RAG, but also include features such as hybrid search, recursive retrieval and query units, step-back approaches, subqueries and hypothetical document embedding.

[0247] As another example, GraphRAG can use knowledge graphs as a source of contextual or factual information. GraphRAG can be implemented using a graph database as a source of contextual information, which is then sent to the LLM / VLM / MMLM, etc. Instead of (or in addition to) providing the model with blocks of data from larger documents—which can lead to a lack of context, factual accuracy, linguistic precision, etc.—GraphRAG can also provide the LLM / VLM / MMLM, etc., with structured entity information by combining the structured text description of the entity with its many properties and relationships, thus enabling deeper insights for the model. In implementing GraphRAG, the systems and procedures described here use a graph as a content store, extract relevant document sections, and request the LLM / VLM / MMLM, etc., to use them to answer questions.In such embodiments, the knowledge graph can contain relevant text content and metadata about the knowledge graph and be integrated into a vector database. In some embodiments, GraphRAG can use a graph as a domain expert, extracting descriptions of concepts and entities relevant to a query / prompt and passing them to the model as semantic context. These descriptions can include relationships between the concepts. In other examples, the graph can be used as a database, where part of a query / prompt can be mapped to a graph query, the graph query can be executed, and the LLM / VLM / MMLM, etc., can summarize the results.In such an example, the graph can store relevant factual information, and a query (natural language query) to a graph query tool (NL-to-Graph query tool) as well as entity joining can be used. In some implementations, GraphRAG (e.g., using a graph database) can be combined with standard RAG (e.g., vector database) and / or other RAG types to benefit from multiple approaches.

[0248] In all embodiments, the RAG component 892 can implement a plug-in, an API, a user interface, and / or other functions to execute RAG. For example, a GraphRAG plug-in can be used by the LLM / VLM / MMLM, etc., to query the knowledge graph and extract relevant information for feeding into the model, and a standard or vector RAG plug-in can be used to query a vector database. For example, the graph database can interact with a plug-in's REST interface, thus decoupling the graph database from the vector database and / or the embedding models.

[0249] The Tokenizer 810 can segment (e.g., processed) text data into smaller units (tokens) for subsequent analysis and processing. Depending on the implementation, the tokens can represent individual words, partial words, characters, portions of audio / video image data, etc. Word-based tokenization divides the text into individual words, with each word treated as a separate token. Part-word tokenization breaks words down into smaller meaningful units (e.g., prefixes, suffixes, word stems), enabling the generative LM 830 to understand morphological variations and process words outside the vocabulary more effectively. Character-based tokenization represents each character as a separate token, allowing the generative LM 830 to process text at a fine-grained level.The choice of tokenization strategy can depend on factors such as the language to be processed, the specific task, and / or the characteristics of the training dataset. Therefore, the Tokenizer 810 can convert the (e.g., processed) text into a structured format according to the tokenization scheme implemented in the respective implementation.

[0250] The Embedding Component 820 can use any known embedding technique to convert discrete tokens into (e.g., dense, continuous vector) representations of semantic meaning. For example, the Embedding Component 820 can use pre-trained word embeddings (e.g., Word2Vec, GloVe, or FastText), one-hot coding, term frequency-inverse document frequency (TF-IDF) coding, one or more neural network embedding layers, and / or other techniques.

[0251] In some implementations where the input 801 contains image / video data, etc., the input processor 805 can scale the data to a standard size compatible with the format of a corresponding input channel and / or scale the pixel values ​​to a common range (e.g., 0 to 1) to ensure consistent display, and the embedding component 820 can encode the image data using any known technique (e.g., using one or more convolutional neural networks, CNNs, to extract visual features). In some implementations where the input 801 contains audio data, the input processor 805 can resample an audio file to a consistent sampling rate for uniform processing, and the embedding component 820 can use any known technique to extract and encode audio features—for example, in the form of a spectrogram (e.g.,(of a Mel spectrogram). In some implementations where the input 801 contains video data, the input processor 805 can extract individual frames or apply resizing to extracted frames, and the embedding component 820 can extract features such as optical flow embeddings or video embeddings and / or encode temporal information or sequences of frames. In some implementations where the input 801 contains multimodal data, the embedding component 820 can merge representations of the different data types (e.g., text, image, audio, USD, video, design, etc.) using techniques such as early fusion (concatenation), late fusion (sequential processing), attention-based fusion (e.g., self-attention, cross-attention), etc.

[0252] The generative LM 830 and / or other components of the generative LM System 800 can employ various types of neural network architectures, depending on the implementation. For example, transformer-based architectures, as used in models such as GPT, can be implemented and may include self-attentional mechanisms that weight the meaning of different words or tokens in the input sequence, and / or feedforward networks that process the output of the self-attentional layers, apply nonlinear transformations to the input representations, and extract higher-level features. Some non-restrictive example architectures include transformers (e.g.,Encoder-decoder, decoder-only, multimodal), RNNs, LSTMs, fusion models, diffusion models, crossmodal embedding models that learn shared embedding spaces, graph neural networks (GNNs), hybrid architectures that combine different types of architectures, adversarial networks such as generative adversarial networks (GANs) or adversarial autoencoders (AAEs) for shared distributional learning, linear time sequence modeling with selective state space modeling (SSM) architectures (e.g., Mamba-LLM architectures), and / or others. Depending on the implementation and architecture, the embedding component 820 can apply a coded representation of the input 801 to the generative LM 830, and the generative LM 830 can process the coded representation of the input 801 to generate an output 890 that may contain response texts and / or other types of data.

[0253] As described herein, in some embodiments the generative LM 830 can be configured to access, use, or be able to access or use plug-ins / APIs 895 (which may include one or more plug-ins, application programming interfaces (APIs), databases, data stores, directories, etc.). For example, for certain tasks or operations for which the generative LM 830 is not ideally suited, the model may access one or more plug-ins / APIs 895 (e.g., third-party plug-ins) to obtain assistance in processing the current input (e.g., as a result of training and / or based on instructions in a particular prompt, such as those retrieved using the RAG component 892).In such an example, where at least part of a prompt relates to restaurants or the weather, the model can access one or more restaurant or weather plugins (e.g., via one or more APIs), send at least part of the prompt relating to the respective plugin / API 895 to the plugin / API 895, the plugin / API 895 can process the information and return a response to the generative LM 830, and the generative LM 830 can use the response to generate the output 890. This process can be repeated—e.g., recursively—for any number of iterations and using any number of plugins / APIs 895 until an output 890 can be generated that covers every request / question / request / process / operation, etc., of the input 801.Thus, one or more models may rely not only on their own knowledge from training with one or more large datasets and / or from data retrieved using the RAG component 892, but also on the expertise or optimization of one or more external resources such as the plug-ins / APIs 895.

[0254] In some embodiments, one or more transformer engines (TEs) can be implemented. The transformer engine can use microtensor scaling to optimize performance and accuracy—for example, to enable 16-bit floating-point (FP16), 8-bit floating-point (FP8), and / or 4-bit floating-point (FP4) processing for artificial intelligence. For instance, the transformer engine can use 16-bit or 8-bit floating-point precision and an 8-bit or 4-bit floating-point data format in combination with software algorithms to enhance AI performance and capabilities. By reducing mathematical operations to 8 bits or 4 bits, the TE enables faster training of larger networks without sacrificing accuracy.For example, the TEs can include a library for accelerating transformer models on processing devices—such as GPUs—to achieve better performance with reduced memory usage during both training and inference. When combined with other technologies, such as high-speed interconnects between nodes (e.g., using switches like NVLink switches) and Tensor Cores (which enable mixed-precision computing, such as support for microscale precision), server clusters can potentially train enormous networks (e.g., with billions of parameters) at high speed. This allows for support of Tensor Core precisions of FP64, TF32, BF16, FP16, FP8, INT8, FP6, and FP4, as well as CUDA Core precisions of FP64, FP32, FP16, and BF16.

[0255] These and other architectures for LLMs / VLMs / MMLMSs / VLAs etc. described herein serve only as examples, and other suitable architectures may be implemented within the scope of this disclosure. EXAMPLE CALCULATION DEVICE

[0256] Fig. Figure 9 is a block diagram of an exemplary computing device 900 suitable for use in implementing some embodiments of the present disclosure. The computing device 900 may include a connection system 902 that directly or indirectly couples the following devices: main memory 904, one or more central processing units (CPUs) 906, one or more graphics processing units (GPUs) 908, a communication interface 910, input / output (I / O) ports 912, input / output components 914, a power supply 916, one or more presentation components 918 (e.g., display(s), speakers, etc.), and one or more logic units 920. In at least one embodiment, the one or more computing devices 900 may comprise one or more virtual machines (VMs), and / or each of the components thereof may comprise virtual components (e.g., virtual hardware components).As non-restrictive examples, one or more of the GPUs 908 can comprise one or more vGPUs, one or more of the CPUs 906 can comprise one or more vCPUs, and / or one or more of the logic units 920 can comprise one or more virtual logic units. Thus, a compute device 900 can contain discrete components (e.g., a complete GPU allocated to the compute device 900), virtual components (e.g., a portion of a GPU allocated to the compute device 900), or a combination thereof.

[0257] Although the various blocks of Fig. Where components 9 are shown connected via the connection system 902, this is not to be understood as a limitation and serves only for clarity. In some embodiments, for example, a presentation component 918, such as a display device, can be considered an I / O component 914 (e.g., if the display is a touchscreen). As another example, the CPUs 906 and / or GPUs 908 can contain memory (e.g., the memory 904 can represent a memory device in addition to the memory of the GPUs 908, the CPUs 906, and / or other components). In other words, the computing device of Fig. Figure 9 is for illustrative purposes only. No distinction is made between categories such as "workstation", "server", "laptop", "desktop", "tablet", "client device", "mobile device", "handheld device", "game console", "electronic control unit (ECU)", "virtual reality system" and / or other device or system types, as all fall within the scope of the computing device of Fig. 9 are being considered.

[0258] The 902 interconnection system can represent one or more connections or buses, such as an address bus, a data bus, a control bus, or a combination thereof. The 902 interconnection system can include one or more bus or connection types, such as an Industry Standard Architecture (ISA) bus, an Extended ISA bus, a Video Electronics Standards Association (VESA) bus, a Peripheral Component Interconnect (PCI) bus, a PCI Express (PCIe) bus, and / or another type of bus or connection. In some embodiments, there are direct connections between components. For example, the CPU 906 can be directly connected to the memory 904. Furthermore, the CPU 906 can be directly connected to the GPU 908.In a direct or point-to-point connection between components, the 902 connection system can include a PCIe link to establish the connection. In these examples, a PCI bus does not need to be included in the 900 computing device.

[0259] The Memory 904 can contain a variety of computer-readable media. Computer-readable media can be any available media to which the Computing Device 900 can access. Computer-readable media can include both volatile and non-volatile media, as well as removable and non-removable media. By way of example, and without limitation, computer-readable media can include computer storage media and communication media.

[0260] Computer storage media can include both volatile and non-volatile media, and / or removable and non-removable media, implemented in any method or technology for storing information, such as computer-readable instructions, data structures, program modules, and / or other data types. For example, Memory 904 can store computer-readable instructions (e.g., representing one or more programs and / or one or more program elements, such as an operating system). Computer storage media can include, but are not limited to, RAM, ROM, EEPROM, flash memory, or other storage technologies; CD-ROM, Digital Versatile Discs (DVDs), or other optical disk storage; magnetic cartridges, magnetic tapes, magnetic disk storage, or other magnetic storage devices; or any other medium that can be used to store the desired information and that the Computing Device 900 can access.As used herein, computer storage media do not inherently contain signals.

[0261] Computer storage media can embody computer-readable instructions, data structures, program modules, and / or other data types in a modulated data signal, such as a carrier wave or other transport mechanism, and include any media for transmitting information. The term "modulated data signal" can refer to a signal in which one or more of its properties are set or modified to encode information within the signal. Computer storage media can include, but are not limited to, wired media, such as a wired network or a direct-wired connection, and wireless media, such as acoustic, RF, infrared, and other wireless media. Combinations of the above should also be included in the scope of protection of the computer-readable media.

[0262] The one or more CPUs 906 can be configured to execute at least some of the computer-readable instructions to control one or more components of the Computing Device 900 to perform one or more of the procedures and / or processes described herein. The one or more CPUs 906 can each contain one or more cores (e.g., one, two, four, eight, twenty-eight, seventy-two, etc.) capable of processing a multitude of software threads simultaneously. The CPUs 906 can contain any type of processor and can contain different types of processors depending on the type of Computing Device 900 implemented (e.g., processors with fewer cores for mobile devices and processors with more cores for servers).Depending on the type of computing device 900, the processor can be, for example, an Advanced RISC Machine (ARM) processor implemented with Reduced Instruction Set Computing (RISC), or an x86 processor implemented with Complex Instruction Set Computing (CISC). The computing device 900 can contain one or more CPUs 906 in addition to one or more microprocessors or additional coprocessors, such as mathematical coprocessors.

[0263] In addition to or as an alternative to the CPU(s) 906, the GPU(s) 908 may be configured to execute at least some of the computer-readable instructions to control one or more components of the computing device 900 to perform one or more of the procedures and / or processes described herein. One or more of the GPUs 908 may be an integrated GPU (e.g., with one or more of the CPUs 906) and / or one or more of the GPUs 908 may be a discrete GPU. In embodiments, one or more of the GPUs 908 may be a coprocessor of one or more of the CPUs 906. The one or more GPUs 908 may be used by the computing device 900 to render graphics (e.g., 3D graphics) or to perform general-purpose computing. For example, the one or more GPUs 908 may be used for general-purpose computing on GPUs (GPGPU).The one or more GPUs 908 can contain hundreds or thousands of cores capable of processing hundreds or thousands of software threads simultaneously. The one or more GPUs 908 can generate pixel data for output images in response to rendering commands (e.g., rendering commands from the one or more CPUs 906 received through a host interface). The one or more GPUs 908 can include graphics memory, such as display memory, for storing pixel data or other suitable data, such as GPGPU data. The display memory can be included as part of the main memory 904. The one or more GPUs 908 can contain two or more GPUs operating in parallel (e.g., via a link). The link can connect the GPUs directly (e.g., using NVLINK) or connect them via a switch (e.g., using NVSwitch).When combined, each GPU can generate 908 pixel data or GPGPU data for different sections of an output or for different outputs (e.g., a first GPU for a first image and a second GPU for a second image). Each GPU can have its own dedicated memory or share memory with other GPUs.

[0264] In addition to or as an alternative to the one or more CPUs 906 and / or the one or more GPUs 908, the one or more logic units 920 may be configured to execute at least some of the computer-readable instructions to control one or more components of the computing device 900 to perform one or more of the methods and / or processes described herein. In embodiments, the one or more CPUs 906, the GPUs 908, and / or the one or more logic units 920 may discretely or jointly execute any combination of the methods, processes, and / or sections thereof. One or more of the logic units 920 may be part of and / or integrated into one or more of the CPUs 906 and / or one or more of the GPUs 908, and / or one or more of the logic units 920 may be discrete components or otherwise separate from the CPUs 906 and / or the GPUs 908.In embodiments, one or more of the logic units 920 can be a co-processor of one or more of the CPUs 906 and / or one or more of the GPUs 908.

[0265] Examples of the one or more Logic Units 920 include one or more processing cores and / or components thereof, such as Data Processing Units (DPUs), Tensor Cores (TCs), Tensor Processing Units (TPUs), Pixel Visual Cores (PVCs), Vision Processing Units (VPUs), Graphics Processing Clusters (GPCs), Texture Processing Clusters (TPCs), Streaming Multiprocessors (SMs), Tree Traversal Units (TTUs), Artificial Intelligence Accelerators (AIAs), Deep Learning Accelerators (DLAs), Deep Learning Accelerator Clusters (XNNs), and Neural Processing Units. Units (NPUs), neural network accelerators (NNAs),Programmable Vision Accelerators (PVAs) – which may contain one or more Direct Memory Access (DMA) systems, one or more Vision / Vector Processing Units (VPUs), one or more Pixel Processing Engines (PPEs) – which may contain, for example, a 2D array of processing elements, each communicating with one or more other processing elements in the array in the north, south, east, and west, one or more decoupled accelerators or units (e.g., decoupled lookup table (DLUT) accelerators or units), etc., Vision Processing Units (VPUs), Optical Flow Accelerators (OFAs), Field-Programmable Gate Arrays (FPGAs), neuromorphic chips, Quantum Processing Units (QPUs),Associative Process Units (APUs), Arithmetic-Logic Units (ALUs), Application-Specific Integrated Circuits (ASICs), Floating-Point Units (FPUs), Input / Output (I / O) Elements, Peripheral Component Interconnect (PCI) or Peripheral Component Interconnect Express (PCIe) Elements, and / or the like.

[0266] The Communication Interface 910 can include one or more receivers, transmitters, and / or transceivers that enable the Computing Device 900 to communicate with other Computing Devices over an electronic network, including wired and / or wireless communication. The Communication Interface 910 can include components and functions that enable communication over a variety of different networks, such as wireless networks (e.g., Wi-Fi, Z-Wave, Bluetooth, Bluetooth LE, ZigBee, etc.), wired networks (e.g., Ethernet or InfiniBand communication), low-power wide area networks (e.g., LoRaWAN, SigFox, etc.), and / or the Internet.In one or more embodiments, the one or more logic units 920 and / or the communication interface 910 may contain one or more data processing units (DPUs) to transfer data received via a network and / or via the connection system 902 directly to one or more GPUs 908 (e.g., a working memory thereof).

[0267] The I / O ports 912 enable the computing device 900 to be logically coupled with other devices, including the I / O components 914, one or more presentation components 918, and / or other components, some of which may be built into (e.g., integrated with) the computing device 900. Examples of I / O components 914 include a microphone, mouse, keyboard, joystick, gamepad, game controller, satellite dish, scanner, printer, wireless device, etc. The I / O components 914 can provide a natural user interface (NUI) that processes air gestures, speech, or other physiological inputs generated by a user. In some cases, the inputs can be transmitted to a suitable network element for further processing.A NUI can implement any combination of speech capture, stylus capture, facial recognition, biometric capture, gesture capture (both on-screen and off-screen), air gestures, head and eye tracking, and touch capture (as further described below) associated with a display of the Computing Device 900. The Computing Device 900 can include depth cameras, such as stereoscopic camera systems, infrared camera systems, RGB camera systems, touchscreen technology, and combinations thereof, for gesture capture and recognition. Additionally, the Computing Device 900 can include accelerometers or gyroscopes (e.g., as part of an inertial measurement unit (IMU)) that enable motion detection. In some examples, the output from the accelerometers or gyroscopes can be used by the Computing Device 900 to render immersive augmented reality or virtual reality.

[0268] The power supply 916 can include a hardwired power supply, a battery power supply, or a combination of both. The power supply 916 can supply power to the computing device 900 to enable the operation of the computing device 900's components.

[0269] The one or more Presentation Components 918 can include a display (e.g., a monitor, a touchscreen, a television screen, a heads-up display (HUD), other display types, or a combination thereof), speakers, and / or other presentation components. The one or more Presentation Components 918 can receive data from other components (e.g., the one or more GPUs 908, the one or more CPUs 906, DPUs, etc.) and output the data (e.g., as an image, video, sound, etc.). EXEMPLARY NETWORK ENVIRONMENTS

[0270] Network environments suitable for implementing embodiments of the disclosure may include one or more client devices, servers, network-attached storage (NAS), other backend devices, and / or other device types. The client devices, servers, and / or other device types (e.g., each device) may run on one or more instances of the one or more computing devices 900 of Fig. 9 can be implemented – for example, each device can contain similar components, features, and / or functionality to one or more computing devices 900. If backend devices (e.g., servers, NAS, etc.) are implemented, the backend devices can also be included as part of a data center (such as, but not limited to, the one described herein).

[0271] The components of a network environment can communicate with each other over one or more networks, which can be wired, wireless, or both. The network can contain multiple networks or a network of networks. For example, the network can include one or more wide area networks (WANs), one or more local area networks (LANs), one or more public networks such as the internet and / or a public switched telephone network (PSTN), and / or one or more private networks. If the network includes a wireless telecommunications network, components such as a base station, a communications tower, or even access points (as well as other components) can provide wireless connectivity.

[0272] Compatible network environments can contain one or more peer-to-peer network environments—in which case a server cannot be included in a network environment—and one or more client-server network environments—in which case one or more servers can be included in a network environment. In peer-to-peer network environments, the functionality described here can be implemented with respect to one or more servers on any number of client devices.

[0273] In at least one embodiment, a network environment can include one or more cloud-based network environments, a distributed computing environment, a combination thereof, etc. A cloud-based network environment can include a framework layer, a job scheduler, a resource manager, and a distributed file system implemented on one or more servers, which may include one or more core network servers and / or edge servers. A framework layer can include a framework to support software of a software layer and / or one or more applications of an application layer. The software or the one or more applications may each include web-based service software or applications. In embodiments, one or more of the client devices can utilize the web-based service software or applications (e.g.,by accessing the service software and / or applications via one or more application programming interfaces (APIs). The framework layer can be a type of free and open-source software web application framework that uses, for example, a distributed file system for processing large amounts of data (e.g., "Big Data"), but is not limited to this.

[0274] A cloud-based network environment can provide cloud computing and / or cloud storage, performing any combination (or one or more parts thereof) of the computing and / or data storage functions described herein. Each of these different functions can be distributed across multiple locations from central or core servers (e.g., one or more data centers, which may be distributed across a state, region, country, the globe, etc.). If a connection to a user (e.g., a client device) is relatively close to one or more edge servers, one or more core servers can offload at least one part of the functionality to the one or more edge servers. A cloud-based network environment can be private (e.g., restricted to a single organization), public (e.g., available to many organizations), and / or a combination thereof (e.g., a hybrid cloud environment).

[0275] The one or more client devices can have at least some of the components, features, and functions of the one or more mentioned here in relation to Fig.The exemplary computing devices described in Section 9 include 900. By way of example, and not as a limitation, a client device may be a personal computer (PC), a laptop, a mobile device, a smartphone, a tablet computer, a smartwatch, a portable computer, a personal digital assistant (PDA), an MP3 player, a virtual reality headset, a global positioning system (GPS) or global positioning device, a video player, a video camera, a surveillance device or surveillance system, a vehicle, a boat, a hydrofoil, a virtual machine, a drone, a robot, a handheld communication device, a hospital device, a gaming device or gaming system, an entertainment system, a vehicle computer system, an embedded system controller, a remote control, a device, a consumer electronics device, a workstation, an edge device,any combination of these described devices or any other suitable device may be embodied. SAMPLE CLAUSES 1. In some embodiments, a method comprises: applying a plurality of masks to a plurality of target state attributes for an articulated object to generate one or more masked target state attributes; generating one or more actions based on at least the one or more masked target state attributes by executing a machine learning model; and updating one or more parameters of the machine learning model based on at least the one or more actions and one or more reference actions for the articulated object to generate a trained machine learning model. 2. A procedure according to clause 1, further comprising: generating one or more additional actions based on at least one or more additional target state attributes by executing the trained machine learning model; and performing a task using the articulated object based on at least the one or more additional actions. 3. Method according to one of clauses 1-2, wherein the task involves at least one of bimanual manipulation, bipedal locomotion or navigation. 4. A method according to any one of clauses 1-3, further comprising: determining a first subset of the plurality of masks at least based on a selection of one or more command spaces associated with the one or more actions; and determining a second subset of the plurality of masks at least based on a subset of the plurality of target state attributes associated with the one or more command spaces. 5. A procedure according to one of clauses 1-4, wherein the first subset of the plurality of masks is used to filter a second subset of the plurality of target state attributes that is not associated with the one or more command spaces. 6. A method according to one of clauses 1-5, wherein the second subset of the plurality of masks is determined by taking from a distribution that is assigned to each state that is contained in the subset of the plurality of target state attributes. 7. A method according to one of clauses 1-6, which further includes generating the one or more reference actions based at least on the multitude of target state attributes by executing a second trained machine learning model. 8. Method according to any of clauses 1-7, wherein the plurality of target state attributes includes at least one from a group of joint positions, a group of joint angles or a group of body center attributes. 9. Method according to any of clauses 1-8, wherein the plurality of target state attributes is assigned to at least one of a command space for kinematic position tracking, a command space for joint angle tracking or a command space for tracking the body center. 10. Method according to any one of clauses 1-9, wherein the articulated object comprises a humanoid robot. 11. In some embodiments, at least one processor comprises a processing circuit for performing operations comprising the following: applying a plurality of masks to a plurality of target state attributes, which is mapped to a plurality of instruction spaces for an articulated object, to generate one or more masked target state attributes; generating one or more actions based at least on the one or more masked target state attributes and a proprioception associated with the articulated object by executing a machine learning model; and updating one or more parameters of the machine learning model based at least on the one or more actions and one or more reference actions for the articulated object, to generate a trained machine learning model. 12. The at least one processor according to clause 11, wherein the operations further comprise: generating one or more additional actions based on at least one or more additional target state attributes associated with the articulated object by executing the trained machine learning model; and causing the articulated object to perform a movement based on at least the one or more additional actions, wherein the movement is associated with at least one of bimanual manipulation, bipedal locomotion, or navigation. 13. The at least one processor according to one of clauses 11-12, wherein the operations further comprise determining one or more additional target state attributes based at least on a control input and one or more control modes associated with the articulated object. 14. The at least one processor according to one of clauses 11-13, wherein the operations further comprise: updating one or more additional parameters of a second machine learning model based at least on a set of reference movements and a set of rewards to generate a second trained machine learning model; and generating the one or more reference actions based at least on an additional proprioception and the variety of target state attributes by executing the second trained machine learning model. 15. The at least one processor according to one of clauses 11-14, wherein the group of rewards is calculated at least on the basis of at least one of a penalty term, a regularization term or a group of task rewards. 16. The at least one processor according to any of clauses 11-15, wherein the proprioception includes at least one of a joint position, a joint velocity, a base angular velocity, a gravity vector or an action sequence. 17. The at least one processor according to any of clauses 11-16, wherein the plurality of masks comprises: a first set of masks associated with one or more selected instruction spaces contained in the plurality of instruction spaces; and a second set of masks applied to a subset of the plurality of target state attributes associated with the one or more selected instruction spaces. 18. The at least one processor according to any one of clauses 11-17, wherein the at least one processor is included in at least one of the following: a system for performing simulation operations; a system for performing digital twin operations; a system for performing collaborative content creation for 3D assets; a system for performing one or more deep learning operations; a system implemented using an edge device; a system for generating or presenting at least one type of virtual reality content, augmented reality content, or mixed reality content; a system implemented using a robot; a system for performing one or more conversational AI operations; a system implemented using one or more large language models (LLMs);a system implemented using one or more Small Language Models (SLMs); a system implementing one or more Vision Language Models (VLMs); a system implementing one or more multimodal language models; a system for generating synthetic data; a system for performing one or more operations with generative AI; a system that uses or employs one or more inference microservices; a system that includes or employs one or more machine learning models in a service or microservice together with an operating system-level virtualization package (e.g., a container); a system that includes one or more virtual machines (VMs); a system that is at least partially deployed in a data center; or a system that is at least partially deployed using cloud computing resources. 19. In some embodiments, a system comprises one or more processors for generating a trained machine learning model by training a machine learning model using a plurality of masked target state attributes, wherein the plurality of masked target state attributes is generated by applying one or more masks to a plurality of target state attributes that are associated with a plurality of instruction spaces for an articulated object. 20. System according to clause 19, wherein the system is included in at least one of the following: a system for performing simulation operations; a system for performing digital twin operations; a system for performing collaborative content creation for 3D assets; a system for performing one or more deep learning operations; a system implemented using an edge device; a system for generating or presenting at least one type of virtual reality, augmented reality, or mixed reality content; a system implemented using a robot; a system for performing one or more conversational AI operations; a system implemented using one or more large language models (LLMs); a system implemented using one or more small language models (SLMs);a system that implements one or more Vision Language Models (VLMs); a system that implements one or more multimodal language models; a system for generating synthetic data; a system for performing one or more operations with generative AI; a system that contains one or more virtual machines (VMs); a system that uses or employs one or more inference microservices; a system that contains or employs one or more machine learning models in a service or microservice together with an operating system-level virtualization package (e.g., a container); a system that is at least partially deployed in a data center; or a system that is at least partially deployed using cloud computing resources. 21. In some embodiments, a method comprises: determining a plurality of masks associated with one or more control modes for an articulated object; applying the plurality of masks to a plurality of target state attributes for the articulated object to generate one or more masked target state attributes; generating one or more actions based on at least the one or more masked target state attributes by executing a trained machine learning model; and performing a task using the articulated object based on at least the one or more actions. 22. The procedure according to clause 21, further comprising: generating one or more additional actions based on at least one or more additional target state attributes by executing the trained machine learning model; and updating one or more parameters of the machine learning model based on at least the one or more additional actions and one or more reference actions to generate the trained machine learning model. 23. A method according to one of clauses 21-22, which further comprises generating the one or more reference actions based at least on the plurality of target state attributes by executing a second trained machine learning model. 24. A method according to any of clauses 21-23, further comprising: determining the plurality of masks at least based on a selection of one or more control modes; and determining the plurality of target state attributes at least based on a control input received via a control interface. 25. A method according to any of clauses 21-24, wherein the plurality of masks is used to filter a subset of the plurality of target state attributes that is not associated with the one or more control modes. 26. Method according to any one of clauses 21-25, wherein the one or more control modes comprise a first control mode for a first part of the articulated object and a second control mode for a second part of the articulated object. 27. A method according to any of clauses 21-26, wherein the task involves at least one of bimanual manipulation, bipedal locomotion or navigation. 28. Method according to one of clauses 21-27, wherein the plurality of target state attributes includes at least one from a group of joint positions, a group of joint angles or a group of body center attributes. 29. Method according to one of clauses 21-28, wherein the plurality of target state attributes is assigned to at least one of a command space for kinematic position tracking, a command space for joint angle tracking or a command space for tracking the body center. 30. A method according to any of clauses 21-29, wherein the articulated object comprises a humanoid robot. 31. In some embodiments, at least one processor comprises a processing circuit for performing operations that include: applying a plurality of masks associated with one or more control modes for an articulated object to a plurality of target state attributes associated with a plurality of instruction spaces for an articulated object to generate one or more masked target state attributes; generating one or more actions by executing a machine learning model based at least on the one or more masked target state attributes and proprioception associated with the articulated object; and performing a task using the articulated object based at least on the one or more actions. 32. The at least one processor according to clause 31, wherein the operations further comprise: generating one or more additional actions based on at least one or more additional target state attributes associated with the articulated object by executing the trained machine learning model; and updating one or more parameters of the machine learning model based on at least the one or more additional actions and one or more reference actions for the articulated object to generate a trained machine learning model. 33. The at least one processor according to one of clauses 31-32, wherein the operations further comprise: updating one or more additional parameters of a second machine learning model based at least on a set of reference movements and a set of rewards to generate a second trained machine learning model; and generating the one or more reference actions based at least on an additional proprioception and the one or more additional target state attributes by executing the second trained machine learning model. 34. The at least one processor according to one of clauses 31-33, wherein the group of rewards is calculated at least on the basis of at least one of a penalty term, a regularization term or a group of task rewards. 35. The at least one processor according to any one of clauses 31-34, wherein the operations further comprise: determining the one or more additional target state attributes based at least on a first set of masks associated with one or more selected instruction spaces contained in the plurality of instruction spaces; and a second set of masks applied to a subset of the plurality of target state attributes associated with the one or more selected instruction spaces. 36. The at least one processor according to one of clauses 31-35, wherein the operations further include determining the plurality of masks and the plurality of target state attributes based at least on one control input and one or more control modes. 37. The at least one processor according to any of clauses 31-36, wherein the proprioception includes at least one of a joint position, a joint velocity, a base angular velocity, a gravity vector or an action sequence. 38. The at least one processor according to any of clauses 31-37, wherein the at least one processor is included in at least one of the following: a system for performing simulation operations; a system for performing digital twin operations; a system for performing collaborative content creation for 3D assets; a system for performing one or more deep learning operations; a system implemented using an edge device; a system for generating or presenting at least one type of virtual reality content, augmented reality content, or mixed reality content; a system implemented using a robot; a system for performing one or more conversational AI operations; a system implemented using one or more large language models (LLMs);a system implemented using one or more Small Language Models (SLMs); a system implementing one or more Vision Language Models (VLMs); a system implementing one or more multimodal language models; a system for generating synthetic data; a system for performing one or more operations with generative AI; a system that uses or employs one or more inference microservices; a system that includes or employs one or more machine learning models in a service or microservice together with an operating system-level virtualization package (e.g., a container); a system that includes one or more virtual machines (VMs); a system that is at least partially deployed in a data center; or a system that is at least partially deployed using cloud computing resources. 39. In some embodiments, a system comprises one or more processors for generating motion on an articulated object based at least on one or more actions generated by executing a trained machine learning model, wherein the trained machine learning model is generated by training a machine learning model using a plurality of masked target state attributes, and wherein the plurality of masked target state attributes are generated by applying one or more masks to a plurality of target state attributes that are associated with a plurality of instruction spaces for the articulated object. 40. System according to clause 39, wherein the system is included in at least one of the following: a system for performing simulation operations; a system for performing digital twin operations; a system for performing collaborative content creation for 3D assets; a system for performing one or more deep learning operations; a system implemented using an edge device; a system for generating or presenting at least one type of virtual reality, augmented reality, or mixed reality content; a system implemented using a robot; a system for performing one or more conversational AI operations; a system implemented using one or more large language models (LLMs); a system implemented using one or more small language models (SLMs);a system that implements one or more Vision Language Models (VLMs); a system that implements one or more multimodal language models; a system for generating synthetic data; a system for performing one or more operations with generative AI; a system that contains one or more virtual machines (VMs); a system that uses or employs one or more inference microservices; a system that contains or employs one or more machine learning models in a service or microservice together with an operating system-level virtualization package (e.g., a container); a system that is at least partially deployed in a data center; or a system that is at least partially deployed using cloud computing resources.

[0276] The revelation can be described in the general context of computer code or machine-usable instructions, including computer-executable instructions such as program modules, which are executed by a computer or other machine, such as a personal data assistant or other handheld device. Generally, program modules contain routines, programs, objects, components, data structures, etc., and refer to code that performs specific tasks or implements certain abstract data types. The revelation can be practiced in a variety of system configurations, including handheld devices, consumer electronics, general-purpose computers, more specialized computing devices, etc.The revelation can also be practiced in distributed computing environments, where tasks are performed by remote processing devices that are connected to each other via a network for communication.

[0277] As used herein, any mention of "and / or" in relation to two or more elements should be interpreted as referring to only one element or a combination of elements. For example, "element A, element B and / or element C" may include only element A, only element B, only element C, element A and element B, element A and element C, element B and element C, or elements A, B, and C. Furthermore, "at least one of element A or element B" may include at least one of element A, at least one of element B, or at least one of element A and at least one of element B. Additionally, "at least one of element A and element B" may include at least one of element A, at least one of element B, or at least one of element A and at least one of element B.

[0278] The subject matter of this disclosure is specifically described herein to satisfy legal requirements. However, the description itself is not intended to limit the scope of protection afforded by this disclosure. Rather, the inventors have considered that the claimed subject matter may also be embodied in other ways to include various steps or combinations of steps similar to those described in this document, in conjunction with other present or future technologies. Although the terms “step” and / or “block” may be used herein to denote various elements of the methods employed, these terms should not be interpreted as implying any particular sequence among or between the various steps disclosed herein, except where the sequence of each step is expressly described.

[0279] It is understood that the aspects and embodiments described above are purely exemplary and that modifications of details may be made within the scope of protection of the claims.

[0280] Each device, each method and each feature disclosed in the description, and (where applicable) the claims and drawings, may be provided independently or in any suitable combination.

[0281] Reference numerals appearing in the claims are for illustrative purposes only and do not restrict the scope of protection of the claims. QUOTES INCLUDED IN THE DESCRIPTION

[0000] This list of documents cited by the applicant was automatically generated and is included solely for the reader's convenience. The list is not part of the German patent or utility model application. The DPMA accepts no liability for any errors or omissions. Cited non-patent literature

[0000] Taxonomy and Definitions for Terms Related to Driving Automation Systems for On-Road Motor Vehicles” of the Society of Automotive Engineers (SAE) (Standard No. J3016-201806, published on June 15, 2018, Standard No. J3016-201609, published on September 30, 2016

[0102]

Claims

[1] Procedure comprising the following: Determining a multitude of masks that are assigned to one or more control modes for an articulated object; Applying the multitude of masks to a multitude of target state attributes for the articulated object to create one or more masked target state attributes; Generating one or more actions based on at least one or more masked target state attributes by executing a trained machine learning model; and Performing a task using the articulated object, based on at least one or more actions. [2] The method of claim 1, further comprising: Generating one or more additional actions based on at least one or more additional target state attributes by executing a machine learning model; and Updating one or more parameters of the machine learning model based on at least one or more additional actions and one or more reference actions to generate the trained machine learning model. [3] The method of claim 2, further comprising generating the one or more reference actions based at least on the plurality of target state attributes by executing a second trained machine learning model. [4] A method according to any of the preceding claims, further comprising: Determining the multitude of masks based at least on a selection of one or more control modes; and Determining the multitude of target state attributes, at least based on a control input received via a control interface. [5] Method according to any of the preceding claims, wherein the plurality of masks is used to filter a subset of the plurality of target state attributes that is not assigned to the one or more control modes. [6] Method according to one of the preceding claims, wherein the one or more control modes comprise a first control mode for a first part of the articulated object and a second control mode for a second part of the articulated object. [7] Method according to any of the preceding claims, wherein the object comprises at least one of bimanual manipulation, bipedal locomotion or navigation. [8] Method according to any of the preceding claims, wherein the plurality of target state attributes comprises at least one from a group of joint positions, a group of joint angles or a group of attributes of the body center. [9] Method according to any of the preceding claims, wherein the plurality of target state attributes is assigned to at least one of a command space for kinematic position tracking, a command space for joint angle tracking or a command space for tracking the body center. [10] Method according to any of the preceding claims, wherein the articulated object comprises a humanoid robot. [11] At least one processor comprising the following: a processing circuit for performing operations that include the following: Applying a variety of masks, associated with one or more control modes for an articulated object, to a variety of target state attributes, associated with a variety of command spaces for the articulated object, to create one or more masked target state attributes; Generating one or more actions based on at least one or more masked target state attributes and proprioception associated with the articulated object by executing a trained machine learning model; and Performing a task using the articulated object, based on at least one or more actions. [12] The at least one processor according to claim 11, wherein the operations further comprise: Generating one or more additional actions based on at least one or more additional target state attributes associated with the articulated object by executing a machine learning model; and Updating one or more parameters of the machine learning model based on at least one or more additional actions and one or more reference actions for the articulated object to create a trained machine learning model. [13] The at least one processor according to claim 12, wherein the operations further comprise: Updating one or more additional parameters of a second machine learning model, based at least on a set of reference movements and a set of rewards, to generate a second trained machine learning model; and Generating one or more reference actions based on at least one additional proprioception and one or more additional target state attributes by executing the second trained machine learning model. [14] The at least one processor according to claim 13, wherein the group of rewards is calculated at least on the basis of at least one of a penalty term, a regularization term or a group of task rewards. [15] The at least one processor according to any one of claims 12 to 14, wherein the operations further comprise determining one or more additional target state attributes based at least on one of the following: a first group of masks that is assigned to one or more selected command spaces contained within the multitude of command spaces; and a second group of masks that is applied to a subset of the multitude of target state attributes that is associated with one or more selected command spaces. [16] The at least one processor according to any one of claims 11 to 15, wherein the operations further comprise determining the plurality of masks and the plurality of target state attributes based at least on a control input and the one or more control modes. [17] The at least one processor according to one of claims 11 to 16, wherein the proprioception includes at least one of a joint position, a joint velocity, a base angular velocity, a gravity vector or an action sequence. [18] The at least one processor according to any one of claims 11-17, wherein the at least one processor is included in at least one of the following: a system for performing simulation operations; a system for performing digital twin operations; a system for conducting collaborative content creation for 3D assets; a system for performing one or more deep learning operations; a system implemented using an edge device; a system for generating or presenting at least one type of virtual reality content, augmented reality content, or mixed reality content; a system that is implemented using a robot; a system for performing one or more operations using conversational AI, a system that is implemented using one or more large language models (LLMs); a system that is implemented using one or more Small Language Models (SLMs); a system that implements one or more Vision Language Models (VLMs); a system that implements one or more multimodal language models; a system for generating synthetic data; a system for performing one or more operations using generative AI; a system that uses or employs one or more inference microservices; a system that includes or uses one or more machine learning models in a service or microservice together with an operating system-level virtualization package (e.g., a container); a system that contains one or more virtual machines (VMs); a system that is at least partially implemented in a data center; or a system that is implemented at least partially using cloud computing resources. [19] System comprising the following: one or more processors for generating a motion on an articulated object based on at least one or more actions generated by executing a trained machine learning model, wherein the trained machine learning model is generated by training a machine learning model using a plurality of masked target state attributes, and wherein the plurality of masked target state attributes are generated by applying one or more masks to a plurality of target state attributes that are mapped to a plurality of command spaces for the articulated object. [20] System according to claim 19, wherein the system is included in at least one of the following: a system for performing simulation operations; a system for performing digital twin operations; a system for conducting collaborative content creation for 3D assets; a system for performing one or more deep learning operations; a system implemented using an edge device; a system for generating or presenting at least one type of virtual reality content, augmented reality content, or mixed reality content; a system that is implemented using a robot; a system for performing one or more operations using conversational AI, a system that is implemented using one or more large language models (LLMs); a system that is implemented using one or more Small Language Models (SLMs); a system that implements one or more Vision Language Models (VLMs); a system that implements one or more multimodal language models; a system for generating synthetic data; a system for performing one or more operations using generative AI; a system that contains one or more virtual machines (VMs); a system that uses or employs one or more inference microservices; a system that includes or uses one or more machine learning models in a service or microservice together with an operating system-level virtualization package (e.g., a container); a system that is at least partially implemented in a data center; or a system that is implemented at least partially using cloud computing resources.