Automated intervention interpreter for robotics systems and applications

An automated robotics intervention interpreter using machine learning models addresses the inefficiencies of manual event analysis by detecting and classifying events, generating alerts, and updating control routines, improving robotic system performance.

DE102025148205A1Pending Publication Date: 2026-05-21NVIDIA CORP
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
DE · DE
Patent Type
Applications
Current Assignee / Owner
NVIDIA CORP
Filing Date
2025-11-20
Publication Date
2026-05-21

AI Technical Summary

Technical Problem

Manual analysis of intervention or disengagement events in autonomous or semi-autonomous robotic systems is expensive, time-consuming, prone to human error, and does not scale well with large volumes of data, and does not allow for automatic updating of control routines.

Method used

An automated robotics intervention interpreter using machine learning models analyzes sensor data to detect events, generate alerts, classify events, create training datasets, and automatically retrain control routines.

Benefits of technology

Efficiently processes large volumes of event data, reduces human error, and automatically updates control routines, enhancing the performance of robotic systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

In various examples, one or more systems use received sensor data to detect engagement or disengagement events associated with one or more autonomous or semi-autonomous robotic systems included in a fleet of robotic systems. Based on a detected engagement or disengagement event, the one or more systems can generate an alert requesting human assistance or guidance. The one or more systems can also analyze a detected engagement or disengagement event and generate a natural language description of the event, classify the event into one or more categories and / or failure modes, create new entries in a training dataset based on the detected event, or initiate automated retraining of one or more autonomous or semi-autonomous control routines based on the detected event.
Need to check novelty before this filing date? Find Prior Art

Description

BACKGROUND

[0001] An organization operating a fleet of autonomous or semi-autonomous robotic systems (e.g., autonomous mobile robots, humanoids, forklifts, vehicles, watercraft, drones, etc.) typically faces a large number of intervention or disengagement events. An intervention or disengagement event can include situations where a human or other operator (e.g., on-site and / or remotely) is required to provide assistance or guidance to a robotic system that has become stuck or is otherwise unable to continue operating in autonomous or semi-autonomous mode. The human operator may be present within the robotic system or monitor its operation on-site or remotely.Intervention or disengagement events provide valuable data points for analyzing the operation of the robot system and identifying necessary improvements to autonomous or semi-autonomous control routines. Data collected in connection with intervention or disengagement events can also be adapted for use as additional training data when retraining or otherwise refining autonomous or semi-autonomous control routines.

[0002] Typically, the review and analysis of data related to invention or disengagement events is performed manually by a human analyst. One or more human analysts must categorize events (e.g., by selecting from a pre-defined list of failure categories or failure modes). These analysts must also manually generate a natural language description of the events and conduct root cause analysis to determine the underlying cause of the intervention or disengagement event. Manual human review is expensive, time-consuming, prone to human error, and does not scale well with large volumes of event data.Furthermore, any knowledge gained through manual human revision is not automatically incorporated into updates or other changes to autonomous or semi-autonomous control routines.

[0003] Therefore, there is a need for more efficient techniques for analyzing recorded data associated with intervention or disengagement events recorded by autonomous or semi-autonomous robotic systems, and for automatically refining autonomous or semi-autonomous control based on the recorded event data. SUMMARY

[0004] The invention is defined by the claims. For the purpose of illustrating the invention, aspects and embodiments that may not fall within the scope of the claims may be described herein.

[0005] One or more systems are disclosed that use received sensor data to detect engagement or disengagement events associated with one or more autonomous or semi-autonomous robotic systems belonging to a fleet of robotic systems. Based on a detected engagement or disengagement event, the one or more systems can generate an alert requesting human assistance or guidance. The one or more systems can also analyze a detected engagement or disengagement event and generate a natural language description of the event, classify the event into one or more categories and / or failure modes, create new entries in a training dataset based on the detected event, or initiate automated retraining of one or more autonomous or semi-autonomous control routines based on the detected event.

[0006] Embodiments of the present disclosure relate to an automated robotics intervention interpreter and active learning for a machine learning model (e.g., a foundational model). Systems and methods are disclosed that collect and analyze image data or other sensor data (e.g., LiDAR, radar, ultrasound, motor control sensors, inertial measurement unit (IMU), self-motion, etc.). Based on the image or other sensor data, the disclosed systems and methods can perform online monitoring of a fleet of robotic systems, detect an intervention or disengagement event, and generate an alert requesting support or guidance from a human or robotic operator or supervisor, who may be located, for example, on-site or remotely with respect to the robot or machine experiencing the intervention or disengagement request.The disclosed systems and methods can also classify an engagement or disengagement event into one of several categories and / or failure modes, generate a natural language description of the event, or select one or more elements of event data for inclusion in an updated training or validation dataset. The disclosed systems and methods can further initiate automatic retraining of one or more autonomous or semi-autonomous control routines based on the updated training or validation dataset. The disclosed systems and methods can also identify specific features of the autonomous or semi-autonomous control routines for improvement based on the relative frequencies of different categories or failure modes of engagement or disengagement events.

[0007] In contrast to conventional approaches, the systems and methods of this disclosure employ a trained machine learning model (e.g., a "foundational machine learning model") (MLM), such as a vision language model (VLM), large language model (LLM), or multimodal language model (MMLM), to monitor a fleet of autonomous or semi-autonomous robotic systems. In an online mode, the MLM receives sensor data associated with the operation of a robotic system and analyzes this data to detect an engagement or disengagement event. The MLM can then generate an alert based on the detected event and transmit a request for assistance or intervention from a human operator or supervisor.In offline mode, the MLM, or one or more additional MLMs, analyzes event data associated with an engagement or disengagement event, generates a natural language description of the event, and assigns the event to a cluster of similar engagement types. In offline mode, one or more MLMs can also identify one or more events to include in an updated training or validation dataset and can automatically trigger the retraining of one or more autonomous or semi-autonomous control routines based on the updated training or validation dataset. The disclosed techniques can also select a subset of event data for permanent storage, reducing memory requirements by discarding redundant event data.

[0008] Further features of the disclosure are characterized by the independent and dependent claims.

[0009] Any feature of one aspect of the disclosure can be applied to other aspects of the disclosure in any suitable combination. In particular, procedural aspects can be applied to device or system aspects and vice versa.

[0010] Furthermore, features implemented in hardware can be implemented in software and vice versa. Any reference to software and hardware features herein should be interpreted accordingly.

[0011] Each system or device feature as described herein can also be provided as a process feature, and vice versa. System and / or device aspects described functionally (including means plus functional features) can alternatively be expressed in terms of their corresponding structure, such as a suitably programmed processor and allocated working memory.

[0012] It is also understood that special combinations of the various features described and defined in aspects of the revelation can be implemented and / or supplied and / or used independently.

[0013] The disclosure also provides computer programs and computer program products comprising software code which, when executed on a data processing device, is adapted to perform one of the procedures and / or to embody one of the device and system features described herein, including all component steps of any procedure.

[0014] The disclosure also provides a computer or computing system (including networked or distributed systems) comprising an operating system that provides a computer program for performing the procedures described herein and / or for embodying any device or system features described herein.

[0015] The revelation also provides a computer-readable medium on which one or more of the aforementioned computer programs are stored.

[0016] The revelation also provides a signal that carries one or more of the aforementioned computer programs.

[0017] The disclosure extends to methods and / or devices and / or systems as described herein with reference to the accompanying drawings.

[0018] Aspects and embodiments of the present disclosure will now be described purely by way of example with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] The present systems and methods for an automated robotics intervention interpreter and active learning via a fundamental model are described in detail below with reference to the attached drawings, in which: Fig. 1 is a block diagram of an exemplary computing device suitable for use in implementing some embodiments of the present disclosure; Fig. 2 a more detailed illustration of the monitoring engine of the Fig. 1 according to different embodiments; Fig. 3. A flowchart of a procedure for monitoring a fleet of robot systems according to various embodiments is illustrated; Fig. 4 a more detailed illustration of the inference engine of the Fig. 1 according to different embodiments; Fig. 5 illustrates a flowchart of a procedure for analyzing event data according to different embodiments; Fig. 6A is an illustration of an exemplary autonomous vehicle according to some embodiments of the present disclosure; Fig. 6B is an example of camera positions and fields of view for the exemplary autonomous vehicle of the Fig. 6A according to some embodiments of the present disclosure is; Fig. 6C a block diagram of an exemplary system architecture for the exemplary autonomous vehicle of the Fig. 6A according to some embodiments of the present disclosure is; Fig. 6D a system diagram for communication between (a) cloud-based server(s) and the exemplary autonomous vehicle of the Fig. 6A according to some embodiments of the present disclosure. Fig. 7 is a block diagram of an exemplary computing device suitable for use in implementing some embodiments of the present disclosure, and Fig. 8 is a block diagram of an exemplary data center suitable for use in implementing some embodiments of the present disclosure. Fig. 9A is a block diagram of an exemplary generative language model system suitable for use in implementing at least some embodiments of the present disclosure. Fig. 9B is a block diagram of an exemplary generative language model that includes a transformer-encoder-decoder suitable for use in implementing at least some embodiments of the present disclosure. Fig. 9C is a block diagram of an exemplary generative language model that includes a decoder-transformer-only architecture suitable for use in implementing at least some embodiments of the present disclosure. DETAILED DESCRIPTION

[0020] Systems and methods relating to automated robotics intervention design and active learning with a fundamental model for autonomous or semi-autonomous systems and applications are disclosed. Although the present disclosure may be described with reference to an exemplary autonomous or semi-autonomous robot, an exemplary autonomous or semi-autonomous vehicle, or an exemplary autonomous or semi-autonomous machine 600 (hereinafter alternatively referred to as "Vehicle 600", "Robot 600", "Machine 600", or "Ego-Machine 600"), for which an example is given with reference to the Fig. The fact that the systems and methods described in sections 6A-6D are limited is not intended to be restrictive. For example, the systems and methods described herein may be used without restriction by robotics systems that include non-autonomous vehicles or machines, semi-autonomous vehicles or machines (e.g., in one or more adaptive driver assistance systems (ADAS)), autonomous vehicles or machines, steered and unsteered robots, warehouse vehicles, all-terrain vehicles, vehicles coupled with one or more trailers, hydrofoils, boats, shuttles, emergency vehicles, motorcycles, electric or motorized bicycles, aircraft, construction vehicles, underwater vehicles, drones, and / or other types of vehicles.Additionally, although the present disclosure may be described with reference to autonomous or semi-autonomous operation of a robotic system, this is not intended to be limiting, and the systems and methods described herein may be used in augmented reality, virtual reality, mixed reality, security and surveillance, autonomous or semi-autonomous machine applications, and / or any other technology domains where the detection and analysis of intervention or disengagement events may be applicable. While the present disclosure is mainly described using examples of sensors in the form of cameras or motor-driven sensors, the disclosed techniques may also be used with any suitable type of sensor, e.g., audio, LiDAR, radar, ultrasound, etc.

[0021] As explained herein, conventional techniques may require human analysis or interpretation of intervention or disengagement events. An intervention or disengagement event may include a situation in which an autonomous or semi-autonomous machine has become stuck or otherwise requires human assistance or guidance to return to autonomous or semi-autonomous operation.

[0022] Human analysis can include classifying the event into one or more predefined event types. Interpretation can include generating a natural language description of the event or performing root cause analysis to determine one or more causes of the intervention or disengagement event. These techniques require expensive and time-consuming human interaction and are subject to human error. Furthermore, manual review and analysis techniques do not scale adequately to process large volumes of intervention or disengagement event data. Traditional manual review techniques also do not allow for the automatic updating of test or validation datasets or the automatic retraining of autonomous or semi-autonomous control routines based on newly captured event data.

[0023] To improve the monitoring of a fleet of autonomous or semi-autonomous robotic systems, the disclosed techniques receive real-time sensor data from a robotic or other autonomous or semi-autonomous system and analyze the sensor data using one or more trained machine learning models to detect an intervention or disengagement event. The disclosed techniques can then generate an alert associated with the event and transmit an intervention request to a human operator or other operator or supervisor (e.g., robotics, computing, etc.). In an offline mode, the disclosed techniques can analyze event data associated with an intervention or disengagement event, including event data associated with actions taken by humans or others, to resolve the event.The disclosed techniques can then classify the event into one or more event types and / or failure modes and generate a natural language description of the event. The disclosed techniques can select data associated with an event to include in an updated test or validation dataset and automatically trigger the retraining of one or more autonomous or semi-autonomous control routines based on the updated training or validation dataset. The disclosed techniques can further select a subset of relevant event data for permanent or other long-term storage. Selecting a subset of event data for storage reduces computational and storage costs because the disclosed techniques can choose to discard data associated with ordinary or redundant events.

[0024] In online mode, a monitoring engine receives real-time sensor data from an autonomous or semi-autonomous robotic system. This robotic system could be, for example, one of several robotic systems included in a fleet operated by a single organization or company. The sensor data can include, without limitation, image, audio, LiDAR, radar, sonar, or ultrasonic sensor data. It can also include motor control sensor data, such as the position or movement of a motorized component within the robotic system. Furthermore, the motor control sensor data can include feedback data, such as measured position, movement, weight, load, stress, or torque associated with a component within the robotic system.For example, one or more IMUs can be used together with wheel or other motion information to track a machine's self-movement over time and to understand the pose and position of the robotic system.

[0025] The monitoring engine includes a machine learning model (e.g., a "foundational machine learning model") such as a large language model (LLM), vision language model (VLM), and / or multimodal language model (MMLM). The machine learning model analyzes real-time sensor data and detects an intervention or disengagement event, indicating that a robot system has become stuck or requires human or other assistance (e.g., robotics, computational assistance, etc.) or guidance to resume normal autonomous or semi-autonomous operation. Based on the detection of an intervention or disengagement event, the monitoring engine can generate an alert and transmit it to a system that can be operated by a human, robotics, computer-based, and / or other type of operator or supervisor.The operator or supervisor can be located within the robot system or geographically distant from it. The monitoring engine also records any sensor data received by the robot system that results from human input necessary to resolve the engagement or disengagement event. For example, the monitoring engine can record motor control sensor data generated by human input to the robot system. The monitoring engine can also record audio data, such as a human narration of the engagement or disengagement event or any corrective steps taken to resolve it.The monitoring engine records all received sensor data associated with the event or human corrective actions and creates an entry in an event database for later offline analysis by an inference engine.

[0026] The monitoring engine can be implemented on-premises within an organization's enterprise computing environment. Alternatively, the monitoring engine can be implemented as a remote cloud service. When implemented as a remote cloud service, the monitoring engine can perform monitoring and inference as a microservice that is made available to multiple different organizations, such as the inference microservice described in more detail herein.

[0027] In offline mode, the inference engine retrieves an entry from the event database. This event database entry includes sensor data associated with a specific historical intervention or disengagement event and sensor data associated with any human corrective action taken to resolve the event. A machine learning model included in the inference engine can perform one or more actions in offline mode based on the event database entry. As explained herein, the machine learning model can be the same as, or similar to, the foundational machine learning model included in the monitoring engine, if implemented herein.

[0028] The fundamental model can generate a natural language description of the event associated with the retrieved event database entry. The fundamental model can also cluster the event into one or more predefined intervention types and / or failure modes, such as navigation failure, sensor failure, obstacle avoidance failure, or motor control failure.

[0029] The foundational model can update a training and validation dataset based on the retrieved event database entry. The foundational model can compare one or more existing entries included in the training and validation dataset with the retrieved event database entry and determine whether the retrieved event database entry represents a unique or unusual engagement or disengagement event. The inference engine can update the training and validation dataset based on a unique or unusual event, thus providing a greater variety of event data in the training and validation dataset.

[0030] The inference engine can trigger automated retraining of an autonomy stack, where the autonomy stack includes one or more autonomous or semi-autonomous control routines used to control a fleet of robotic systems. The inference engine can trigger automated retraining on every update of the training and validation dataset, after a predetermined number of updates to the training and validation dataset, or based on a predetermined schedule.

[0031] The inference engine can evaluate the received event database entry for permanent or long-term storage. The underlying model can compare the received event database entry with one or more previously stored events and select or reject the received event database entry for storage based on similarities between the received event database entry and the one or more previously stored events. The inference engine may select an entry for storage if it determines that the entry is unique, unusual, or underrepresented compared to the previously stored events. Conversely, the inference engine may reject an entry for storage if the entry would be unnecessarily duplicated in light of the previously stored events.Selective event evaluation for long-term storage reduces the computing and storage requirements of an organization's computing environment.

[0032] Fig. Figure 1 is a block diagram of an exemplary computing device suitable for use in implementing some embodiments of the present disclosure. It is understood that this and other arrangements described herein are presented only as examples. Other arrangements and elements (e.g., machines, interfaces, functions, sequences, groupings of functions, etc.) may be used in addition to or instead of those shown, and some elements may be omitted entirely. Furthermore, many of the elements described herein are functional units that may be implemented as separate or distributed components, or in conjunction with other components, in any suitable combination and at any suitable location. Various functions described herein as being performed by units may be performed by hardware, firmware, and / or software.For example, various functions can be performed by a processor executing instructions stored in the working memory. In some embodiments, the systems, methods, and processes described herein can be implemented using similar components, features, and / or functionality to those of the exemplary autonomous vehicle 600 of the [reference to relevant document]. Fig. 6A-6D, the exemplary computing device 700 of the Fig. 7, of the exemplary data center 800 of the Fig. 8 and / or machine learning models of Fig. 9A-9C will be executed.

[0033] In one embodiment, the computing device 100 includes a desktop computer, a laptop computer, a smartphone, a personal digital assistant (PDA), a tablet computer, a vehicle-mounted or robot-integrated computing device, or any other type of computing device configured to receive input, process data, and optionally display images, and capable of performing one or more embodiments. The computing device 100 is configured to run a monitoring engine 122 and an inference engine 124, which reside in a working memory 116.

[0034] It should be noted that the computing device described herein is for illustrative purposes only, and any other technically feasible configurations fall within the scope of this disclosure. For example, multiple instances of the Monitoring Engine 122 and the Inference Engine 124 could run on a set of nodes in a distributed and / or cloud computing system to implement the functionality of the computing device 100. In other examples, the Monitoring Engine 122 and the Inference Engine 124 could run on different sets of hardware, types of devices, or in different environments to adapt the Monitoring Engine 122 or the Inference Engine 124 to different use cases or applications. In a third example, the Monitoring Engine 122 and the Inference Engine 124 could run on different computing devices and / or different sets of computing devices.

[0035] In one embodiment, the computing device 100 includes, without limitation, a connection (bus) 112 connecting one or more processors 102, an input / output device interface 104 (I / O device interface) coupled to one or more input / output devices (I / O devices) 108, a working memory 116, a storage device 114, and a network interface 106. The processor(s) 102 can be any suitable processor, implemented as a central processing unit (CPU), graphics processing unit (GPU), application-specific integrated circuit (ASIC), field-programmable gate array (FPGA), artificial intelligence accelerator (AI accelerator), any other type of processing unit, or a combination of different processing units, such as a CPU configured to work in conjunction with a GPU.In general, the processor(s) 102 can be any technically feasible hardware unit capable of processing data and / or executing software applications. Furthermore, in the context of this disclosure, the computing elements shown in the computing device 100 can correspond to a physical computing system (e.g., a system in a data center) or can be a virtual computing instance running within a computing cloud.

[0036] I / O devices 108 include devices capable of providing input, such as a keyboard, mouse, touchscreen, and so forth, as well as devices capable of providing output, such as a display device. Additionally, the I / O devices 108 may include devices capable of both receiving input and providing output, such as a touchscreen, a universal serial bus port (USB port), and so forth. I / O devices 108 can be configured to receive various types of input from an end user (e.g., a designer) of the computing device 100 and also to output various types of output to the end user of the computing device 100, such as displayed digital images, digital videos, or text. In some embodiments, one or more of the I / O devices 108 are configured to connect the computing device 100 to a network 110.

[0037] Network 110 is a technically feasible type of communication network that allows data to be exchanged between the computing device 100 and external units or devices, such as a web server or another networked computing device. Network 110 can include, for example, a Wide Area Network (WAN), a Local Area Network (LAN), a wireless network (WiFi network), and / or the Internet.

[0038] Memory 114 includes non-volatile memory for applications and data and can include fixed or removable hard disk drives, flash memory devices, and CD-ROM, DVD-ROM, Blu-ray, HD-DVD, or other magnetic, optical, or solid-state storage devices. The monitoring engine 122 and inference engine 124 can be stored in memory 114 and loaded into main memory 116 when executed.

[0039] The main memory 116 includes a random access memory module (RAM module), a flash memory unit, or any other type of main memory unit, or a combination thereof. The processor(s) 102, the I / O device interface 104, and the network interface 106 are configured to read from and write data to the main memory 116. The main memory 116 includes various software programs that can be executed by the processor(s) 102 and application data associated with the software programs, including the monitoring engine 122 and the inference engine 124.

[0040] Fig. 2 a more detailed illustration of the monitoring engine 122 of the Fig. 1 according to various embodiments. It is understood that these and other arrangements described herein are presented only as examples. Other arrangements and elements (e.g., machines, interfaces, functions, sequences, groupings of functions, etc.) may be used in addition to or instead of those shown, and some elements may be omitted entirely. Furthermore, many of the elements described herein are functional units that may be implemented as separate or distributed components, or in conjunction with other components, in any suitable combination and at any suitable location. Various functions described herein as being performed by units may be performed by hardware, firmware, and / or software. For example, various functions may be performed by a processor executing instructions stored in memory.

[0041] The monitoring engine 122 can receive sensor data 210, detect an engagement or disengagement event associated with an autonomous or semi-autonomous robot system based on the received sensor data, and generate an alert based on the detected engagement or disengagement event. The monitoring engine can also transmit the generated alert to one or more human or other operators (e.g., computer system operators, robotics operators, etc.) and / or supervisors. The monitoring engine 122 can also create an entry in the event database 240 based on the detected engagement or disengagement event and any human or other guidance or support associated with the detected event. The monitoring engine 122 includes, among other elements, the underlying model 220 (or, more generally, a machine learning model) and an alert generator 230.

[0042] As an overview, the training engine 122 can be configured to generate, process, preprocess, augment, and / or otherwise prepare sensor data 210 for analysis by the trained foundational model 220. In embodiments that include the warning generator 230, the warning generator 230 can be configured to generate a warning when it is determined that the sensor data 210 indicates an engagement or disengagement event associated with an autonomous or semi-autonomous robot system.

[0043] In various embodiments, the sensor data 210 can include sensor data generated by a robotic system using any number of sensors (physical and / or virtual or simulated), such as LiDAR sensor(s) 664, RADAR sensor(s) 660, ultrasonic sensor(s) 662, microphone(s) 696, and / or other sensor types. The sensor data can represent fields of view and / or sensor fields of sensors (e.g., LiDAR sensor(s) 664, RADAR sensor(s) 660, etc.) and / or can represent a perception of the environment by one or more sensors (e.g., microphone(s) 696). Sensors such as image sensors (e.g., from cameras), LiDAR sensors, RADAR sensors, SONAR sensors, ultrasonic sensors and / or the like may herein be referred to as perception sensors or perception sensor devices, and the sensor data produced by the perception sensors may herein be referred to as perception sensor data.In some examples, an instance or representation of the sensor data can be represented by an image (e.g., the image data) captured by an image sensor, a depth map generated by a LiDAR sensor, and / or the like. Audio data, LiDAR data, SONAR data, RADAR data, and / or other sensor data types can be correlated with or associated with image data generated using one or more image sensors. Image data representing one or more images can, for example, be updated to include data relating to LiDAR sensors, SONAR sensors, RADAR sensors, and / or the like, so that the sensor data used as input to the fundamental model 220 has more information or detail than image data alone.

[0044] In embodiments where the sensor data 210 are used, the sensors can be calibrated such that the sensor data are assigned pixel coordinates in the image data. In some embodiments, such as when the sensor data specifies depth (e.g., radar data, LiDAR data, etc.), the depth values ​​can be correlated with pixel coordinates in the image data and then used as an additional (or, in some examples, alternative) input to the fundamental model 220. For example, one or more of the pixels may have an additional value that is representative of depth, as determined from the sensor data.

[0045] In various embodiments, the basic model 220 includes a trained machine learning model that is a vision language model (VLM), a multimodal language model (MMLM), or any of the models described herein. Fig. The basic model 220 includes, but is not limited to, the machine learning model explained in 9A-9C. It analyzes sensor data 210 and detects an engagement or disengagement event associated with an autonomous or semi-autonomous robot system. As explained herein, the basic model 220 can be used to describe the Fig. 4. Also analyze sensor data associated with an engagement or disengagement event to generate a natural language description of the event, classify the event into one or more predefined categories and / or failure modes, generate new entries for a training dataset, or initiate automatic retraining of an autonomy stack that includes one or more autonomous or semi-autonomous control routines.

[0046] In various embodiments, the foundational model 220 can be trained "from scratch" to detect intervention or disengagement events based on domain-specific training data generated by an organization from the historical operation of one or more robotic systems included in a fleet of robotic systems. Domain-specific training data, encompassing thousands or tens of thousands of hours of annotated historical sensor data, can be used to fully train the foundational model 220 to detect and analyze intervention or disengagement events in a fleet of autonomous or semi-autonomous robotic systems.

[0047] In other embodiments, the foundational model 220 can include a pre-trained machine learning model, which is then further trained or refined with a relatively smaller set of annotated domain-specific training data. Domain-specific training data, comprising dozens or hundreds of hours of annotated historical sensor data, can be used to refine a pre-trained machine learning model included in the foundational model 220 such that the refined machine learning model can analyze sensor data 210 and detect and analyze engagement or disengagement events in a fleet of autonomous or semi-autonomous systems.

[0048] In embodiments that include the warning generator 230, the warning generator 230 can be configured to generate a warning when the basic model 220 determines that the sensor data 210 indicates an engagement or disengagement event associated with a robot system. The generated warning can include a subset of sensor data 210, such as one or more images and / or audio recordings. The generated warning can also include an identifier associated with a specific robot system and / or a location associated with the robot system.

[0049] The monitoring engine 122 can also transmit the generated alert to one or more human operators and / or supervisors. These operators and / or supervisors may be located at the same location as the robot system, for example, a human operator of an autonomous or semi-autonomous forklift. The operators and / or supervisors may also include teleoperators who are located at the same general location as the robot system or remotely. The generated alert may prompt the operators and / or supervisors to take control of the robot system, either locally or remotely, and to provide guidance or assistance in returning the robot system to normal autonomous or semi-autonomous operation.The monitoring engine 122 can transmit the generated alert via any suitable means, including but not limited to email, text message, organizational messaging application, or a dedicated receiver located in the same place as the operator or supervisor. In various embodiments, dedicated receivers can be operated to generate audible and / or visual alerts, such as graphical and / or text displays associated with an alert.

[0050] In embodiments that include the event database 240, the monitoring engine 122 can generate an entry in the event database 240 associated with a detected engagement and disengagement event. The generated entry can include sensor data 210 which, when processed by the basic model 220, cause the basic model 220 to detect the event. The generated entry can also include sensor data 210 associated with support or guidance provided to a robot system by a human operator or supervisor in response to the detected event. In various embodiments, the sensor data 210 associated with human support or guidance can include sensor data associated with motor control devices included in the robot system.Supportive human control inputs provided to the robot system can be recorded as positions, movements, orientations, loads, stresses, or torques, as indicated or learned by one or more motor-controlled arrangements enclosed within the robot system and included in sensor data 210. The sensor data 210 can also include audio recordings, such as a description by a human operator or supervisor of the engagement or disengagement event, or a narration of the support actions taken to return the robot system to normal autonomous or semi-autonomous operation. The monitoring engine 122 transmits the event database 240 to the inference engine 124, which is described herein.

[0051] Although examples relating to the use of language models and specifically vision language models or multimodal language models are described here as the basic model 220, this should not be understood as restrictive. For example, and without limitation, the fundamental model 220 described herein can include one or more of any type of machine learning model, such as one or more machine learning models that use linear regression, logistic regression, decision trees, support vector machines (SVM), Naive Bayes, k-nearest neighbor (Knn), K-means clustering, random forest, dimensionality reduction algorithms, gradient boosting algorithms, neural networks (e.g., auto-encoder, convolution, recursion, perceptrons, long / short term memory (LSTM), Hopfield, Boltzmann, deep belief, deconvolutional, generative adversarial, liquid state machine, etc.) and / or other types of machine learning models.

[0052] The monitoring engine 122 is described by way of example and without limitation with reference to one or more MLMs trained for use in computer vision and / or perception operations to monitor the operation of a robot system. However, aspects of the disclosure apply more broadly to any form of MLM trained and / or used to make predictions based on sensor data. In some examples, the basic model 220 can be trained to predict path points, vehicle orientation (e.g., with respect to environmental features such as lane markings), and / or vehicle state (e.g., with respect to an object maneuver such as a lane change, turn, merge, etc.) that can be used to control an autonomous vehicle. This is not intended to be limiting, however.

[0053] Additionally, the monitoring engine 122 is an example of a machine that can be used in at least one embodiment, such as to perform an MLM for use in computer vision and / or perception operations, for navigating a vehicle, or for other purposes. However, the monitoring engine 122 can be varied to include more, fewer, and / or different components and / or processing paths than those described in [reference to be added]. Fig. 2 is shown, to include.

[0054] The monitoring engine 122 can be implemented in a cloud computing environment and made available to one or more customers or clients as a monitoring microservice for fleets of autonomous or semi-autonomous robot systems. In various embodiments where the monitoring engine is run as a monitoring microservice, each of the one or more instances of the fundamental model 220 can be trained or refined based on different historical client fleet operational data. In other embodiments, one or more instances of the fundamental model can be trained or refined on the same historical fleet operational data.

[0055] With reference to the Fig. 3 and Fig. 5. Each operation of Procedures 300 and 500 described herein comprises a computational process that can be performed using any combination of hardware, firmware, and / or software. For example, various functions can be performed by a processor executing instructions stored in memory. The process can also be embodied as computer-usable instructions stored on computer storage media. The procedures can be provided by a standalone application, a service or hosted service (standalone or in combination with another hosted service), or a plug-in to another product, to name only a few. Additionally, Procedures 300 and 500 are illustrated by reference to the systems of Fig. 1-2 and 4. However, these procedures can additionally or alternatively be performed by any system or combination of systems, including, but not limited to, any of the systems described herein.

[0056] Fig. Figure 3 illustrates a flowchart of a method for recording engagement or disengagement events according to various embodiments. As shown in Fig. As shown in Figure 3, the procedure 300 begins with operation 302, in which the monitoring engine 122 receives sensor data 210 from a robotic system, such as the vehicle described in the descriptions of the Fig. 6A-6C as explained herein, receives. In various embodiments, the sensor data 210 includes, but is not limited to, image data, video data, audio recordings, data from one or more LiDAR, RADAR, SONAR, ultrasonic or infrared sensors.

[0057] Method 300, in operation 304, includes detecting, via the trained fundamental model 220 and based on sensor data 210, an engagement or disengagement event associated with the robot system, where the robot system has become stuck or otherwise requires human intervention to return to a normal autonomous or semi-autonomous operating mode. In various embodiments, the fundamental model 220 includes a trained machine learning model that is a vision-language model, a multimodal language model, or any of those described herein. Fig. 9A-9C includes, but is not limited to, the machine learning models explained.

[0058] In various embodiments, the foundational model 220 can be trained entirely on a relatively large corpus of annotated domain-specific training data generated by an organization from the historical operation of one or more robot systems included in a fleet of robot systems. In other embodiments, the foundational model 220 can include a pre-trained machine learning model that is then further trained or refined with a relatively smaller set of annotated domain-specific training data.

[0059] Procedure 300, in Operation 306, includes generating and transmitting an alert based on the detected intervention or disengagement event. The generated alert indicates a need for human intervention or assistance to return the robot system to a normal mode of autonomous or semi-autonomous operation. The monitoring engine 122 can transmit the generated alert to a human operator located at the same location as the robot system. Alternatively or additionally, the monitoring engine 122 can transmit the generated alert to a human supervisor located in the same geographical area as the robot system or at a distance from the robot system.

[0060] Method 300, in operation 308, includes receiving additional sensor data from the robot system associated with the human intervention or assistance provided to the robot system in response to the generated warning. In various embodiments, the additional sensor data 210 associated with human assistance or guidance may include sensor data associated with motor control devices included in the robot system. Assistive human control inputs provided to the robot system may be recorded as positions, movements, orientations, loads, stresses, or torques indicated or learned by one or more motor-controlled arrangements included in the robot system and included in the sensor data 210.The sensor data 210 can also include audio recordings, such as a description by a human operator or supervisor of the intervention or disengagement event, or a narration of the support actions taken to return the robot system to normal autonomous or semi-autonomous operation.

[0061] Procedure 300, in operation 310, includes generating an entry in the event database 240 based on the detected intervention or disengagement event of sensor data 210 and additional sensor data 210. The generated entry may include sensor data 210 which, when processed by the fundamental model 220, causes the fundamental model 220 to detect the intervention or disengagement event. The generated entry may also include additional sensor data 210 associated with support or guidance provided to a robot system by a human operator or supervisor in response to the detected event. The monitoring engine 122 transmits the event database 240 to the inference engine 124, which is described herein.

[0062] Fig. Figure 4 is a more detailed illustration of the inference engine 124 of the Fig. 1 according to various embodiments. It is understood that these and other arrangements described herein are presented only as examples. Other arrangements and elements (e.g., machines, interfaces, functions, sequences, groupings of functions, etc.) may be used in addition to or instead of those shown, and some elements may be omitted entirely. Furthermore, many of the elements described herein are functional units that may be implemented as separate or distributed components, or in conjunction with other components, in any suitable combination and at any suitable location. Various functions described herein as being performed by units may be performed by hardware, firmware, and / or software. For example, various functions may be performed by a processor executing instructions stored in memory.

[0063] As an overview, the inference engine 124 can receive an entry from the event database 240 and perform one or more of several actions, including, but not limited to, generating a natural language description of an engagement or disengagement event, classifying an engagement or disengagement event into one or more categories and / or failure modes, and selecting an entry received from the event database 240 for permanent or long-term storage in memory 114. The inference engine 124 can also select one or more elements of event data to include in the training dataset 400. Furthermore, based on the updated training or validation dataset, the inference engine 124 can initiate automatic retraining of an autonomy stack that includes one or more autonomous or semi-autonomous control routines.The inference engine 124 can also identify specific features of autonomous or semi-autonomous control routines for improvement based on the relative frequencies of different categories and / or failure modes of engagement or disengagement events. The inference engine 124 can include, among other elements, the foundational model 220, a natural language generator 410, an event classifier 420, a training data generator 430, the autonomy stack trainer 440, and an event selection module 450. Each of the natural language generator 410, the event classifier 420, the training data generator 430, and the event selection module 450 can represent a different input / output mechanism included in the foundational model 220.

[0064] The inference engine 124 receives an event entry contained in the event database 240 and transmitted by the monitoring engine 122. As explained herein, an entry contained in the event database 240 can include sensor data associated with an engagement or disengagement event detected by the monitoring engine 122. In embodiments that include the natural language generator 410, the inference engine 124 can, via the underlying model 220 and the natural language generator 410, generate a natural language description of the engagement or disengagement event associated with the received event entry.The Natural Language Generator 410 can provide the Basic Model 220 with a text prompt instructing the Basic Model 220 to generate the natural language description based on sensor data included in the received entry and associated with the engagement or disengagement event itself, or sensor data included in the received entry and associated with the human assistance input provided to a robot system in response to the engagement or disengagement event. The Natural Language Generator 410 can receive the natural language description from the Basic Model 220 and transmit the natural language description to the Inference Engine 124.

[0065] In embodiments that include the event classifier 420, the inference engine 124 can classify an event entry into one or more predefined categories of intervention types and / or failure modes via the foundational model 220 and the event classifier 420. In various embodiments, the one or more predefined categories of intervention types can include, but are not limited to, navigation failure, sensor failure, obstacle avoidance failure, or motor control failure. The event classifier 420 can generate a text prompt that instructs the foundational model 220 to classify an event entry and transfer the classification from the foundational model 220 to the inference engine 124.In various embodiments, the event classifier 420 can also generate a prompt instructing the basic model 220 to determine a failure mode associated with the event entry, the failure mode providing more specificity than a general category of intervention types. For example, the basic model 220 can classify an event entry as a navigation failure and further determine a failure mode that specifies a failure or deterioration of a particular sensor included in the robot system.

[0066] In embodiments that include the training data generator 430, the inference engine 124 can generate a new entry in the training data set 400 via the training data generator 430. The training data set 400 includes one or more entries, each input including sensor data associated with an engagement or disengagement event, and sensor data associated with human assistance and / or guidance inputs provided in response to the engagement or disengagement event. Inputs included in the training data set 400 include test and validation data used to train an autonomy stack, which includes one or more autonomous or semi-autonomous control routines used to control a fleet of robotic systems.The training data generator 430 calculates similarities between the event entry received from the event database 240 and one or more inputs included in the training data set 400. Based on these calculated similarities, the training data generator 430 determines whether the received event entry is unusual or atypical compared to the one or more entries included in the training data set 400. Based on this determination, the training data generator 430 can instruct the foundational model 220 to create and format a new entry for the training data set 400 based on the received event entry. The inference engine 124 can then record the new entry in the training data set 400.If the training data generator 430 determines that the received event entry is duplicate or redundant in view of one or more existing inputs included in the training data set 400, the training data generator 430 may stop analyzing the received event entry.

[0067] In embodiments that include the autonomy stack trainer 440, the inference engine 124 can initiate automated retraining of one or more autonomous or semi-autonomous control routines contained in an autonomy stack via the autonomy stack trainer 440, based on one or more new inputs generated by the training data generator 430 and added to the training data set 400. In some embodiments, the inference engine 124 can initiate automated retraining of the autonomy stack immediately after generating a new entry in the training data set 400. In other embodiments, the inference engine 124 can initiate automated retraining upon generating a predetermined number of new inputs in the training data set 400 or according to a predetermined schedule.

[0068] In embodiments that include the event selection module 450, the event selection module 450 can be configured to record an event entry received from the event database 240 in permanent or other long-term storage, such as memory 114. The event selection module 450 can compare the event entry with existing entries included in memory 114 based on the sensor data 210 included in the event entry, which is generated by the natural language generator 410 and associated with the event entry, and / or a category into which the event entry has been classified by the event classifier 420.Based on the comparison, the event selection module 450 can select the event entry for storage in memory 114 if the event entry and its associated natural language and / or classification description are unique, unusual, or otherwise underrepresented in memory 114. Alternatively, the event selection module 450 can reject an event entry for permanent or other long-term storage if the event would be duplicated or redundant given the existing event entries included in memory 114. Evaluating individual event entries for permanent or long-term storage reduces the computational and memory requirements compared to storing all received event entries in memory 114.

[0069] The Inference Engine 124 can analyze multiple event entries over time and suggest specific areas for improvement in autonomous or semi-autonomous control routines included in an autonomy stack, based on these multiple event entries. For example, the Inference Engine 124 can determine that a threshold percentage of event entries has been classified as navigation failures by the Foundation Model 220 and the Event Classifier 420, and can suggest a review of the one or more navigation routines included in the autonomy stack. In various embodiments, the Inference Engine 124 can also suggest a revision of one or more control or navigation routines, or robot / vehicle components, based on specific failure modes determined by the Foundation Model 220, rather than solely based on event classification categories.The Inference Engine 124, for example, can propose a revision of both navigation routines and robot system navigation sensors based on a threshold number of event entries classified as navigation failures with a corresponding specific failure mode indicating a failure or deterioration of one or more navigation sensors included in a robot system, such as LiDAR or RADAR sensors.

[0070] Fig. Figure 5 illustrates a flowchart of a procedure for analyzing event data according to various embodiments. As in Fig. As shown in Figure 5, procedure 500 begins with operation 502, in which the inference engine 124 receives an event entry from the event database 240. The event entry is associated with an engagement or disengagement event that is linked to a robot system and is captured by the monitoring engine 122. The event entry includes sensor data 210 associated with the engagement or disengagement event, as well as sensor data 210 associated with human assistance or guidance inputs provided to the robot system to reset the robot system to normal autonomous or semi-autonomous operation.

[0071] Method 500 includes, in operation 504, the generation of a natural language description of an engagement or disengagement event associated with the event entry. In various embodiments, the natural language generator 410 can provide a text prompt in the base model 220 that instructs the base model 220 to generate the natural language description based on sensor data included in the received entry and associated with the engagement or disengagement event itself, or sensor data included in the received entry and associated with the human assistance input provided to a robot system in response to the engagement or disengagement event.The natural language generator 410 can receive the natural language description from the basic model 220 and transmit the natural language description to the inference engine 124.

[0072] In operation 506, method 500 includes classifying the engagement or disengagement event associated with the event entry into one or more categories and / or determining a failure mode associated with the event entry. The event classifier 420 can generate a text prompt that instructs the basic model 220 to classify an event entry into one or more predefined categories of event types. In various embodiments, the one or more predefined categories of engagement types may include, but are not limited to, navigation failure, sensor failure, obstacle avoidance failure, or motor control failure. The event classifier 420 can transfer the classification from the basic model 220 to the inference engine 124.In various embodiments, the event classifier 420 can also generate a prompt instructing the basic model 220 to determine a failure mode associated with the event entry, the failure mode providing more specificity than a general category of intervention types. For example, the basic model 220 can classify an event entry as a navigation failure and further determine a failure mode that specifies a failure or deterioration of a LiDAR or RADAR sensor included in the robot system.

[0073] Procedure 500, in Operation 508, includes evaluating an event entry for inclusion as a new entry in the training dataset 400. Inputs included in the training dataset 400 include test and validation data used to train an autonomy stack, which includes one or more autonomous or semi-autonomous control routines used to control a fleet of robotic systems. The training data generator 430 calculates similarities between the event entry received from the event database 240 and one or more inputs included in the training dataset 400. Based on the calculated similarities, the training data generator 430 determines whether the received event entry is unusual or atypical compared to the one or more entries included in the training dataset 400.Based on this determination, the training data generator 430 can instruct the foundational model 220 to create and format a new entry for the training dataset 400 based on the received event entry. The inference engine 124 can record the new entry in the training dataset 400. If the training data generator 430 determines that the received event entry is duplicate or redundant given one or more existing inputs included in the training dataset 400, the training data generator 430 can stop analyzing the received event entry.

[0074] Procedure 500 includes, in operation 510, initiating automated retraining of an autonomy stack based on a new entry in the training dataset 400. The inference engine 124 can initiate automated retraining of one or more autonomous or semi-autonomous control routines contained in an autonomy stack via the autonomy stack trainer 440, based on one or more new inputs generated by the training data generator 430 and added to the training dataset 400. In some embodiments, the inference engine 124 can initiate automated retraining of the autonomy stack immediately after generating a new entry in the training dataset 400. In other embodiments, the inference engine 124 can initiate automated retraining upon generating a predetermined number of new inputs in the training dataset 400 or according to a predetermined schedule.

[0075] Procedure 500 includes, in operation 510, the evaluation of an event entry for storage in permanent or other long-term memory, such as memory 114. The event selection module 450 can compare the event entry with existing entries included in memory 114, based on the sensor data 210 included in the event input generated by the natural language generator 410 and associated with the event entry, and / or a category into which the event entry has been classified by the event classifier 420. Based on this comparison, the event selection module 450 can select the event entry for storage in memory 114 if the event entry and its associated natural language description and / or classification are unique, unusual, or otherwise underrepresented in memory 114.Alternatively, the event selection module 450 may not select an event entry for permanent or other long-term storage if the event would be duplicated or redundant in view of the existing event entries included in memory 114.

[0076] The systems and methods described herein can be used for a variety of purposes, including, but not limited to, machine control (e.g., robots, vehicles, construction equipment, warehouse vehicles / machines, automatic, semi-automatic, and / or other machine types), machine locomotion, machine driving, synthetic data generation, model training (e.g., using real, augmented, and / or synthetic data, such as synthetic data generated using a simulation platform or system, synthetic data generation techniques such as, but not limited to, those described herein, etc.), perception, augmented reality (AR), virtual reality (VR), mixed reality (MR), robotics, security, and surveillance (e.g.,in a smart cities implementation), autonomous or semi-autonomous machine applications, deep learning, environmental simulation, object or actuator simulation and / or digital twinning, data center processing, conversational AI, light transport simulation (e.g., ray tracing, path tracing, etc.), distributed collaborative content generation for 3D assets (e.g., using Universal Scene Descriptor (USD) data, such as OpenUSD, and / or other data types), cloud computing, generative artificial intelligence (e.g., using one or more diffusion models, transformer models, etc.), and / or other suitable applications.

[0077] Disclosed embodiments can consist of a variety of different systems, such as automotive systems (e.g., a control system for an autonomous or semi-autonomous machine, a perception system for an autonomous or semi-autonomous machine), systems implemented using a robot or robotics platform, aviation systems, media systems, boating systems, intelligent area monitoring systems, systems for performing deep learning operations, systems for performing simulation operations (e.g., in a driving or vehicle simulation, in a robotics simulation, in a smart city or surveillance simulation, etc.), systems for performing digital twinning operations (e.g.,in conjunction with a collaborative content creation platform or system, such as, but not limited to, NVIDIA's OMNIVERSE and / or any other platform, system or service that uses USD or OpenUSD data types, systems implemented using an edge device, systems containing one or more virtual machines (VMs), systems for performing operations to generate synthetic data (e.g., using one or more neural rendering fields (NERFs), Gaussian splat techniques, diffusion models, transformer models, etc.).), systems that are at least partially implemented in a data center, systems for performing conversational AI operations, systems that implement one or more language models, such as one or more large language models (LLMs), one or more vision language models (VLMs), one or more multimodal language models, etc., systems for performing light transport simulation, systems for performing collaborative content creation for 3D assets (e.g., using Universal Scene Descriptor (USD) data, such as OpenUSD, computer-aided design (CAD) data, 2D and / or 3D graphics or design data, and / or other data types), systems that are at least partially implemented using cloud computing resources, and / or other types of systems.

[0078] In some embodiments, the systems and methods described herein can be performed within a simulation environment (e.g., NVIDIA's DriveSIM, NVIDIA's ISAAC GYM, NVIDIA's ISAAC SIM, etc.) using simulated data (e.g., simulated sensor data from simulated sensors of a virtual or simulated machine). Simulated sensor data can be used, for example, (e.g., using one or more machine learning models, neural networks, etc.) to identify, capture, and / or classify lane lines, obstacles, navigation paths, road boundary lines, other lines, vertical structures / features, etc., within the simulation environment using points on a curve and / or one or more curve-fitting algorithms. This information can then be used to perform operations (e.g., control, navigation, obstacle avoidance, planning operations, etc.).The simulation is used to perform operations assigned to the virtual machine within the environment. These simulated operations can be used to test the performance of underlying algorithms, systems, and / or processes before their deployment in the real world. In some cases, the simulation can be used to generate synthetic training data, such as training data that includes regions of interest and / or subregions of interest, from within the simulation. In some embodiments, other methods can be used additionally or alternatively to simulation for generating synthetic training data. The synthetic training data can be generated, for example, using neural rendering fields (NeRFs), Gaussian splat techniques, diffusion models, electrostatic models (e.g., Poisson flow generative models (PFGMs)), and so on.The synthetic training data (in addition to or as an alternative to real-world data) can then be processed to determine geometry, curvature, semantic information, classification information, and / or other information regarding features of interest, such as lines, obstacles, paths, longitudinal features (e.g., posts), and / or other features within a driving environment, warehouse, etc. In each example, such as when a simulation environment is used for testing, validation, training, etc., the simulation environment and / or associated training data can be rendered or otherwise generated using one or more light transport algorithms, such as ray tracing and / or path tracing algorithms.In some embodiments, the simulation environment and / or one or more objects, features, or components thereof can be created or managed within a three-dimensional content collaboration platform (3D content collaboration platform) (e.g., NVIDIA's OMNIVERSE) for industrial digitization, generative physical AI, and / or other use cases, applications, or services. The content collaboration platform or system may include a system that incorporates universal scene descriptor data (USD data) (e.g., OpenUSD data) for managing objects, features, scenes, etc., within a simulated environment, digital environment, etc. The platform may include real-world physics simulation, such as using NVIDIA's PhysX SDK, to simulate real-world physics and physical interactions with simulations hosted by the platform.The platform can integrate OpenUSD together with ray tracing / path tracing / light transport simulation (e.g. NVIDIA's RTX rendering technologies) into software tools and simulation workflows for building, training, deploying or testing AI systems, such as systems for testing, validating, training (e.g. machine learning model(s), neural networks, etc.) and / or other tasks related to automotive, robotics, machinery or other applications.

[0079] In some embodiments, remote operation or control of a vehicle or other machine can be performed using a remote control or operation system. The systems and methods described herein can be used to identify lane lines, road boundary lines, obstacles, paths, longitudinal features, etc., included in a visualization or representation of an environment to assist a remote operator in controlling or providing waypoints or other information for controlling or navigating an autonomous or semi-autonomous machine through an environment.

[0080] In some embodiments, the systems and methods described herein can be used in a robotics application. For example, a robot or robotics system may include one or more onboard processors (CPUs, GPUs, hardware-based deep learning accelerators (DLAs), hardware-based programmable vision accelerators (PVAs), which may include one or more vector processing units (VPUs), direct memory access systems (DMA systems) and / or pixel processing engines (PPEs), hardware-based optical flow accelerators (OFAs), SoCs, etc.), and working memory and / or storage (e.g., for storing control algorithms, sensor data, and one or more machine learning models). The robotics systems may use these processors to execute one or more machine learning models (e.g.,The system uses language models that allow it to perform complex tasks autonomously or semi-autonomously, such as interacting with and / or manipulating static and / or dynamic objects, or navigating environments using sensors like cameras, LiDAR, radar, ultrasonic sensors, and more. The system can combine sensor fusion techniques to integrate data from multiple sensors (e.g., cameras, infrared, LiDAR, radar, accelerometers) to create a comprehensive model of the robot's environment. This data can be processed locally on the robot or sent to remote servers for more computationally intensive tasks, such as 3D mapping or SLAM (simultaneous localization and mapping). In one or more embodiments, data from individual robots (e.g.,Sensor data, task status, or environmental conditions are uploaded to the cloud, where central AI models analyze optimized commands and distribute them to the entire fleet. In some embodiments, the machine learning model(s) (e.g., language models, VLMs, LLMs, MMLMs, diffusion models, NeRF models, DNNs, etc.) described herein can be used to allow the robot to perceive and analyze its environment and / or communicate with one or more other robots and / or people in that environment. In some embodiments, the robot can communicate with one or more locally hosted servers / computing devices and / or with one or more remote servers / computing devices (e.g., in one or more data centers), for example, using one or more network interface cards (NICs) and / or data processing units (DPUs).

[0081] In some examples, the machine learning model(s) (e.g., deep neural networks, language models, LLMs, VLMs, multimodal language modules, perceptual models, tracking models, fusion models, transformer models, diffusion models, encoder-only models, decoder-only models, encoder-decoder models, neural rendering field models (NeRF models), etc.) described herein may be packaged as a microservice, such as an inference microservice (e.g., NVIDIA NIMs), which may include a container (e.g., an operating system (OS)-level virtualization package) that may contain an application programming interface (API) layer, a server layer, a runtime layer, and / or model “engines”. The inference microservice can, for example, include the container itself and the model(s) (e.g., weights and biases). In some cases, such as when the machine learning model(s) is small enough (e.g.,With a sufficiently small number of parameters, the model(s) can be enclosed within the container itself. In other examples, such as when the model(s) is large, it can be hosted / stored in the cloud (e.g., in a data center) and / or hosted on-premises and / or at the edge (e.g., on a local server or local computing device but outside the container). In such embodiments, the model(s) can be accessible via one or more APIs, such as REST APIs. Therefore, in some embodiments, the machine learning model(s) described herein can be used as an inference microservice to accelerate the deployment of one or more models on any cloud, in any data center, or at an edge computing system, while ensuring data security.The inference microservice can, for example, include one or more APIs, a pre-configured container for simplified deployment, an optimized inference engine (e.g., built using standardized AI model deployment and execution software, such as NVIDIA's Triton Inference Server), and / or one or more APIs for high-performance deep learning inference, which may include inference runtime and model optimizations that deliver low latency and high throughput for production applications (such as NVIDIA's TensorRT) and / or enterprise management data for telemetry (e.g., including identity, metrics, health checks, and / or monitoring).The machine learning model(s) described herein may be included as part of the microservice along with an accelerated infrastructure capable of single-command deployment and / or orchestration and automatic scaling using a container orchestration system on an accelerated infrastructure (e.g., from a single device to a data center-sized environment). Therefore, the inference microservice may include the machine learning model(s) (e.g., optimized for high-performance inference), inference runtime software for executing the machine learning model(s) and providing outputs / responses to inputs (e.g., user queries, prompts, etc.), and enterprise management software for providing health checks, identity verification, and / or other monitoring.In some embodiments, the inference microservice may include software for performing in-place replacement and / or updates of the machine learning model(s). During replacement or update, the software performing the replacement / update may retain user configurations of the inference runtime software and enterprise management software. EXEMPLARY AUTONOMOUS VEHICLE

[0082] Fig. Figure 6A is an illustration of an exemplary autonomous or semi-autonomous vehicle 600 according to some embodiments of the present disclosure. The autonomous vehicle 600 (hereinafter referred to alternatively as "vehicle 600") may, without limitation, include a passenger vehicle such as a passenger car, a truck, a bus, an emergency service vehicle, a shuttle, an electric or motorized bicycle, a motorcycle, a fire engine, a police vehicle, an ambulance, a boat, a construction vehicle, an underwater vehicle, a robotic vehicle, a drone, an aircraft, a vehicle coupled to a trailer (e.g., a semi-trailer truck used for transporting cargo), and / or another type of vehicle (e.g., one that is unmanned and / or carries one or more occupants).Autonomous vehicles are generally described in terms of automation levels defined by the National Highway Traffic Safety Administration (NHTSA), a division of the U.S. Department of Transportation, and the Society of Automotive Engineers (SAE) standard "Taxonomy and Definitions for Terms Related to Driving Automation Systems for On-Road Motor Vehicles" (Standard No. J3016-201806, published June 15, 2018, Standard No. J3016-201609, published September 30, 2016, and earlier and future versions of this standard). The Vehicle 600 may exhibit functionality corresponding to one or more of the Levels 3 through 5 of autonomous driving levels.The Vehicle 600 can exhibit functionality corresponding to one or more of the Levels 1 to 5 of autonomous driving. For example, depending on its embodiment, the Vehicle 600 may be capable of driver assistance (Level 1), partial automation (Level 2), conditional automation (Level 3), high automation (Level 4), and / or full automation (Level 5). The term "autonomous," as used herein, may include any and / or all types of autonomy for the Vehicle 600 or any other machine, such as fully autonomous, highly autonomous, conditionally autonomous, partially autonomous, assistive autonomy, semi-autonomous, primary autonomous, or any other designation.

[0083] The vehicle 600 can include components such as a chassis, a vehicle body, wheels (e.g., 2, 4, 6, 8, 18, etc.), tires, axles, and other vehicle components. The vehicle 600 can include a propulsion system 650, such as an internal combustion engine, a hybrid electric power plant, a pure electric motor, and / or another type of propulsion. The propulsion system 650 can be connected to a drivetrain of the vehicle 600, which may include a transmission to enable the propulsion of the vehicle 600. The propulsion system 650 can be controlled in response to signals received from the throttle valve or accelerator device 652.

[0084] A steering system 654, which may include a steering wheel, can be used to steer the vehicle 600 (e.g., along a desired path or route) when the drive system 650 is in operation (e.g., when the vehicle is in motion). The steering system 654 can receive signals from a steering actuator 656. The steering wheel can be optional for full automation (level 5).

[0085] The brake sensor system 646 can be used to actuate the vehicle brakes in response to receiving signals from the brake actuators 648 and / or the brake sensors.

[0086] The one or more controllers 636, the one or more systems-on-chips (SoCs) 604 ( Fig. 6C) and / or GPUs, can provide signals (e.g., representing instructions) to one or more components and / or systems of the vehicle 600. For example, the one or more controllers 636 can send signals to actuate the vehicle brakes via one or more brake actuators 648, to actuate the steering system 654 via one or more steering actuators 656, and to actuate the propulsion system 650 via one or more throttle / accelerator devices 652. The one or more controllers 636 can include one or more built-in (e.g., integrated) computing devices (e.g., supercomputers) that process sensor signals and issue operating commands (e.g., signals representing commands) to enable autonomous driving and / or to assist a human driver in driving the vehicle 600.The one or more controllers 636 can include a first controller 636 for autonomous driving functions, a second controller 636 for functional safety functions, a third controller 636 for artificial intelligence functions (e.g., computer vision), a fourth controller 636 for infotainment functions, a fifth controller 636 for emergency redundancy, and / or other controllers. In some examples, a single controller 636 can perform two or more of the above functionalities, two or more controllers 636 can perform a single functionality, and / or any combination thereof.

[0087] The one or more controllers 636 can provide the signals to control one or more components and / or systems of the vehicle 600 in response to sensor data received from one or more sensors (e.g. sensor inputs). The sensor data can be received, for example, without restriction, by one or more of the following: Global Navigation Satellite Systems (GNSS) sensor(s) 658 (e.g., Global Positioning System sensor(s)), RADAR sensor(s) 660, Ultrasonic sensor(s) 662, LiDAR sensor(s) 664, Inertial Measurement Unit (IMU) sensor(s) 666 (e.g., accelerometer(s), gyroscope(s), magnetic compass(s), magnetometer(s), etc.), microphone(s) 696, stereo camera(s) 668, wide-angle camera(s) 670 (e.g., fisheye cameras), infrared camera(s) 672, ambient camera(s) 674 (e.g.,360-degree cameras), long-range and / or medium-range camera(s) 698, speed sensor(s) 644 (e.g. for measuring the speed of the vehicle 600), vibration sensor(s) 642, steering sensor(s) 640, brake sensor(s) (e.g. as part of the brake sensor system 646), and / or other sensor types.

[0088] One or more of the controllers 636 can receive inputs (e.g., in the form of input data) from an instrument cluster 632 of the vehicle 600 and provide outputs (e.g., in the form of output data, display data, etc.) via a human-machine interface (HMI) display 634, an audible alarm, a loudspeaker, and / or via other components of the vehicle 600. The outputs can include information such as vehicle speed, engine speed, time, map data (e.g., the high-definition map (HD map) 622 of the vehicle). Fig. 6C), location data (e.g., the location of vehicle 600, e.g., on a map), direction, location of other vehicles (e.g., an occupancy grid), information about objects and the status of objects as perceived by the one or more controllers 636, etc. For example, the HMI display 634 can show information about the presence of one or more objects (e.g., a road sign, a warning sign, a changing traffic light, etc.) and / or information about driving maneuvers that the vehicle has performed, is currently performing, or will perform (e.g., changing lanes now, taking exit 34B in two miles, etc.).

[0089] The vehicle 600 further includes a network interface 624, which can use one or more wireless antennas 626 and / or modems for communication over one or more networks. The network interface 624 can be suitable, for example, for communication over Long-Term Evolution (LTE), Wideband Code Division Multiple Access (WCDMA), Universal Mobile Telecommunications System (UMTS), Global System for Mobile Communications (GSM), IMT-CDMA Multi-Carrier (CDMA2000), etc. The one or more wireless antennas 626 can also enable communication between objects in the environment (e.g., vehicles, mobile devices, etc.) using local area networks such as Bluetooth, Bluetooth Low Energy (LE), Z-Wave, ZigBee, etc.and / or low power wide area networks (LPWANs), such as LoRaWAN, SigFox, etc.

[0090] Fig. 6B is an example of camera locations and fields of view for the exemplary autonomous vehicle 600. Fig. 6A according to some embodiments of the present disclosure; The cameras and respective fields of view are an exemplary embodiment and are not intended to be limiting. For example, additional and / or alternative cameras may be included and / or the cameras may be located at different locations on the vehicle 600.

[0091] The camera types may include, but are not limited to, digital cameras designed for use with the components and / or systems of the vehicle 600. The one or more cameras may operate at Automotive Safety Integrity Level (ASIL) B and / or another ASIL. Depending on the configuration, the camera types may be capable of any frame rate, such as 60 frames per second (fps), 120 fps, 240 fps, etc. The cameras may use roller shutters, global shutters, another type of shutter, or a combination thereof.In some examples, the color filter array may include a red-clear-clear-clear (RCCC) color filter array, a red-clear-clear-blue (RCCB) color filter array, a red-blue-green (RBGC) color filter array, a Foveon X3 color filter array, a Bayer sensor color filter array (RGGB), a monochrome sensor color filter array, and / or another type of color filter array. In some embodiments, cameras with clear pixels, such as cameras with an RCCC, RCCB, and / or RBGC color filter array, may be used to increase light sensitivity.

[0092] In some examples, one or more of the cameras can be used to perform advanced driver assistance system (ADAS) functions (e.g., as part of a redundant or fail-safe design). For instance, a multi-function monocular camera can be installed to provide features including lane departure warning, traffic sign recognition, and intelligent headlight control. One or more of the cameras (e.g., all cameras) can simultaneously record and provide image data (e.g., video).

[0093] One or more cameras can be mounted in a bracket, such as a specially designed (three-dimensional ("3D") printed bracket, to eliminate stray light and reflections from inside the vehicle (e.g., reflections of the dashboard in the windshield) that could interfere with the camera's image acquisition. With regard to exterior mirror brackets, the exterior mirrors can be individually 3D printed so that the camera mounting plate is shaped to fit the mirror. In some cases, the one or more cameras can be integrated into the exterior mirror itself. For side cameras, the one or more cameras can also be integrated into the four pillars at each corner of the cabin.

[0094] Cameras with a field of view that includes portions of the environment in front of the vehicle (e.g., forward-facing cameras) can be used for surround view to help identify forward paths and obstacles and to provide, with the aid of one or more controllers and / or control SoCs, information critical for generating an occupancy grid and / or determining preferred vehicle paths. Forward-facing cameras can be used to perform many of the same ADAS functions as LiDAR, including emergency braking, pedestrian detection, and collision avoidance. Forward-facing cameras can also be used for ADAS functions and systems that include lane departure warnings (LDW), autonomous cruise control (ACC), and / or other functions, such as traffic sign recognition.

[0095] A variety of cameras can be used in a forward-facing configuration, including, for example, a monocular camera platform incorporating a complementary metal oxide semiconductor (CMOS) color imager. Another example is the 670 wide-angle camera, which can be used to detect objects moving into the field of view from the periphery (e.g., pedestrians, crossing traffic, or bicycles). Although in Fig. While Figure 6B illustrates only one wide-angle camera, the vehicle 600 can contain any number (including zero) of wide-angle cameras 670. Furthermore, any number of long-range cameras 698 (e.g., a pair of long-range stereo cameras) can be used for depth-based object detection, particularly for objects for which a neural network has not yet been trained. The one or more long-range cameras 698 can also be used for object detection and classification, as well as basic object tracking.

[0096] Any number of stereo cameras 668 can also be included in a forward-facing configuration. In at least one embodiment, one or more of the stereo cameras 668 can include an integrated control unit comprising a scalable processing unit that can provide programmable logic (“FPGA”) and a multicore microprocessor with an integrated controller area network (“CAN”) or Ethernet interface on a single chip. Such a unit can be used to create a 3D map of the vehicle's surroundings, including a distance estimate for all points in the image.One or more alternative stereo cameras 668 can include a compact stereo vision sensor, which may incorporate two camera lenses (one on the left and one on the right) and an image processing chip that can measure the distance between the vehicle and the target object and use the generated information (e.g., metadata) to activate the autonomous emergency braking and lane departure warning functions. Other types of stereo cameras 668 can be used in addition to or as alternatives to those described herein.

[0097] Cameras with a field of view that includes sections of the environment to the sides of the vehicle 600 (e.g., side cameras) can be used for the surround view and provide information that is used to create and update the occupancy grid and to generate side-impact collision warnings. For example, one or more surround cameras 674 (e.g., four surround cameras 674, as in Fig. (Figure 6B illustrates) are positioned on the vehicle 600. The one or more surround-view cameras 674 can include one or more wide-angle cameras 670, one or more fisheye cameras, one or more 360-degree cameras, and / or the like. For example, four fisheye cameras can be mounted at the front, rear, and sides of the vehicle. In an alternative arrangement, the vehicle can use three surround-view cameras 674 (e.g., left, right, and rear) and one or more other cameras (e.g., a forward-facing camera) as a fourth surround-view camera.

[0098] Cameras with a field of view that includes sections of the area behind the vehicle 600 (e.g., reversing cameras) can be used for parking assistance, surround view, rear-impact warnings, and creating and updating the occupancy grid. A variety of cameras can be used, including, but not limited to, cameras that are also suitable as one or more forward-facing cameras (e.g., one or more long-range and / or medium-range cameras 698, one or more stereo cameras 668, one or more infrared cameras 672, etc.), as described herein.

[0099] Fig. 6C is a block diagram of an exemplary system architecture for the exemplary autonomous vehicle 600. Fig. 6A according to some embodiments of the present disclosure. It is understood that these and other arrangements described herein are presented only as examples. Other arrangements and elements (e.g., machines, interfaces, functions, sequences, groupings of functions, etc.) may be used in addition to or instead of those shown, and some elements may be omitted entirely. Furthermore, many of the elements described herein are functional units that may be implemented as separate or distributed components, or in conjunction with other components, in any suitable combination and at any suitable location. Various functions described herein as being performed by units may be performed by hardware, firmware, and / or software.For example, various functions can be performed by a processor that executes instructions stored in the main memory.

[0100] Each of the components, features and systems of the 600 vehicle in Fig. 6C is illustrated as being connected via bus 602. Bus 602 may include a Controller Area Network Data Interface (CAN Data Interface) (referred to herein alternatively as "CAN bus"). A CAN can be a network within the vehicle 600, used to support the control of various features and functions of the vehicle 600, such as the operation of brakes, acceleration, braking, steering, windshield wipers, etc. A CAN bus may be configured to have dozens or even hundreds of nodes, each with its own unique identifier (e.g., a CAN ID). The CAN bus can be read to determine the steering wheel angle, vehicle speed, engine speed (rpm), button positions, and / or other vehicle status indicators. The CAN bus may be ASIL B compliant.

[0101] Although the 602 bus is described herein as a CAN bus, this is not intended as a limitation. For example, FlexRay and / or Ethernet can be used in addition to or as an alternative to the CAN bus. Furthermore, while a single line is used to represent the 602 bus, this is not meant as a limitation. For example, there can be any number of 602 buses, which may include one or more CAN buses, one or more FlexRay buses, one or more Ethernet buses, and / or one or more other types of buses using a different protocol. In some examples, two or more 602 buses may be used to perform different functions and / or for redundancy. For example, a first 602 bus may be used for collision avoidance functionality, and a second 602 bus may be used for actuation control.In each example, each bus 602 can communicate with one of the vehicle 600 components, and two or more buses 602 can communicate with the same components. In some examples, each SoC 604, each controller 636, and / or each computer within the vehicle can have access to the same input data (e.g., inputs from vehicle 600 sensors) and be connected to a common bus, such as the CAN bus.

[0102] The vehicle 600 can include one or more controllers 636, as described herein with reference to Fig. 6A are described. The one or more 636 controllers can be used for a variety of functions. The one or more 636 controllers can be coupled with one or more of the various other components and systems of the vehicle 600 and can be used for controlling the vehicle 600, for the artificial intelligence of the vehicle 600, for infotainment for the vehicle 600, and / or the like.

[0103] The vehicle 600 can include one or more systems-on-a-chip (SoC) 604. The SoC 604 can include one or more CPUs 606, one or more GPUs 608, one or more processors 610, one or more caches 612, one or more accelerators 614, one or more data storage devices 616, and / or other components and features not illustrated. The one or more SoCs 604 can be used to control the vehicle 600 in a variety of platforms and systems. For example, the one or more SoCs 604 in a system (e.g., the system of the vehicle 600) can be combined with an HD card 622, which is accessed via a network interface 624 by one or more servers (e.g., the one or more servers 678 of the Fig. 6D) may receive map refreshes and / or updates.

[0104] The one or more CPUs 606 can include a CPU cluster or a CPU complex (hereafter referred to as "CCPLEX"). The one or more CPUs 606 can include multiple cores and / or L2 caches. In some embodiments, the one or more CPUs 606 can, for example, include eight cores in a coherent multiprocessor configuration. In some embodiments, the one or more CPUs 606 can include four dual-core clusters, each cluster having a dedicated L2 cache (e.g., a 2 MB L2 cache). The one or more CPUs 606 (e.g., the CCPLEX) can be configured to support the concurrent operation of clusters, so that any combination of clusters of the one or more CPUs 606 can be active at any given time.

[0105] The one or more CPUs 606 can implement power management functions that include one or more of the following features: individual hardware blocks can be automatically clocked when idle to dynamically save power; each core clock can be controlled when the core is not actively executing instructions due to the execution of WFI / WFE instructions; each core can be independently power-controlled; each core cluster can be independently clock-controlled when all cores are clock-controlled or power-controlled; and / or each core cluster can be independently power-controlled when all cores are power-controlled.The one or more CPUs 606 can also implement an improved power state management algorithm where permissible power states and expected wake-up times are defined, and the hardware / microcode determines the best power state to input for the core, cluster, and CCPLEX. The processing cores can support simplified sequences for inputting the power state to software, offloading the work to the microcode.

[0106] The one or more GPUs 608 can include an integrated GPU (referred to herein alternatively as an "iGPU"). The one or more GPUs 608 can be programmable and can be efficient for parallel workloads. The one or more GPUs 608 can use an extended Tensor instruction set in some examples. The one or more GPUs 608 can include one or more streaming microprocessors, each of which can include an L1 cache (for example, an L1 cache of at least 96 KB), and two or more of the streaming microprocessors can share an L2 cache (for example, an L2 cache of 512 KB). In some embodiments, the one or more GPUs 608 can include at least eight streaming microprocessors. The one or more GPUs 608 can use one or more application programming interfaces (APIs) for computation.Furthermore, the one or more GPUs 608 can use one or more parallel computing platforms and / or programming models (e.g., NVIDIA's CUDA).

[0107] The one or more GPUs 608 can be power-optimized for best performance in automotive and embedded applications. The one or more GPUs 608 can be manufactured, for example, on a FinFET field-effect transistor. However, this is not a limitation, and the one or more GPUs 608 can also be manufactured using other semiconductor manufacturing processes. Each streaming microprocessor can contain an array of mixed-precision processing cores, divided into multiple blocks. For example, and without limitation, 64 PF32 cores and 32 PF64 cores can be divided into four processing blocks.In such an example, each processing block can be allocated 16 FP32 cores, 8 FP64 cores, 16 INT32 cores, two mixed-precision NVIDIA TENSOR COREs for deep learning matrix arithmetic, an L0 instruction cache, a warp scheduler, a dispatch unit, and / or a 64 KB register file. Furthermore, the streaming microprocessors can include independent parallel integer and floating-point data paths to provide efficient execution of workloads with a mix of computations and addressing calculations. The streaming microprocessors can include an independent thread scheduling function to enable fine-grained synchronization and cooperation between parallel threads. The streaming microprocessors can include a combined L1 data cache and a shared memory unit to improve performance while simplifying programming.

[0108] The one or more GPUs 608 can include high-bandwidth memory (HBM) and / or a 16 GB HBM2 memory subsystem to provide a peak memory bandwidth of approximately 900 GB / second in some examples. In some examples, synchronous graphics random access memory (SGRAM), such as double-data-rate type five synchronous graphics random access memory (GDDR5), can be used in addition to or as an alternative to HBM memory.

[0109] The one or more GPUs 608 can include a unified memory technology that incorporates access counters to enable more accurate migration of memory pages to the processor that accesses them most frequently, thereby improving the efficiency of memory areas shared by processors. In some examples, Address Translation Services (ATS) support can be used so that the one or more GPUs 608 can directly access the page tables of the one or more CPUs 606. In such examples, if the Memory Management Unit (MMU) of the one or more GPUs 608 fails, an address translation request can be sent to the one or more CPUs 606.In response, the one or more CPUs 606 can search their page tables for the virtual-physical mapping for the address and send the translation back to the one or more GPUs 608. This unified memory technology thus enables a single, unified virtual address space for the memory of both the one or more CPUs 606 and the one or more GPUs 608, thereby simplifying the programming of the one or more GPUs 608 and the porting of applications to the one or more GPUs 608.

[0110] Additionally, the one or more GPUs 608 can include an access counter that tracks the frequency of accesses by the one or more GPUs 608 to the memory of other processors. The access counter can help ensure that memory pages are moved to the physical memory of the processor that accesses them most frequently.

[0111] The one or more SoCs 604 can include any number of caches 612, including those described herein. For example, the one or more caches 612 can include an L3 cache available to both the one or more CPUs 606 and the one or more GPUs 608 (e.g., one connected to both the one or more CPUs 606 and the one or more GPUs 608). The one or more caches 612 can include a write-back cache capable of tracking line states, for example, by using a cache coherence protocol (e.g., MEI, MESI, MSI, etc.). The L3 cache can be 4 MB or larger, depending on the embodiment, although smaller cache sizes are also possible.

[0112] The one or more SoCs 604 can include one or more Arithmetic Logic Units (ALUs) that can be used to perform processing related to one of the many tasks or operations of the Vehicle 600—such as DNN processing. Additionally, the one or more SoCs 604 can include one or more Floating Point Units (FPUs)—or other mathematical or numerical coprocessors—for performing mathematical operations within the system. For example, the one or more SoCs 604 can include one or more FPUs integrated as execution units into one or more CPUs 606 and / or one or more GPUs 608.

[0113] The one or more SoCs 604 can include one or more Accelerators 614 (e.g., hardware accelerators, software accelerators, or a combination thereof). For example, the one or more SoCs 604 can include a hardware acceleration cluster, which may include optimized hardware accelerators and / or a large amount of on-chip memory. The large on-chip memory (e.g., 4 MB SRAM) can enable the hardware acceleration cluster to accelerate neural networks and other computations. The hardware acceleration cluster can be used in conjunction with the one or more GPUs 608 and offload some of the tasks performed by the one or more GPUs 608 (e.g., to free up more cycles of the one or more GPUs 608 for other tasks). The one or more Accelerators 614 can, for example, be used for specific workloads (e.g.,Perception, convolutional neural networks (CNNs), etc., are used that are stable enough to be suitable for acceleration. The term "CNN" as used herein can include all types of CNNs, including region-based or regional convolutional neural networks (RCNNs) and fast RCNNs (e.g., for object detection).

[0114] The one or more Accelerators 614 (e.g., the Hardware Acceleration Cluster) can include a Deep Learning Accelerator (DLA). The one or more DLAs can include one or more Tensor Processing Units (TPUs) configured to provide an additional ten trillion operations per second for deep learning applications and inference. The TPUs can be accelerators configured and optimized to perform image processing functions (e.g., for CNNs, RCNNs, etc.). The one or more DLAs can also be optimized for a specific set of neural network types and floating-point operations, as well as for inference. The design of the one or more DLAs can deliver more performance per millimeter than a general-purpose GPU and far surpasses the performance of a CPU.The one or more TPUs can perform multiple functions, including a convolution function for a single instance that supports, for example, INT8, INT16 and FP16 data types for both features and weights, as well as post-processor functions.

[0115] One or more DLAs can quickly and efficiently run neural networks, especially CNNs, on processed or unprocessed data for a variety of functions, including, but not limited to: a CNN for object identification and detection using camera sensor data; a CNN for distance estimation using camera sensor data; a CNN for emergency vehicle detection and identification using microphone data; a CNN for facial recognition and vehicle owner identification using camera sensor data; and / or a CNN for security and / or protection-related events.

[0116] The one or more DLAs can perform any function of the one or more GPUs 608, and by using an inference accelerator, a developer can, for example, allocate either the one or more DLAs or the one or more GPUs 608 to each function. For example, the developer can concentrate the processing of CNNs and floating-point operations on the one or more DLAs and leave other functions to the one or more GPUs 608 and / or other accelerators 614.

[0117] The one or more Accelerators 614 (e.g., the Hardware Acceleration Cluster) can include one or more Programmable Vision Accelerators (PVAs), which may also be referred to herein as Vision Accelerators. The one or more PVAs can be designed and configured to accelerate computer vision algorithms for Advanced Driver Assistance Systems (ADAS), autonomous driving, and / or augmented reality (AR) and / or virtual reality (VR) applications. The one or more PVAs can provide a balance between performance and flexibility. For example, each PVA can include, without limitation, any number of Reduced Instruction Set Computer (RISC) cores, Direct Memory Access (DMA) cores, and / or any number of vector processors.

[0118] The RISC cores can interact with image sensors (e.g., the image sensors of one of the cameras described herein), image signal processors, and / or the like. Each RISC core can include any amount of memory. Depending on the implementation, the RISC cores can use any number of protocols. In some examples, the RISC cores can run a real-time operating system (RTOS). The RISC cores can be implemented with one or more integrated circuits, application-specific integrated circuits (ASICs), and / or memory devices. For example, the RISC cores can include an instruction cache and / or tightly coupled RAM.

[0119] The DMA can allow components of the PVA(s) to access the system's memory independently of the single or multiple CPUs. The DMA can support any number of features that optimize the PVA, including, but not limited to, support for multidimensional and / or circular addressing. In some examples, the DMA can support up to six or more dimensions of addressing, which may include block width, block height, block depth, horizontal block stepping, vertical block stepping, and / or depth stepping.

[0120] Vector processors can be programmable processors designed to efficiently and flexibly execute programming for computer vision algorithms and provide signal processing capabilities. In some examples, the PVA may include a PVA core and two vector processing subsystem partitions. The PVA core may include a processor subsystem, one or more DMA machines (e.g., two DMA machines), and / or other peripheral devices. The vector processing subsystem may act as the primary processing engine of the PVA and may include a vector processing unit (VPU), an instruction cache, and / or main memory (e.g., VMEM).A VPU core can include a digital signal processor, such as a single instruction, multiple data (SIMD) or a very long instruction word (VLIW). The combination of SIMD and VLIW can increase throughput and speed.

[0121] Each vector processor can include an instruction cache and can be coupled to dedicated memory. Therefore, in some examples, each vector processor can be configured to operate independently of the others. In other examples, the vector processors contained in a particular PVA can be configured to use data parallelism. For example, in some embodiments, the multitude of vector processors contained in a single PVA can execute the same computer vision algorithm, but on different regions of an image. In other examples, the vector processors contained in a particular PVA can simultaneously execute different computer vision algorithms on the same image, or even different algorithms on successive images or sections of an image.Among other things, the hardware acceleration cluster can include any number of PVAs and any number of vector processors in each of the PVAs. Furthermore, one or more PVAs can include additional memory for error-correcting code (ECC) to increase overall system security.

[0122] The one or more Accelerator 614 units (e.g., the hardware acceleration cluster) can include an on-chip computer vision network and SRAM to provide high-bandwidth, low-latency SRAM for the Accelerator 614. In some examples, the on-chip memory can include at least 4 MB of SRAM, consisting, for example, and without limitation, of eight field-configurable memory blocks accessible to both the PVA and the DLA. Each pair of memory blocks can include an Advanced Peripheral Bus (APB) interface, configuration circuitry, a controller, and a multiplexer. Any type of memory can be used. The PVA and the DLA can access the memory via a backbone, enabling high-speed memory access for both the PVA and the DLA.The backbone can include an on-chip computer vision network that connects the PVA and DLA to the main memory (e.g., using the APB).

[0123] The on-chip computer vision network can include an interface that, prior to the transmission of control signals / addresses / data, ensures that both the PVA and the DLA provide ready-to-use and valid signals. Such an interface can provide separate phases and channels for the transmission of control signals / addresses / data, as well as burst communication for continuous data transmission. This type of interface can conform to ISO 26262 or IEC 61508 standards, although other standards and protocols can also be used.

[0124] In some examples, one or more 604 SoCs can include a real-time ray-tracing hardware accelerator as described in U.S. Patent Application No. 16 / 101,232, filed on August 10, 2018. The real-time ray-tracing hardware accelerator can be used to quickly and efficiently determine the positions and extents of objects (e.g., within a world model), to generate real-time visualization simulations, for radar signal interpretation, for sound propagation synthesis and / or analysis, for the simulation of SONAR systems, for general wave propagation simulation, for comparison with LiDAR data for localization purposes, and / or for other functions and / or purposes. In some embodiments, one or more Tree Traversal Units (TTUs) can be used to perform one or more operations related to ray tracing.

[0125] The single or multiple Accelerator 614 (e.g., the hardware accelerator cluster) have a wide range of applications for autonomous driving. The PVA can be a programmable vision accelerator used for critical processing steps in ADAS and autonomous vehicles. The PVA's capabilities are well-suited to algorithmic domains requiring predictable processing with low power consumption and low latency. In other words, the PVA is well-suited for semi-dense or dense regular computations, even with small datasets, that require predictable runtimes with low latency and low power consumption. Therefore, in the context of autonomous vehicle platforms, PVAs are designed to execute classic computer vision algorithms, as they are efficient at object detection and operate with integer mathematics.

[0126] According to one embodiment of the technology, the PVA is used, for example, to perform computer stereovision. In some examples, a semi-global matching-based algorithm can be used, although this is not intended as a limitation. Many applications for Level 3-5 autonomous driving require spontaneous motion estimation or stereo matching (e.g., structure of motion, pedestrian detection, lane detection, etc.). The PVA can perform computer stereovision on input from two monocular cameras.

[0127] In some examples, the PVA can be used to perform dense optical flow processing. This involves processing raw radar data (e.g., using a 4D Fast Fourier Transform) to provide processed radar data. In other examples, the PVA is used for time-of-flight depth processing, for instance, by processing raw time-of-flight data to provide processed time-of-flight data.

[0128] The DLA can be used to power any type of network to improve control and driving safety; this includes, for example, a neural network that outputs a confidence score for each object detection. Such a confidence score can be interpreted as a probability or as providing a relative "weighting" of each detection compared to other detections. This confidence score allows the system to make further decisions about which detections should be considered true positives and not false positives. For example, the system can set a confidence threshold and consider only those detections that exceed the threshold as true positives.In an automatic emergency braking (AEB) system, false positive detections would cause the vehicle to automatically initiate emergency braking, which is obviously undesirable. Therefore, only the safest detections should be considered as triggers for AEB. The DLA can employ a neural network for confidence regression. The neural network can take as input at least a subset of parameters, such as the dimensions of the boundary frame, the ground plane estimate (obtained, for example, from another subsystem), the output of the inertial measurement unit (IMU) sensor 666 correlated with the vehicle's orientation 600, distance, and 3D position estimates of the object obtained from the neural network and / or other sensors (e.g., one or more LiDAR sensors 664 or one or more radar sensors 660).

[0129] The one or more SoCs 604 can include the one or more data stores 616 (e.g., main memory). The one or more data stores 616 can be on-chip main memory on the one or more SoCs 604, where neural networks can be stored to run on the GPU and / or the DLA. In some examples, the one or more data stores 616 can be large enough to store multiple instances of neural networks for redundancy and security. The one or more data cache(s) 612 can include one or more L2 or L3 caches 612. The reference to the one or more data stores 616 can include a reference to main memory allocated to the PVA, the DLA, and / or one or more other accelerators 614, as described herein.

[0130] The one or more SoCs 604 can include one or more processors 610 (e.g., embedded processors). The one or more processors 610 can include a boot and power management processor, which can be a dedicated processor and subsystem to handle boot power and management functions and the associated security enforcement. The boot and power management processor can be part of the boot sequence of the one or more SoCs 604 and can provide runtime power management services. The boot and power management processor can provide clock and voltage programming, support for system transitions to a low-power state, management of the thermals and temperature sensors of the one or more SoCs 604, and / or management of the power states of the one or more SoCs 604.Each temperature sensor can be implemented as a ring oscillator whose output frequency is proportional to the temperature, and the one or more SoCs 604 can use the ring oscillators to detect the temperatures of the one or more CPUs 606, the one or more GPUs 608, and / or the one or more accelerators 614. If it is determined that the temperatures exceed a threshold, the boot and power management processor can enter a temperature fault routine and put the one or more SoCs 604 into a reduced-power state and / or put the vehicle 600 into a chauffeur-to-safe-stop mode (e.g., bring the vehicle 600 to a safe stop).

[0131] The one or more 610 processors can also include a number of embedded processors that can serve as an audio processing engine. The audio processing engine can be an audio subsystem that provides full hardware support for multi-channel audio across multiple interfaces and a wide and flexible range of audio I / O interfaces. In some examples, the audio processing engine is a dedicated processor core with a digital signal processor and dedicated RAM.

[0132] The one or more 610 processors can also include an always-on processor engine, which provides the necessary hardware functions to support low-power sensor management and wake-up of use cases. The always-on processor engine can include a processor core, tightly coupled RAM, supporting peripherals (such as timers and interrupt controllers), various I / O controller peripherals, and routing logic.

[0133] The single or multiple 610 processors can also include a security cluster machine, which contains a dedicated processor subsystem for the security management of automotive applications. The security cluster machine can include two or more processor cores, tightly coupled RAM, supporting peripherals (such as timers, an interrupt controller, etc.), and / or routing logic. In a security mode, the two or more cores can operate in lockstep mode, functioning as a single core with comparison logic to detect any differences between their operations.

[0134] The one or more 610 processors can also include a real-time camera engine, which may include a dedicated processor subsystem for managing the real-time camera.

[0135] The one or more 610 processors can further include a high dynamic range signal processor, which can include an image signal processor, which is a hardware machine that is part of the camera processing pipeline.

[0136] The one or more 610 processors can include a video image compositor, which can be a processing block (e.g., implemented on a microprocessor) that implements video post-processing functions required by a video playback application to generate the final image for the player window. The video image compositor can perform lens distortion correction on wide-angle camera(s) 670, ambient camera(s) 674, and / or on the sensors of the in-cabin surveillance camera. The in-cabin surveillance camera sensor is preferably monitored by a neural network running on a separate instance of the extended SoC and configured to detect events in the cabin and respond accordingly.A system in the cabin can lip-read to activate mobile service and make a call, dictate emails, change the destination, activate or change the infotainment system and vehicle settings, or enable voice-controlled internet browsing. Certain functions are only available to the driver when the vehicle is operating in autonomous mode and are otherwise deactivated.

[0137] The video image compositor can include enhanced temporal noise reduction for both spatial and temporal noise reduction. For example, if there is motion in a video, the noise reduction weights the spatial information accordingly and reduces the impact of information provided by adjacent frames. If a frame or portion of a frame does not contain motion, the temporal noise reduction performed by the video image compositor can use information from the previous frame to reduce noise in the current frame.

[0138] The video image compositor can also be configured to perform stereo equalization of the input stereo lens images. Furthermore, the video image compositor can be used for user interface design when the operating system desktop is in use and the one or more GPUs 608 do not need to constantly render new surfaces. Even when the one or more GPUs 608 are powered on and actively performing 3D rendering, the video image compositor can be used to offload the load on the one or more GPUs 608, thus improving performance and responsiveness.

[0139] The one or more 604 SoCs can also include a serial camera interface with a Mobile Industry Processor Interface (MIPI) for receiving video and camera input, a high-speed interface, and / or a video input block that can be used for camera and related pixel input functions. The one or more 604 SoCs can also include one or more software-controlled input / output controllers that can be used to receive I / O signals not assigned to a specific role.

[0140] The one or more 604 SoCs can also include a wide range of peripheral interfaces to enable communication with peripheral devices, audio codecs, power management, and / or other devices. The one or more 604 SoCs can be used to process data from cameras (e.g., via Gigabit Multimedia Serial Link and Ethernet), sensors (e.g., one or more 664 LiDAR sensors, one or more 660 radar sensors, etc., which can be connected via Ethernet), data from the 602 bus (e.g., vehicle speed 600, steering wheel position, etc.), and data from one or more 658 GNSS sensors (e.g., connected via Ethernet or CAN bus).Furthermore, the one or more SoCs 604 can include dedicated high-performance mass storage controllers, which can include their own DMA machines and can be used to offload routine data management tasks from the one or more CPUs 606.

[0141] The single or multiple 604 SoCs can form an end-to-end platform with a flexible architecture spanning automation levels 3-5, thereby providing a comprehensive functional safety architecture that supports and efficiently utilizes computer vision and ADAS techniques for diversity and redundancy, and provides a platform for a flexible, reliable driving software stack along with deep learning tools. The single or multiple 604 SoCs can be faster, more reliable, and even more energy- and space-efficient than conventional systems. For example, the single or multiple 614 accelerators, in combination with the single or multiple 606 CPUs, the single or multiple 608 GPUs, and the single or multiple 616 data stores, can form a fast, efficient platform for level 3-5 autonomous vehicles.

[0142] This technology thus provides capabilities and functions that cannot be achieved by conventional systems. For example, computer vision algorithms can be run on CPUs that can be configured using a high-level programming language, such as C, to execute a wide variety of processing algorithms on a wide range of visual data. However, CPUs are often unable to meet the performance requirements of many computer vision applications, such as execution time and power consumption. In particular, many CPUs are unable to execute complex object detection algorithms in real time, which is a requirement for in-vehicle ADAS applications and a prerequisite for practical Level 3-5 autonomous vehicles.

[0143] Unlike conventional systems, the technology described herein, by providing a CPU complex, a GPU complex, and a hardware acceleration cluster, enables the simultaneous and / or sequential execution of multiple neural networks and the combination of their results to enable Level 3-5 autonomous driving functionality. For example, a CNN running on the DLA or the dGPU (e.g., one or more GPUs 620) can include text and word recognition, which allows the supercomputer to believe it can read and understand traffic signs, including signs for which the neural network has not been specifically trained. The DLA can further include a neural network capable of identifying and interpreting the sign, providing a semantic understanding, and passing this semantic understanding to the path planning modules running on the CPU complex.

[0144] Another example is that multiple neural networks can run simultaneously, as required for driving at levels 3, 4, or 5. For instance, a warning sign reading "Caution: Flashing lights indicate black ice" accompanied by an electric light can be interpreted independently or jointly by several neural networks. The sign itself can be identified as a traffic sign by a first neural network (e.g., a trained one), while the text "Flashing lights indicate black ice" can be interpreted by a second neural network, which then informs the vehicle's path planning software (preferably running on the CPU) that the presence of black ice indicates the presence of flashing lights.The turn signal can be identified across multiple images by a third neural network, which informs the vehicle's path planning software about the presence (or absence) of turn signals. All three neural networks can run simultaneously, for example, within the DLA and / or on one or more GPUs 608.

[0145] In some examples, a CNN for facial recognition and vehicle owner identification can use data from camera sensors to identify the presence of an authorized driver and / or owner of the vehicle 600. The always-on sensor processing engine can be used to unlock the vehicle and turn on the lights when the owner approaches the driver's door, and to disable the vehicle in security mode when the owner leaves. In this way, one or more SoCs 604 provide security against theft and / or carjacking.

[0146] In another example, a CNN for detecting and identifying emergency vehicles can use data from microphones 696 to detect and identify emergency vehicle sirens. Unlike conventional systems that use general classifiers to detect sirens and manually extract features, the one or more SoCs 604 use the CNN to classify ambient and urban noise as well as visual data. In a preferred embodiment, the CNN running on the DLA is trained to detect the relative approach speed of the emergency vehicle (e.g., by using the Doppler effect). The CNN can also be trained to identify emergency vehicles specific to the local area in which the vehicle is operating, as identified by one or more GNSS sensors 658.Therefore, for example, the CNN will attempt to detect European sirens when it is operating in Europe, and when it is operating in the United States, the CNN will attempt to identify only North American sirens. Once an emergency vehicle is detected, a controller can be used to execute an emergency vehicle safety routine, slowing the vehicle down, pulling over to the side of the road, parking the vehicle, and / or letting the vehicle idle, using the 662 ultrasonic sensors, until one or more emergency vehicles pass.

[0147] The vehicle can include one or more CPUs 618 (e.g., one or more discrete CPUs or one or more dCPUs) that can be coupled to the one or more SoCs 604 via a high-speed connection (e.g., PCIe). The CPUs 618 can, for example, include an x86 processor. The CPUs 618 can be used, for example, to perform a variety of functions, including reconciling potentially inconsistent results between ADAS sensors and the one or more SoCs 604 and / or monitoring the status and health of the one or more Controllers 636 and / or the Infotainment SoC 630.

[0148] The Vehicle 600 can include one or more GPUs 620 (e.g., one or more discrete GPUs or one or more dGPUs) that can be coupled to the one or more SoCs 604 via a high-speed connection (e.g., NVIDIA's NVLINK). The one or more GPUs 620 can provide additional artificial intelligence capabilities, such as running redundant and / or distinct neural networks, and can be used to train and / or update neural networks based on input (e.g., sensor data) from sensors in the Vehicle 600.

[0149] The vehicle 600 can further include the network interface 624, which can include one or more wireless antennas 626 (e.g., one or more wireless antennas for different communication protocols, such as a cellular antenna, a Bluetooth antenna, etc.). The network interface 624 can be used to enable a wireless connection via the internet to the cloud (e.g., to the one or more servers 678 and / or other network devices), to other vehicles, and / or to computing devices (e.g., passenger client devices). To communicate with other vehicles, a direct connection between the two vehicles and / or an indirect connection (e.g., via networks and the internet) can be established. Direct connections can be established via vehicle-to-vehicle communication.Vehicle-to-vehicle communication can provide the vehicle 600 with information about vehicles in its vicinity (e.g., vehicles in front of, beside, and / or behind the vehicle 600). This functionality can be part of a cooperative adaptive cruise control function of the vehicle 600.

[0150] The 624 network interface can include a system-on-a-chip (SoC) that provides modulation and demodulation capabilities, enabling one or more 636 controllers to communicate over wireless networks. The 624 network interface can include a high-frequency (RF) front end for upconversion from baseband to RF and downconversion from RF to baseband. Frequency conversions can be performed using known processes and / or superheterodyne processes. In some examples, the RF front-end functionality can be provided by a separate chip. The network interface can include wireless functionality for communication via LTE, WCDMA, UMTS, GSM, CDMA2000, Bluetooth, Bluetooth LE, Wi-Fi, Z-Wave, ZigBee, LoRaWAN, and / or other wireless protocols.

[0151] The vehicle 600 may further include one or more data storage devices 628, which may be located outside the chip (e.g., outside the SoCs 604). The one or more data storage devices 628 may include one or more memory elements, including RAM, SRAM, DRAM, VRAM, flash memory, hard disks, and / or other components and / or devices capable of storing at least one bit of data.

[0152] The Vehicle 600 can also include one or more GNSS Sensors 658. The one or more GNSS Sensors 658 (e.g., GPS, supported GPS sensors, differential GPS (DGPS) sensors, etc.) assist with mapping, perception, occupancy grid creation, and / or path planning. Any number of GNSS Sensors 658 can be used, including, for example, a GPS unit that uses a USB connection with an Ethernet-to-serial bridge (RS-232 bridge).

[0153] The vehicle 600 can further include one or more RADAR sensors 660. The one or more RADAR sensors 660 can be used by the vehicle 600 for long-range vehicle detection, even in darkness and / or adverse weather conditions. The functional safety level of the RADAR can be ASIL B. The one or more RADAR sensors 660 can use the CAN bus and / or the 602 bus (e.g., for transmitting the data generated by the one or more RADAR sensors 660) for control and access to object tracking data, with some examples using Ethernet for access to the raw data. A variety of RADAR sensor types can be used. The one or more RADAR sensors 660 can be suitable for front, rear, and side RADAR applications without restriction. In some examples, one or more pulse-Doppler RADAR sensors are used.

[0154] The RADAR sensor(s) 660 can include different configurations, such as long-range with a narrow field of view, short-range with a wide field of view, side coverage with short-range coverage, etc. In some examples, long-range radar can be used for adaptive cruise control. Long-range radar systems can provide a wide field of view, achieved through two or more independent scanning circuits, as in a 250-meter range. The RADAR sensor(s) 660 help distinguish between stationary and moving objects and can be used by ADAS systems for emergency braking assistance and forward collision warning. Long-range radar sensors can include a monostatic multimodal radar with multiple (e.g., six or more) fixed radar antennas and a high-speed CAN and FlexRay interface.In an example with six antennas, the four central antennas can generate a focused beam pattern designed to detect the area around vehicle 600 at higher speeds with minimal interference from traffic in adjacent lanes. The other two antennas can expand the field of view, enabling the rapid detection of vehicles entering or exiting vehicle 600's lane.

[0155] Medium-range radar systems, for example, can have a range of up to 660 m (front) or 80 m (rear) and a field of view of up to 42 degrees (front) or 650 degrees (rear). Short-range radar systems can include, among other things, radar sensors designed for installation at both ends of the rear bumper. When such a radar sensor system is installed at both ends of the rear bumper, it can create two beams that continuously monitor the blind spot behind and to the sides of the vehicle.

[0156] Short-range radar systems can be used in an ADAS system for blind spot detection and / or as a lane change assistant.

[0157] The vehicle 600 can also include one or more ultrasonic sensors 662. The one or more ultrasonic sensors 662, which can be mounted on the front, rear, and / or sides of the vehicle 600, can be used for parking assistance and / or for creating and updating an occupancy grid. A variety of ultrasonic sensors 662 can be used, and different ultrasonic sensors 662 can be used for different detection ranges (e.g., 2.5 m, 4 m). The one or more ultrasonic sensors 662 can operate with functional safety levels of ASIL B.

[0158] The vehicle 600 can include one or more LiDAR sensors 664. The one or more LiDAR sensors 664 can be used for object and pedestrian detection, emergency braking, collision avoidance, and / or other functions. The one or more LiDAR sensors 664 can meet the functional safety level ASIL B. In some examples, the vehicle 600 can include multiple LiDAR sensors 664 (e.g., two, four, six, etc.) that can use Ethernet (e.g., to deliver data to a Gigabit Ethernet switch).

[0159] In some examples, one or more LiDAR sensors 664 may be able to provide a list of objects and their distances for a 360-degree field of view. Commercially available LiDAR sensors 664 may, for example, have a specified range of approximately 600 m, with an accuracy of 2-3 cm and support for a 600 Mbit / s Ethernet connection. In some examples, one or more non-protruding LiDAR sensors 664 may be used. In such examples, the one or more LiDAR sensors 664 may be implemented as a small device that can be embedded in the front, rear, sides, and / or corners of the vehicle 600. In such examples, one or more LiDAR 664 sensors can provide a horizontal field of view of up to 120 degrees and a vertical field of view of up to 35 degrees, with a range of 200 m, even with objects of low reflectivity.The one or more front-mounted LiDAR 664 sensors can be configured for a horizontal field of view between 45 degrees and 135 degrees.

[0160] In some examples, LiDAR technologies, such as 3D flash LiDAR, can also be used. 3D flash LiDAR uses a laser pulse as a transmission source to illuminate the vehicle's surroundings up to approximately 200 m. A flash LiDAR unit includes a sensor that records the travel time of the laser pulse and the reflected light at each pixel, corresponding to the distance between the vehicle and objects. Flash LiDAR can generate highly accurate and distortion-free images of the environment with each laser pulse. In some examples, four flash LiDAR sensors can be used, one on each side of the vehicle. Available 3D flash LiDAR systems include a solid-state 3D focal plane array LiDAR camera that contains no moving parts other than a fan (e.g., a non-scanning LiDAR device).The flash LiDAR device can use a 5-nanosecond pulse of a Class I (eye-safe) laser per frame and capture the reflected laser light in the form of 3D distance point clouds and co-registered intensity data. Because flash LiDAR is a solid-state device with no moving parts, the single or multiple LiDAR sensors may be less susceptible to motion blur, vibration, and / or shock.

[0161] The vehicle may also include one or more IMU sensors 666. In some examples, the one or more IMU sensors 666 may be located in the center of the rear axle of the vehicle 600. The one or more IMU sensors 666 may, for example, and without limitation, include one or more accelerometers, one or more magnetometers, one or more gyroscopes, one or more magnetic compasses, and / or other sensor types. In some examples, such as six-axis applications, the one or more IMU sensors 666 may include accelerometers and gyroscopes, while in nine-axis applications, the one or more IMU sensors 666 may include accelerometers, gyroscopes, and magnetometers.

[0162] In some embodiments, the one or more IMU sensors 666 can be implemented as a miniaturized, high-performance GPS-aided inertial navigation system (GPS / INS) that combines inertial sensors of a microelectromechanical system (MEMS), a highly sensitive GPS receiver, and advanced Kalman filter algorithms to provide estimates of position, velocity, and orientation. Therefore, in some examples, the one or more IMU sensors 666 can enable the vehicle 600 to estimate its course without requiring input from a magnetic sensor by directly observing and correlating velocity changes from the GPS with the one or more IMU sensors 666.In some examples, one or more IMU sensors 666 and one or more GNSS sensors 658 can be combined in a single integrated unit.

[0163] The vehicle can include one or more microphones 696, which are placed in and / or around the vehicle 600. The one or more microphones 696 can be used, among other things, for the detection and identification of emergency vehicles.

[0164] The vehicle can further include any number of camera types, including one or more stereo cameras 668, one or more wide-angle cameras 670, one or more infrared cameras 672, one or more surround-view cameras 674, one or more long-range and / or medium-range cameras 698, and / or other camera types. The cameras can be used to capture image data around the entire periphery of the vehicle 600. The types of cameras used depend on the embodiment and requirements of the vehicle 600, and any combination of camera types can be used to provide the necessary coverage around the vehicle 600. Furthermore, the number of cameras can vary depending on the embodiment. For example, the vehicle can include six cameras, seven cameras, ten cameras, twelve cameras, and / or any other number of cameras.The cameras can, as an example and without limitation, support Gigabit Multimedia Serial Link (GMSL) and / or Gigabit Ethernet. Each of the cameras is described herein with reference to... Fig. 6A and Fig. 6B is described in more detail.

[0165] The vehicle 600 can further include one or more vibration sensors 642. The one or more vibration sensors 642 can measure vibrations of vehicle components, such as one or more axles. For example, changes in vibrations can indicate a change in the road surface. In another example, if two or more vibration sensors 642 are used, the differences between the vibrations can be used to determine friction or slippage on the road surface (e.g., if the difference in vibration is between a driven axle and a freely rotating axle).

[0166] The vehicle 600 may include an ADAS system 638. The ADAS system 638 may include a SoC in some examples. The ADAS system 638 may include autonomous / adaptive / automatic cruise control (ACC), cooperative adaptive cruise control (CACC), forward crash warning (FCW), automatic emergency braking (AEB), lane departure warning (LDW), lane keep assist (LKA), blind spot warning (BSW), rear cross-traffic warning (RCTW), collision warning systems (CWS), lane centering (LC), and / or other features and functions.

[0167] The ACC systems can use one or more radar sensors 660, one or more LiDAR sensors 664, and / or one or more cameras. The ACC systems can include longitudinal ACC and / or lateral ACC. Longitudinal ACC monitors and controls the distance to the vehicle immediately in front of the vehicle 600 and automatically adjusts the vehicle speed to maintain a safe distance from vehicles ahead. Lateral ACC performs distance control and advises the vehicle 600 to change lanes if necessary. Lateral ACC interacts with other ADAS applications, such as LCA and CWS.

[0168] The CACC uses information from other vehicles, which can be received via the network interface 624 and / or the one or more wireless antennas 626 from other vehicles via a wireless connection or indirectly via a network connection (e.g., via the internet). Direct connections can be provided via a vehicle-to-vehicle (V2V) communication link, while indirect connections can be an infrastructure-to-vehicle (I2V) communication link. In general, the V2V communication concept provides information about the vehicles immediately ahead (e.g., vehicles directly in front of the vehicle 600 and in the same lane), while the I2V communication concept provides information about traffic further ahead.CACC systems can incorporate both I2V and V2V information sources. Given the information about the vehicles ahead, CACC can be more reliable and has the potential to improve traffic flow and reduce congestion.

[0169] FCW systems are designed to warn the driver of a hazard so they can take corrective action. FCW systems use a forward-facing camera and / or one or more RADAR 660 sensors coupled with a dedicated processor, DSP, FPGA, and / or ASIC, which is electrically connected to the driver feedback system, such as a display, speaker, and / or vibrating component. FCW systems can provide a warning, for example, in the form of an audible signal, a visual warning, a vibration, and / or a rapid braking pulse.

[0170] AEB systems detect an impending forward collision with another vehicle or object and can automatically apply the brakes if the driver does not take corrective action within a specified time or distance parameter. AEB systems can use one or more forward-facing cameras and / or one or more RADAR 660 sensors coupled with their own processor, DSP, FPGA, and / or ASIC. When the AEB system detects a hazard, it typically first warns the driver so they can take corrective action to avoid the collision; if the driver does not take corrective action, the AEB system can automatically apply the brakes to prevent or at least mitigate the impact of the predicted collision. AEB systems may include techniques such as dynamic brake assist and / or emergency braking for an impending collision.

[0171] Lane Departure Warning (LDW) systems provide visual, audible, and / or tactile warnings, such as steering wheel or seat vibrations, to alert the driver if the vehicle crosses lane markings. An LDW system will not activate if the driver indicates an intentional lane departure by using a turn signal. LDW systems may utilize forward-facing cameras coupled to a dedicated processor, DSP, FPGA, and / or ASIC, which is electrically connected to the feedback device for the driver, such as a display, speaker, and / or vibrating component.

[0172] LKA systems are a variant of LDW systems. LKA systems provide steering or braking inputs to correct the vehicle 600 if the vehicle 600 begins to leave its lane.

[0173] Blind Spot Warning (BSW) systems detect and warn the driver of vehicles in the car's blind spot. BSW systems can provide visual, audible, and / or tactile warnings to indicate that merging into or changing lanes is unsafe. The system can provide an additional warning when the driver activates a turn signal. BSW systems can use one or more rear-facing camera(s) and / or radar sensors coupled with a dedicated processor, DSP, FPGA, and / or ASIC, which is electrically connected to the driver for feedback, such as a display, speaker, and / or vibrating component.

[0174] RCTW systems can provide visual, audible, and / or tactile alerts when an object is detected outside the reversing camera's field of view while the vehicle is reversing. Some RCTW systems include AEB to ensure the vehicle's brakes are applied to prevent a collision. RCTW systems can utilize one or more rear-facing radar sensors coupled with a dedicated processor, DSP, FPGA, and / or ASIC, which is electrically connected to the driver for feedback, such as a display, speaker, and / or vibrating component.

[0175] Conventional ADAS systems can produce false positives, which, while annoying and distracting for the driver, are generally not catastrophic because the ADAS systems warn the driver and give them the opportunity to decide whether a safety issue truly exists and to act accordingly. However, in an autonomous vehicle 600, the vehicle 600 itself must decide, in the event of conflicting results, whether to follow the result from a primary computer or a secondary computer (e.g., a first controller 636 or a second controller 636). In some embodiments, the ADAS system 638 can, for example, be a backup and / or secondary computer that provides information about perception to a rationality module of the backup computer.The backup computer rationality monitor can run redundant, diverse software on hardware components to detect errors in perception and dynamic driving tasks. The outputs of the ADAS system 638 can be provided to a monitoring MCU. If the outputs of the primary and secondary computers conflict, the monitoring MCU must determine how to resolve the conflict to ensure safe operation.

[0176] In some examples, the primary computer can be configured to provide the supervising MCU with a confidence score indicating its confidence in the chosen outcome. If the confidence score exceeds a threshold, the supervising MCU can follow the primary computer's instruction, regardless of whether the secondary computer provides a conflicting or inconsistent result. If the confidence score does not reach the threshold and the primary and secondary computers report different results (e.g., a conflict), the supervising MCU can mediate between the computers to determine the appropriate outcome.

[0177] The monitoring MCU can be configured to run one or more neural networks that are trained and configured to determine, based on the output of the primary and secondary computers, the conditions under which the secondary computer will trigger false alarms. This allows the one or more neural networks in the monitoring MCU to learn when the output of the secondary computer can be trusted and when it cannot. For example, if the secondary computer is a radar-based FCW system, a neural network in the monitoring MCU can learn to trigger an alarm when the FCW system identifies metallic objects that do not actually pose a threat, such as a drain grate or manhole cover.Similarly, if the secondary computer is a camera-based lane departure warning (LDW) system, a neural network in the supervising MCU can learn to override the LDW system when cyclists or pedestrians are present and leaving the lane is indeed the safest maneuver. In embodiments that include one or more neural networks running on the supervising MCU, the supervising MCU can include at least one DLA or GPU suitable for running the one or more neural networks with associated memory. In preferred embodiments, the supervising MCU can include and / or be contained as a component of the one or more SoCs 604.

[0178] In other examples, the ADAS System 638 can include a secondary computer that executes the ADAS functionality according to the classical rules of computer vision. Thus, the secondary computer can use classical computer vision rules (if-then), and the presence of one or more neural networks in the monitoring MCU can improve reliability, safety, and performance. For example, the diverse implementation and intentional non-identity make the overall system more fault-tolerant, especially to errors caused by software (or software-hardware interfaces).For example, if a software bug or error occurs in the software on the primary computer and the non-identical software code on the secondary computer produces the same overall result, the monitoring MCU can have greater confidence that the overall result is correct and that the bug in the software or hardware on the primary computer is not causing a material error.

[0179] In some examples, the output of the ADAS system 638 can be fed into the perception block of the primary computer and / or the dynamic driving task block of the primary computer. For example, if the ADAS system 638 issues a forward collision warning due to an object immediately in front of the vehicle, the perception block can use this information in object identification. In other examples, the secondary computer may have its own trained neural network, thus reducing the risk of false positives, as described herein.

[0180] The Vehicle 600 may further include the Infotainment SoC 630 (e.g., an in-vehicle infotainment system (IVI)). Although illustrated and described as an SoC, the infotainment system may not actually be an SoC and may include two or more discrete components. The Infotainment SoC 630 may include a combination of hardware and software that can be used to provide the Vehicle 600 with audio (e.g., music, a personal digital assistant, navigation instructions, news, radio, etc.), video (e.g., TV, movies, streaming, etc.), telephone (e.g., hands-free calling), network connectivity (e.g., LTE, Wi-Fi, etc.), and / or information services (e.g., navigation systems, rear parking sensors, a radio data system, vehicle-related information such as fuel level, total distance traveled, brake fluid level, oil level, door open / close status, air filter information, etc.).The Infotainment SoC 630 can include, for example, radios, record players, navigation systems, video players, USB and Bluetooth connectivity, car computers, in-car entertainment, Wi-Fi, steering wheel audio controls, hands-free calling, a heads-up display (HUD), an HMI display 634, a telematics device, a control panel (e.g., for controlling and / or interacting with various components, functions, and / or systems), and / or other components. The Infotainment SoC 630 can also be used to provide information (e.g., visual and / or audible) to one or more vehicle users, such as information from the ADAS system 638, autonomous driving information such as planned vehicle maneuvers, road layouts, environmental information (e.g., intersection information, vehicle information, road information, etc.), and / or other information.

[0181] The infotainment SoC 630 may include GPU functionality. The infotainment SoC 630 can communicate with other devices, systems, and / or components of the vehicle 600 via the bus 602 (e.g., CAN bus, Ethernet, etc.). In some examples, the infotainment SoC 630 may be coupled with a monitoring MCU so that the infotainment system's GPU can perform some self-driving functions if one or more primary controllers 636 (e.g., the primary and / or backup computers of the vehicle 600) fail. In such an example, the infotainment SoC 630 can place the vehicle 600 into a chauffeur-to-safe-stop mode, as described herein.

[0182] The vehicle 600 may further include an instrument cluster 632 (e.g., a digital instrument cluster, an electronic instrument cluster, a digital instrument panel, etc.). The instrument cluster 632 may include a controller and / or supercomputer (e.g., a discrete controller or supercomputer). The instrument cluster 632 may include a number of instruments, such as a speedometer, fuel gauge, oil pressure gauge, tachometer, odometer, turn signals, shift position indicator, seat belt warning light(s), parking brake warning light(s), engine malfunction light(s), airbag system (SRS) information, lighting controls, safety system controls, navigation information, etc. In some examples, information from the infotainment SoC 630 and the instrument cluster 632 may be displayed and / or shared. In other words, the instrument cluster 632 may be included as part of the infotainment SoC 630, or vice versa.

[0183] Fig. 6D a system diagram for communication between (a) cloud-based server(s) and the exemplary autonomous vehicle 600 of the Fig. 6A according to some embodiments of the present disclosure. The system 676 can include the one or more servers 678, the one or more networks 690, and the vehicles, including the vehicle 600. The server(s) 678 can include a plurality of GPUs 684(A)-684(H) (hereinafter collectively referred to as GPUs 684), PCIe switches 682(A)-682(H) (hereinafter collectively referred to as PCIe switches 682), and / or CPUs 680(A)-680(B) (hereinafter collectively referred to as CPUs 680). The GPUs 684, the CPUs 680, and the PCIe switches can be interconnected by high-speed links, for example, and without limitation, the NVIDIA-developed NVLink interfaces 688 and / or PCIe links 686. In some examples, the GPUs 684 are connected via NVLink and / or NVSwitch SoC, and the GPUs 684 and the PCIe switches 682 are connected via PCIe connections.Although eight GPUs 684, two CPUs 680, and two PCIe switches are illustrated, this should not be interpreted as a limitation. Depending on the configuration, each Server 678 can include any number of GPUs 684, CPUs 680, and / or PCIe switches. For example, one or more Server 678s can each include eight, sixteen, thirty-two, and / or more GPUs 684.

[0184] The one or more servers 678 can receive image data from the vehicles via the one or more networks 690. This image data is representative of images showing unexpected or changed road conditions, such as recently started roadworks. The one or more servers 678 can transmit neural networks 692, updated neural networks 692, and / or map information 694 to the vehicles via the one or more networks 690. This map information includes traffic and road condition data. The map information updates 694 can include updates to the HD map 622, such as information about construction sites, potholes, detours, flooding, and / or other obstacles.In some examples, the neural networks 692, the updated neural networks 692 and / or the map information 694 may result from new training and / or experience represented in the data received from any number of vehicles in the environment, and / or may be based on training performed in a data center (e.g. using one or more servers 678 and / or other servers).

[0185] One or more Server 678 systems can be used to train machine learning models (e.g., neural networks) based on training data. The training data can be generated by the vehicles and / or in a simulation (e.g., using a game machine). In some examples, the training data is tagged (e.g., if the neural network benefits from supervised learning) and / or subjected to other preprocessing, while in other examples, the training data is not tagged and / or preprocessed (e.g., if the neural network does not require supervised learning).Training can be performed using one or more classes of machine learning techniques, including, but not limited to, classes such as: supervised training, semi-supervised training, unsupervised training, self-learning, reinforcement learning, federated learning, transfer learning, feature learning (including principal component and cluster analysis), multilinear subspace learning, diverse learning, representational learning (including substitute dictionary learning), rule-based machine learning, anomaly detection, and all variants or combinations thereof. Once the machine learning models are trained, they can be used by the vehicles (e.g., transmitted to the vehicles via one or more networks 690) and / or used by one or more servers 678 for remote monitoring of the vehicles.

[0186] In some examples, one or more Server 678s can receive data from the vehicles and apply that data to current real-time neural networks for intelligent inference. The one or more Server 678s can include deep learning supercomputers and / or dedicated AI computers powered by GPUs 684, such as NVIDIA's DGX and DGX Station machines. However, in some examples, the one or more Server 678s may include a deep learning infrastructure that uses only CPU-powered data centers.

[0187] The deep learning infrastructure of one or more Server 678s can perform fast, real-time inference and can use this capability to assess and verify the state of the processors, software, and / or associated hardware in the Vehicle 600. For example, the deep learning infrastructure can receive periodic updates from the Vehicle 600, such as a sequence of images and / or objects that the Vehicle 600 has located within that sequence of images (e.g., via computer vision and / or other machine learning object classification techniques).The deep learning infrastructure can run its own neural network to identify the objects and compare them with the objects identified by the vehicle 600. If the results do not match and the infrastructure concludes that the AI ​​in the vehicle 600 is not working correctly, one or more servers 678 can send a signal to the vehicle 600, instructing a fail-safe computer in the vehicle 600 to take control, notify the passengers, and perform a safe parking maneuver.

[0188] For inference, one or more Server 678s can include GPUs 684s and one or more programmable inference accelerators (e.g., NVIDIA's TensorRT). The combination of GPU-driven servers and inference accelerators can enable real-time responsiveness. In other scenarios, such as when performance is less critical, servers powered by CPUs, FPGAs, and other processors can be used for inference. EXAMPLE CALCULATION DEVICE

[0189] Fig. Figure 7 is a block diagram of an exemplary computing device 700 suitable for use in implementing some embodiments of the present disclosure. The computing device 700 may include a connection system 702 that directly or indirectly couples the following devices: main memory 704, one or more central processing units (CPUs) 706, one or more graphics processing units (GPUs) 708, a communication interface 710, input / output ports (I / O ports) 712, input / output components 714, a power supply 716, one or more presentation components 718 (e.g., display(s)), and one or more logic units 720. In at least one embodiment, the one or more computing devices 700 may comprise one or more virtual machines (VMs), and / or each of the components thereof may comprise virtual components (e.g., virtual hardware components).As non-restrictive examples, one or more of the GPUs 708 can comprise one or more vGPUs, one or more of the CPUs 706 can comprise one or more vCPUs, and / or one or more of the logic units 720 can comprise one or more virtual logic units. Thus, one or more computing device(s) 700 can include discrete components (e.g., a complete GPU allocated to computing device 700), virtual components (e.g., a portion of a GPU allocated to computing device 700), or a combination thereof.

[0190] Although the various blocks of the Fig. Where components 7 are shown connected by lines via the connection system 702, this is not intended as a limitation and is for clarity only. In some embodiments, for example, a presentation component 718, such as a display device, may be considered an I / O component 714 (if, for example, the display is a touchscreen). As another example, the CPUs 706 and / or GPUs 708 may include memory (for example, the memory 704 may represent a memory device in addition to the memory of the GPUs 708, the CPUs 706, and / or other components). In other words, the computing device of Fig. Section 7 is for illustrative purposes only. No distinction is made between categories such as "workstation", "server", "laptop", "desktop", "tablet", "client device", "mobile device", "handheld device", "game console", "electronic control unit (ECU)", "virtual reality system" and / or other types of device or system, as all are considered to be within the protective scope of the computing device. Fig. 7 are considered.

[0191] The 702 connection system can represent one or more links or buses, such as an address bus, a data bus, a control bus, or a combination thereof. The 702 connection system can include one or more bus or link types, such as an Industry Standard Architecture bus (ISA bus), an Extended Industry Standard Architecture bus (EISA bus), a Video Electronics Standards Association bus (VESA bus), a Peripheral Component Interconnect bus (PCI bus), a Peripheral Component Interconnect Express bus (PCIe bus), and / or any other type of bus or link. In some embodiments, there are direct connections between components. For example, the CPU 706 can be directly connected to the memory 704. Furthermore, the CPU 706 can be directly connected to the GPU 708. In a direct or point-to-point connection between components, the 702 connection system can include a PCIe link to establish the connection.In these examples, a PCI bus does not need to be included in the Computing Device 700.

[0192] The 704 main memory can include a variety of computer-readable media. These computer-readable media can be any available media that the 700 computing device can access. They can include both volatile and non-volatile media, as well as removable and non-removable media. By way of example and without limitation, computer-readable media can include computer storage media and communication media.

[0193] Computer storage media can include both volatile and non-volatile media, and / or removable and non-removable media, implemented in any method or technology for storing information, such as computer-readable instructions, data structures, program modules, and / or other data types. For example, main memory can store 704 computer-readable instructions (e.g., representing one or more programs and / or one or more program elements, such as an operating system).Computer storage media may include, but are not limited to, RAM, ROM, EEPROM, flash memory or other storage technology, CD-ROM, Digital Versatile Discs (DVDs) or other optical disk storage, magnetic cartridges, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to store the desired information and that the Computing Device 700 can access. As used herein, computer storage media do not include signals as such.

[0194] Computer storage media can embody computer-readable instructions, data structures, program modules, and / or other data types in a modulated data signal, such as a carrier wave or other transport mechanism, and include any media for transmitting information. The term "modulated data signal" can refer to a signal in which one or more properties are set or modified such that information is encoded in the signal. Computer storage media can include, by way of example and without limitation, wired media, such as a wired network or a directly wired connection, and wireless media, such as acoustic, RF, infrared, or other wireless media. Combinations of the above should also be included within the scope of protection of computer-readable media.

[0195] The CPU(s) 706 can be configured to execute at least some of the computer-readable instructions to control one or more components of the Computing Device 700 to perform one or more of the procedures and / or processes described herein. The CPU(s) 706 can each include one or more cores (e.g., one, two, four, eight, twenty-eight, seventy-two, etc.) capable of handling a variety of software threads concurrently. The CPU(s) 706 can include any type of processor and may include different types of processors depending on the type of Computing Device 700 implemented (e.g., processors with fewer cores for mobile devices and processors with more cores for servers).Depending on the type of computing device 700, the processor can be, for example, an Advanced RISC Machines (ARM) processor implemented using a Reduced Instruction Set Computing (RISC) processor, or an x86 processor implemented using Complex Instruction Set Computing (CISC). The computing device 700 can include one or more CPUs 706 in addition to one or more microprocessors or supplementary coprocessors, such as mathematical coprocessors.

[0196] In addition to or as an alternative to the one or more CPUs 706, the one or more GPUs 708 may be configured to execute at least some of the computer-readable instructions to control one or more components of the computing device 700 to perform one or more of the procedures and / or processes described herein. One or more of the GPUs 708 may be an integrated GPU (e.g., with one or more of the CPUs 706), and / or one or more of the GPUs 708 may be a separate GPU. In embodiments, one or more of the GPUs 708 may be a coprocessor of one or more of the CPUs 706. The GPUs 708 may be used by the computing device 700 to render graphics (e.g., 3D graphics) or to perform general-purpose computing. For example, the GPUs 708 may be used for general-purpose GPU computing (GPGPU).The GPU(s) 708 can include hundreds or thousands of cores capable of processing hundreds or thousands of software threads simultaneously. The GPU(s) 708 can generate pixel data for output images in response to rendering commands (e.g., rendering commands from the CPU(s) 706 received via a host interface). The GPU(s) 708 can include graphics memory, such as display memory for storing pixel data or other useful data, such as GPGPU data. The display memory can be included as part of the main memory 704. The GPU(s) 708 can include two or more GPUs operating in parallel (e.g., via a link). The link can connect the GPUs directly (e.g., using NVLINK) or connect them via a switch (e.g., using NVSwitch).When combined, each GPU can generate 708 pixel data or GPGPU data for different sections of an output or for different outputs (e.g., a first GPU for a first image and a second GPU for a second image). Each GPU can have its own dedicated memory or share memory with other GPUs.

[0197] In addition to or as an alternative to the CPU(s) 706 and / or the GPU(s) 708, the one or more logic units 720 may be configured to execute at least some of the computer-readable instructions to control one or more components of the computing device 700 to perform one or more of the methods and / or processes described herein. In embodiments, the CPU(s) 706, the GPU(s) 708, and / or the logic units 720 may, separately or jointly, perform any combination of methods, processes, and / or parts thereof. One or more of the logic units 720 may be part of and / or integrated into one or more of the CPUs 706 and / or one or more of the GPUs 708, and / or one or more of the logic units 720 may be discrete components or otherwise separate from the CPUs 706 and / or the GPUs 708.In embodiments, one or more logic units 720 can be a coprocessor of one or more of the CPU(s) 706 and / or one or more of the GPU(s) 708.

[0198] Examples of Logic Unit(s) 720 include one or more processing cores and / or components thereof, such as Data Processing Units (DPUs), Tensor Cores (TCs), Tensor Processing Units (TPUs), Pixel Visual Cores (PVCs), Vision Processing Units (VPUs), Graphics Processing Clusters (GPCs), Texture Processing Clusters (TPCs), Streaming Multiprocessors (SMs), Tree Traversal Units (TTUs), Artificial Intelligence Accelerators (AIAs), Deep Learning Accelerators (DLAs), Arithmetic Logic Units (ALUs), and application-specific integrated circuits. (Application-Specific Integrated Circuits, ASICs), Floating Point Units (FPUs), Input / Output (I / O) elements,Peripheral component interconnects (PCI elements) or PCI Express elements (PCIe elements), and / or the like.

[0199] The 710 communication interface can include one or more receivers, transmitters, and / or transceivers that enable the 700 computing device to communicate with other computing devices over an electronic communication network, including wired and / or wireless communications. The 710 communication interface can include components and functionality to enable communication over a number of different networks, such as wireless networks (e.g., Wi-Fi, Z-Wave, Bluetooth, Bluetooth LE, ZigBee, etc.), wired networks (e.g., communicating over Ethernet or InfiniBand), low-power wide-area networks (e.g., LoRaWAN, SigFox, etc.), and / or the Internet.In one or more embodiments, the one or more logic units 720 and / or the communication interface 710 can include one or more data processing units (DPUs) to directly transfer data received via a network and / or the connection system 702 to one or more GPUs 708 (e.g., a memory thereof).

[0200] The I / O ports 712 enable the computing device 700 to be logically coupled with other devices, including the I / O components 714, the one or more presentation components 718, and / or other components, some of which may be built into (e.g., integrated with) the computing device 700. Illustrative I / O components 714 include a microphone, mouse, keyboard, joystick, gamepad, game controller, satellite dish, scanner, printer, wireless device, etc. The I / O components 714 can provide a natural user interface (NUI) that processes air gestures, speech, or other physiological inputs generated by a user. In some cases, the inputs can be transmitted to a suitable network element for further processing.A NUI can implement any combination of speech capture, stylus capture, facial recognition, biometric recognition, gesture recognition (both on-screen and off-screen), air gestures, head and eye tracking, and touch capture (as further described below) associated with a display of the Computing Device 700. The Computing Device 700 can include depth cameras, such as stereoscopic camera systems, infrared camera systems, RGB camera systems, touchscreen technology, and combinations thereof, for gesture capture and recognition. Additionally, the Computing Device 700 can include accelerometers or gyroscopes (e.g., as part of an inertial measurement unit (IMU)) that enable motion detection. In some examples, the outputs from the accelerometers and gyroscopes can be used by the Computing Device 700 to render immersive augmented reality or virtual reality.

[0201] The power supply 716 can include a wired power supply, a battery power supply, or a combination thereof. The power supply 716 can provide power to the computing device 700 to enable the functioning of the computing device 700's components.

[0202] The 718 presentation component(s) can include a display (e.g., a monitor, touchscreen, television screen, heads-up display (HUD), other display types, or a combination thereof), speakers, and / or other presentation components. The 718 presentation component(s) can receive data from other components (e.g., the one or more 708 GPU(s), 706 CPU(s), DPUs, etc.) and output the data (e.g., as an image, video, sound, etc.). EXEMPLARY DATA CENTER

[0203] Fig. Figure 8 illustrates an exemplary data center 800 that can be used in at least one embodiment of the present disclosure. The data center 800 can include a data center infrastructure layer 810, a framework layer 820, a software layer 830, and / or an application layer 840.

[0204] As in Fig. As shown in Figure 8, the data center infrastructure layer 810 can include a resource orchestrator 812, clustered compute resources 814, and node compute resources (“node RRs”) 816(1)-816(N), where “N” represents any positive integer. In at least one embodiment, the node RRs 816(1)-816(N) can include, but are not limited to, a number of central processing units (CPUs) or other processors (including DPUs, accelerators, field-programmable gate arrays (FPGAs), graphics processing units or graphics processing units (GPUs), etc.), memory devices (e.g., dynamic read-only memories), storage devices (e.g., solid-state or disk drives), network input / output (NW I / O) devices, network switches, virtual machines (VMs), power modules, and / or cooling modules, etc. In some embodiments, one or more node RRs of the node RR can beNode RRs 816(1)-816(N) correspond to a server that has one or more of the computing resources mentioned above. Additionally, in some embodiments, the Node RRs 816(1)-816(N) may include one or more virtual components, such as vGPUs, vCPUs and / or the like, and / or one or more of the Node RRs 816(1)-816(N) may correspond to a virtual machine (VM).

[0205] In at least one embodiment, grouped compute resources 814 can include separate groupings of node RRs 816 located in one or more racks (not shown), or many racks located in data centers at different geographic locations (also not shown). Separate groupings of node RRs 816 within grouped compute resources 814 can include grouped compute, network, memory, or storage resources that can be configured or allocated to support one or more workloads. In at least one embodiment, multiple node RRs 816, including the CPUs, GPUs, DPUs, and / or other processors, can be grouped in one or more racks to provide compute resources to support one or more workloads.The one or more racks can also include any number of power modules, cooling modules and / or network switches in any combination.

[0206] The resource orchestrator 812 can configure or otherwise control one or more node RRs 816(1)-816(N) and / or grouped compute resources 814. In at least one embodiment, the resource orchestrator 812 can include a software design infrastructure management unit (SDI management unit) for the data center 800. The resource orchestrator 812 can include hardware or software, or a combination thereof.

[0207] In at least one embodiment, as in Fig. As shown in Figure 8, the framework layer 820 can include a job scheduler 833, a configuration manager 834, a resource manager 836, and / or a distributed file system 838. The framework layer 820 can include a framework to support software 832 of the software layer 830 and / or one or more applications 842 of the application layer 840. The software 832 or the application(s) 842 can each include web-based service software or applications, such as those provided by Amazon Web Services, Google Cloud, and Microsoft Azure. The framework layer 820 can be a type of free and open-source software web application framework, such as Apache Spark™ (hereinafter "Spark"), which can utilize a distributed file system 838 for handling large amounts of data (e.g., "Big Data"), without being limited to it.In at least one embodiment, the job scheduler 833 can include a Spark driver to facilitate the scheduling of workloads supported by different layers of the data center 800. The configuration manager 834 can be able to configure different layers, such as the software layer 830 and the framework layer 820, including Spark and the distributed file system 838 to support large-scale data processing. The resource manager 836 can be able to manage clustered or grouped compute resources that are mapped to or associated with the distributed file system 838 and the job scheduler 833 for support. In at least one embodiment, the clustered or grouped compute resources can include the grouped compute resources 814 on the data center infrastructure layer 810.The Resource Manager 836 can coordinate with the Resource Orchestrator 812 to manage these mapped or allocated computing resources.

[0208] In at least one embodiment, the software 832 enclosed in software layer 830 may include software used by at least sections of the node RRs 816(1)-816(N), grouped computing resources 814, and / or the distributed file system 838 of framework layer 820. One or more types of software may include, but are not limited to, internet web search software, email virus scanning software, database software, and streaming video content software.

[0209] In at least one embodiment, the application(s) 842 included in the application layer 840 may include one or more types of applications used by at least sections of the node RRs 816(1)-816(N), grouped compute resources 814, and / or the distributed file system 838 of the framework layer 820. One or more types of applications may, but are not limited to, include any number of genome applications, cognitive computation, and machine learning applications, including training or inference software, machine learning framework software (e.g., PyTorch, TensorFlow, Caffe, etc.), and / or other machine learning applications used in conjunction with one or more embodiments.

[0210] In at least one embodiment, the configuration manager 834, the resource manager 836, and / or the resource orchestrator 812 can implement any number and any type of self-modifying actions based on any set and any type of data acquired by any technically feasible means. Self-modifying actions can relieve a data center operator of the data center 800 of potentially making poor configuration decisions and potentially avoiding underutilized and / or poorly functioning parts of a data center.

[0211] The Data Center 800 may include tools, services, software, or other resources for training one or more machine learning models or for predicting or inferring information using one or more machine learning models according to one or more embodiments described herein. For example, one or more machine learning models may be trained by calculating weighting parameters according to a neural network architecture, using software and / or computing resources described above with reference to the Data Center 800.In at least one embodiment, trained or deployed machine learning models corresponding to one or more neural networks can be used to derive or predict information using resources described above with reference to the Computing Center 800, using weighting parameters computed by one or more training techniques such as those described herein.

[0212] In at least one embodiment, the data center can use 800 CPUs, application-specific integrated circuits (ASICs), GPUs, FPGAs, and / or other hardware (or equivalent virtual computing resources) to perform training and / or inference using the resources described above. Furthermore, one or more of the software and / or hardware resources described above can be configured as a service to allow users to perform training or inference on information, such as image recognition, speech recognition, or other AI services. EXEMPLARY LANGUAGE MODELS

[0213] In at least some embodiments, language models such as large language models (LLMs), vision language models (VLMs), multimodal language models (MMLMs), and / or other types of generative artificial intelligence (AI) can be implemented. These models can be capable of understanding, summarizing, translating, and / or otherwise generating text (e.g., natural language text, code, etc.), images, video, computer-aided design (CAD) assets, OMNIVERSE and / or METAVERSE file information (e.g., in USD format, such as OpenUSD), and / or the like, based on the context provided in prompts or queries. These language models can be considered "large" in embodiments based on the models being trained on massive datasets and architectures with a large number of learnable network parameters (weights and biases), such as millions or billions of parameters. The LLMs / VLMs / MMLMs / etc.They can be implemented for summarizing text data, analyzing and extracting insights from data (e.g., text, image, video data, etc.), and generating new text / images / videos, etc., in user-specified styles, nuances, and / or formats. The LLMs / VLMs / MMLMs / etc. of this disclosure can be used exclusively for text processing in some embodiments, while in other embodiments, multimodal LLMs can be implemented to accept, understand, and / or generate text and / or any other type of content, such as images, audio, 2D and / or 3D data (e.g., in USB formats), and / or video. Vision language models (VLMs) or, more generally, multimodal language models (MMLMs) can, for example, be implemented to process image, video, audio, text, 3D design (e.g.,to accept (in CAD) and / or other input data types and / or to generate or output image, video, audio, text, 3D design and / or other output data types.

[0214] Different types of LLM / VLM / MMLM / etc. architectures can be implemented in various embodiments. For example, different architectures can be implemented that use different techniques for understanding and / or generating outputs such as text, audio, video, image, 2D and / or 3D design or asset data, etc. In some embodiments, the LLM / VLM / MMLM / etc. architectures can use recurrent neural networks (RNNs) or long / short-term memory networks (LSTM networks), while in other embodiments, transformer architectures, such as those based on self-attention and / or cross-attention mechanisms (e.g., between context data and text data), can be used to understand and recognize relationships between words or tokens and / or context data (e.g., other text, video, image, design data, USD, etc.).One or more generative processing pipelines, including LLMs, VLMs, MMLMs, etc., may also include one or more diffusion blocks (e.g., noise suppressors). The LLMs, VLMs, MMLMs, etc., of this disclosure may include one or more encoder and / or decoder blocks. Discriminative or pure encoder models, such as BERT (Bidirectional Encoder Representations from Transformers), may be implemented for tasks involving language understanding, such as classification, sentiment analysis, and question answering, and may be referred to as entity recognition. As another example, generative or decoder-only models, such as GPC (Generative Pretrained Transformer), may be implemented for tasks involving language and content generation, such as text completion, story creation, and dialogue generation.Architectures that include both encoder and decoder components, such as the T5 (Text-to-Text Transformer), can be implemented to understand and generate content, for example, for translation and summarization. These examples are intended to be limiting, and any architecture type, including but not limited to those described herein, can be implemented depending on the specific implementation and the task(s) performed using the LLMs / VLMs / MMLMs / etc.

[0215] In various embodiments, LLMs / VLMs / MMLMs / etc. can be trained using unsupervised learning, where an LLM / VLM / MMLM / etc. learns patterns from large amounts of unlabeled text, audio, video, image, design, and USD data. Due to comprehensive training, models in some embodiments may not require task-specific or domain-specific training. LLMs / VLMs / MMLMs / etc. that have undergone comprehensive pretraining on large amounts of unlabeled data can be referred to as foundational models and may be suitable for a variety of tasks, such as question-answering, summarizing, filling in missing information, translation, and image / video / design / USD / data generation. Some LLMs / VLMs / MMLMs / etc. may be tailored for a specific use case using techniques such as prompt tuning, fine-tuning, retrieval augmented generation (RAG), and adding adapters (e.g.,adapted neural networks and / or neural network layers that can tune or adjust input prompts to align the language model towards a particular task or domain), and / or the use of other fine-tuning or adaptation techniques that optimize the models for use in a particular task and / or within particular domains.

[0216] In some embodiments, the LLMs / VLMs / MMLMs / etc. of the present disclosure can be implemented using various model alignment techniques. For example, in some embodiments, safeguards can be implemented to identify inappropriate or unwanted inputs (e.g., prompts) and / or outputs of the models. The system can use the safeguards and / or other model alignment techniques to either prevent a particular unwanted input from being processed using the LLMs / VLMs / MMLMs / etc., and / or to prevent the output or presentation (e.g., display, audio output, etc.) of information from being generated using the LLMs / VLMs / MMLMs / etc. In some embodiments, one or more additional models or layers thereof can be implemented to identify problems with the inputs and / or outputs of the models.These "protective measures" models can, for example, be trained to identify inputs and / or outputs that are "safe" or otherwise OK or desirable, and / or that are "unsafe" or otherwise undesirable for the particular application / implementation. Consequently, the LLMs / VLMs / MMLMs / etc. of this disclosure may be less likely to output speech / text / audio / design data / USD data, etc., that is offensive, vulgar, inappropriate, unsafe, out of scope, and / or otherwise undesirable for the particular application / implementation.

[0217] In some implementations, the LLMs / VLMs / etc. can be configured or enabled to access or use more plug-ins, application programming interfaces (APIs), databases, data stores, data pools, etc. For certain tasks or operations for which the model is not ideally suited, the model can, for example, have instructions (e.g., a training result and / or based on instructions in a given prompt) to access one or more plug-ins (e.g., third-party plug-ins) to assist in processing the current input. In such an example, where at least part of a prompt concerns restaurants or weather, the model can access one or more restaurant or weather plug-ins (e.g., via one or more APIs) to retrieve the relevant information.As another example, where at least part of an answer requires a mathematical calculation, the model can access one or more mathematical plugins or APIs to assist in solving the problem(s) and can then use the answer from the plugin and / or API in the model's output. This process can be repeated recursively over a number of iterations using any number of plugins and / or APIs until an answer to the input prompt can be generated that addresses any task / question / requirement / process / operation, etc. The model(s) can therefore draw not only on its own knowledge acquired through training on one or more large datasets, but also on the expertise and optimized nature of one or more external resources, such as APIs, plugins, and / or the like.

[0218] In some embodiments, multiple language models (e.g., LLMs / VLMs / MMLMs / etc.), multiple instances of the same language model, and / or multiple prompts served by the same language model or instance of the same language model can be implemented, executed, or accessed (e.g., using one or more plug-ins, user interfaces, APIs, databases, data stores, data pools, etc.) to provide output in response to the same query or in response to separate sections of a query. In at least one embodiment, multiple language models, e.g., language models with different architectures, language models trained on different (e.g., updated) datasets, can be served with the same input query and prompt (e.g., set of constraints, conditioners, etc.).In one or more embodiments, the language models can be different versions of the same underlying model. In one or more embodiments, at least one language model can be instantiated as multiple agents; for example, more than one prompt can be provided to restrict, guide, or otherwise influence a style, content, character, etc., of the output provided. In one or more example, non-restrictive embodiments, the same language model can be instructed to provide output corresponding to a different role, perspective, character, or knowledge base, etc., as defined by a provided prompt.

[0219] In each such embodiment, the output of two or more (e.g., each) language models, of two or more versions of at least one language model, of two or more instantiated agents of at least one language model, and / or of two or more prompts provided to at least one language model can be further processed, for example, aggregated, compared, filtered, or used to determine (and provide) a consensus response. In one or more embodiments, the output from a language model or version, instance, or agent can perhaps be provided as input to another language model for further processing and / or validation. In one or more embodiments, the language model can be prompted to produce or otherwise obtain output with respect to input source material, the output being associated with the input source material.Such an assignment might include, for example, the generation of a heading or text passage that is embedded (e.g., as metadata) in an input source text or image. In one or more embodiments, an output from a language model can be used to determine the validity of input source material for further processing or inclusion in a dataset. For example, a language model can be used to assess the presence (or absence) of a target word in a text passage or an object in an image, annotating the text or image to indicate such presence (or lack thereof). Alternatively, the determination from the language model can be used to determine whether the source material should be included in a carefully selected dataset, for example, without restriction.

[0220] Fig. Figure 9A is a block diagram of an exemplary generative language model system 900, suitable for use in implementing at least some embodiments of the present disclosure. In the Fig. 9A illustrated example includes the generative language model system 900 a Retrieval Augmented Generation (RAG) component 992, an input processor 905, a tokenizer 910, an embedding component 920, plug-ins / APIs 995 and a generative language model (LM) 930 (which may include an LLM, a VLM, a multimodal LM, etc.).

[0221] At a high level, the input processor 905 can receive an input 901 that includes text and / or other input data (e.g., audio data, video data, image data, sensor data (e.g., LiDAR, RADAR, ultrasound, etc.), 3D design data, CAD data, Universal Scene Descriptor data (USD data – such as OpenUSD, etc.), depending on the architecture of the generative LM 930 (e.g., LLM / VLM / MMLM / etc.)). In some embodiments, the input 901 includes plain text in the form of one or more sentences, paragraphs, and / or documents. Additionally or alternatively, the input 901 can include number sequences, pre-computed embeddings (e.g., word or sentence embeddings), and / or structured data (e.g., in tabular formats, JSON, or XML).In some implementations where the generative LM 930 is capable of processing multimodal inputs, the LM 930 input processor can combine (or omit) text with image data, audio data, video data, design data, USD data, and / or other types of input data, such as, but not limited to, those described herein. Using raw input text as an example, the LM 930 input processor can prepare raw input text in various ways. For instance, the LM 930 input processor can perform various types of text filtering to remove noise from relevant text content (e.g., special characters, punctuation, HTML tags, stop words, parts of images, parts of audio files, etc.). In an example involving stop words (common words that tend to contain little semantic meaning), the LM 930 input processor can remove stop words to reduce noise and allow the generative LM 930 to focus on more meaningful content.The 905 input processor can normalize text, for example, by converting all characters to lowercase, removing accents, and / or handling special cases such as contractions or abbreviations to ensure consistency. These are just a few examples, and other types of input processing can be applied.

[0222] In such embodiments, a RAG component 992 (which may include one or more RAG models and / or be implemented using the generative LM 930 itself) can be used to retrieve additional information for use as part of the input 901 or prompt. RAG can be used to enhance the input to the LLM / VLM / MMLM / etc. with external knowledge, making answers to specific questions, retrievals, or requests more relevant, as in a case where specific knowledge is required. The RAG component 992 can retrieve this additional information (e.g., background information such as background text / images / videos / audio files / USD / CAD / etc.) from one or more external sources, which can then be passed along with the prompt to the LLM / VLM / MMLM / etc. to improve the accuracy of the model's answers or outputs.

[0223] In some embodiments, for example, the input 901 can be generated using the retrieval or input into the model (e.g., a question, a request, etc.) in addition to data retrieved using the RAG component 992. In some embodiments, the input processor 905 can analyze the input 901 and communicate with the RAG component 992 (or the RAG component 992 can be part of the input processor 905 in some embodiments) to identify relevant text and / or other data to provide to the generative LM 930 as additional context or information sources from which the reaction, response, or output 990 is to be generally identified.For example, if the input indicates that the user is interested in a desired tire pressure for a specific make and model of vehicle, the RAG component 992 can, using a RAG model that performs, for example, a vector search in an embedding space, retrieve the tire pressure information or corresponding text from the digital (embedded) version of the user manual for that specific make and model of vehicle. If a user revisits a chatbot in connection with a specific product offering or service, the RAG component 992 can retrieve a previously stored conversation history, or at least a summary thereof, and include the previous conversation history along with the current query / request as part of the input 901 to the generative LM 930.

[0224] The RAG component 992 can employ various RAG techniques. Naive RAG, for example, can be used for indexing, splitting documents into blocks, and applying it to an embedding model to generate embeddings that correspond to chunks. A user query can also be applied to the embedding model and / or another embedding model of the RAG component 992, and the chunk embeddings can be compared with the query's embeddings to identify the embeddings that are most similar to / correspond to the query. This information can then be passed to the generative LM 930 to generate output.

[0225] In some embodiments, more advanced RAG techniques can be used. For example, before passing chunks to the embedding model, the chunks can be subjected to pre-fetching processes (e.g., routing, rewriting, metadata analysis, expansion, etc.). Additionally, before generating the final embeddings, post-fetching processes (e.g., reordering, prompt compression, etc.) can be performed on the outputs of the embedding model prior to the final embeddings for comparison with an input query.

[0226] Another example is modular RAG techniques, such as those similar to naive and / or advanced RAG, but which also include features like hybrid search, recursive retrieval and query engines, step-back approaches, subqueries, and hypothetical document embedding.

[0227] As another example, Graph RAG can use knowledge graphs as a source of context or factual information. Graph RAG can be implemented using a graph database as a source of contextual information that is sent to the LLM / VLM / MMLM / etc. Instead of populating the model with chunks of data extracted from larger documents (or in addition to this, which can lead to a lack of context, factual accuracy, linguistic precision, etc.), Graph RAG can also provide the LLM / VLM / MMLM / etc. with structured entity information by combining the structured text description of the entity with its numerous properties and relationships, allowing the model to gain deeper insights. When implementing Graph RAG, the systems and procedures described herein use a graph as a content store, extract relevant chunks from documents, and request them from the LLM / VLM / MMLM / etc.to use them when answering questions. In such embodiments, the knowledge graph can contain relevant text content and metadata about the knowledge graph, as well as have an integrated vector database. In some embodiments, Graph RAG can use a graph as a subject matter expert, extracting relevant descriptions of concepts and entities for a query / prompt and passing them to the model as semantic context. These descriptions can include relationships between the concepts. In other examples, the graph can be used as a database, where part of a query / prompt can be mapped to a graph query, the graph query can be executed, and the LLM / VLM / MMLM / etc. can summarize the results.In such an example, the graph can store relevant factual information, and a query (natural language query) to a graph query tool (NL-to-graph query tool) as well as entity joining can be used. In some implementations, graph RAG (e.g., using a graph database) can be combined with standard RAG (e.g., vector database) and / or other RAG types to utilize multiple approaches.

[0228] In any embodiment, the RAG component 992 can implement a plug-in, API, user interface, and / or other functionality for performing RAG. For example, a graph RAG plug-in can be used by the LLM / VLM / MMLM / etc. to query the knowledge graph and extract relevant information for feeding into the model, and a standard or vector RAG plug-in can be used to query a vector database. The graph database can, for example, interact with the plug-in's REST interface in such a way that the graph database is decoupled from the vector database and / or the embedding models.

[0229] The Tokenizer 910 can segment text data into smaller units (tokens) for subsequent analysis and processing. Tokens can represent individual words, partial words, characters, or sections of audio, video, image, etc., depending on the implementation. Word-based tokenization divides the text into individual words, with each word treated as a separate token. Part-word tokenization distributes words into smaller, meaningful units (e.g., prefixes, suffixes, stems), enabling the generative LM 930 to understand morphological variations and handle out-of-vocabulary words more effectively. Character-based tokenization represents each character as a separate token, allowing the generative LM 930 to process text at a fine-grained level. The choice of tokenization strategy may depend on factors such as the processing of the language, the task at hand, and / or characteristics of the training dataset.The Tokenizer 910 can therefore convert the (e.g. processed) text into a structured format according to a tokenization system implemented in the particular embodiment.

[0230] The Embedding Component 920 can use any known embedding technique to transform individual tokens into representations with semantic meaning (e.g., dense continuous vectors). For example, the Embedding Component 920 can use pre-trained word embeddings (e.g., Word2Vec, GloVe, or FastText), one-hot coding, term-frequency-inverse document-frequency coding (TF-IDF coding), one or more neural network embedding layers, and / or other methods.

[0231] In some implementations where the input 901 includes image data / video data / etc., the input processor 901 can resize the data to a standard size compatible with the format of a corresponding input channel and / or normalize pixel values ​​to a common range (e.g., 0 to 1) to ensure consistent display. The embedding component 920 can encode the image data using a known technique (e.g., using one or more convolutional neural networks (CNNs) to extract visual features). In some implementations where the input 901 includes audio data, the input processor 901 can resample an audio file to a consistent sampling rate for uniform processing, and the embedding component 920 can use any known technique to extract and encode audio features, such as in the form of a spectrogram (e.g., a Mel spectrogram).In some implementations where the input 901 includes video data, the input processor 901 can extract or resize images, and the embedding component 920 can extract features such as optical flow embeddings or video embeddings and / or encode temporal information or image sequences. In some implementations where the input 901 includes multimodal data, the embedding component 920 can merge representations of different types of data (e.g., text, image, audio, USD, video, design, etc.) using techniques such as early fusion (concatenation), late fusion (sequential processing), attention-based fusion (e.g., self-attention, cross-attention), etc.

[0232] The generative LM 930 and / or other components of the generative LM system 900 can use different types of neural network architectures depending on the implementation. Transformer-based architectures, such as those used in models like GPT, can be implemented and may include self-attention mechanisms that weight the importance of different words or tokens in the input sequence and / or feedforward networks that process the output of the self-attention layers, apply nonlinear transformations to the input representation, and extract higher-level features. Some non-restrictive example architectures include transformers (e.g.,Encoder-decoder, decoder-only, multimodal), RNNs, LSTMs, fusion models, diffusion models, crossmodal embedding models that learn shared embedding spaces, graph neural networks (GNNs), hybrid architectures that combine different types of architectures, adversarial networks such as generative adversarial networks or GANs, or adversarial autoencoders (AAEs) for shared distributional learning, and others. Therefore, depending on the implementation and architecture, the embedding component 920 can apply a coded representation of the input 901 to the generative LM 930, and the generative LM 930 can process the coded representation of the input 901 to produce an output 990 that may include response text and / or other types of data.

[0233] As described herein, in some embodiments the generative LM 930 may be configured to access or use plug-ins / APIs 995 (which may include one or more plug-ins, application programming interfaces (APIs), databases, data stores, data pools, etc.). For certain tasks or operations for which the generative LM 930 is not ideally suited, the model may have instructions (e.g., as a result of training and / or based on instructions in a given prompt, such as those retrieved using the RAG component 992) to access one or more plug-ins / APIs 995 (e.g., third-party plug-ins) to assist in processing the current input. In such an example, where at least part of the prompt concerns restaurants or weather, the model may access one or more restaurant or weather plug-ins (e.g.,Accessing one or more APIs, at least part of the input request regarding the specific plug-in / API 995 is sent to the plug-in / API 995. The plug-in / API 995 can process the information and send a response back to the generative LM 930, and the generative LM 930 can use the response to generate the output 990. This process can be recursively performed over a number of iterations using any number of plug-ins / APIs 995 until an output 990 is generated that handles every request / question / request / process / operation / etc. from the input 901. The model(s) can therefore not only rely on its own knowledge from training on one or more large datasets and / or on data retrieved using the RAG component 992, but also on expertise and optimized nature of one or more external resources, such as the plug-ins / APIs 995.

[0234] Fig. Figure 9B is a block diagram of an example implementation where the generative LM 930 includes a transformer-encoder-decoder. For example, assume that input text such as "Who discovered gravity" (e.g., through the Tokenizer 910 of the LM 930) is used to create a 930 Transformer-Encoder-Decoder. Fig. 9A) is tokenized into tokens such as words, and each token (e.g., through the embedding component 920 of the Fig. 9A) is encoded into a corresponding embedding (e.g., with size 512). Since these token embeddings typically do not represent the token's position in the input sequence, any known technique can be used to add positional encoding to each token embedding to encode the sequential relationship and context of the tokens in the input sequence. Therefore, the resulting embeddings can be applied to one or more Encoders 935 of the generative LM 930.

[0235] In an exemplary implementation, the Encoder 935 forms an encoder stack, with each encoder including a self-attention layer and a feedforward network. In an exemplary transformer architecture, each token (e.g., each word) flows through a separate path. Each encoder can therefore accept a sequence of vectors that take each vector through the self-attention layer, then the feedforward network, and then up to the next encoder in the stack. Any known self-attention technique can be used.For example, to calculate a self-attention value for each token (word), a query vector, a key vector, and a value vector can be created for each token. A self-attention value can be calculated for token pairs by taking the dot product of the query vector with the corresponding key vectors, normalizing the resulting values, multiplying them by the corresponding value vectors, and summing the weighted value vectors. The encoder can apply multi-headed attention, where the attention mechanism is applied multiple times in parallel with different learned weight matrices. Any number of encoders can be cascaded to generate a context vector that encodes the input. An attention projection layer 940 can convert the context vector into attention vectors (keys and values) for the decoder(s) 945.

[0236] In an exemplary implementation, the Decoder(s) 945 form a Decoder stack, with each Decoder including a Self-Attention layer, an Encoder-Decoder Self-Attention layer that uses the attention vectors (keys and values) from the Encoder to focus on relevant parts of the input sequence, and a feedforward network. As with the Encoder(s) 935, in an exemplary Transformer architecture, each token (e.g., word) flows through a separate path into the Decoder(s) 945. During a first pass, the Decoder(s) 945, a Classifier 950, and a Generating Mechanism 955 can generate an initial token, and the Generating Mechanism 955 can apply the generated token as an input during a second pass. The process can repeat in a loop by successively generating tokens (e.g., words, words, words, etc.).Words) are generated and added to the output from the previous pass, and the token embeddings of the composite sequence with positional encoding are applied as an input to the decoder(s) 945 during a subsequent pass, generating one token at a time (known as autoregression) until a character or token representing the end of the response is predicted. Within a decoder, the self-attention layer is typically restricted to dealing only with previous positions in the output sequence by applying a masking technique (e.g., setting future positions to negative infinity) before the softmax operation. In an exemplary implementation, the encoder-decoder-attention layer works similarly to (e.g., multi-headed) self-attention in the encoder(s) 935, except that it establishes its fetches from the layer below and the keys and values ​​(e.g.,Matrix) from the output of the encoder 935.

[0237] The Decoder(s) 945 can therefore output a somewhat decoded (e.g., vector) representation of the input applied during a given pass. The Classifier 950 can include a multi-class classifier comprising one or more neural network layers that project the decoded (e.g., vector) representation into an appropriate dimensionality (e.g., one dimension for each supported word or token in the output vocabulary), as well as a softmax operation that converts logits into probabilities. The Generation Mechanism 955 can therefore select or extract a word or token based on an appropriate predicted probability (e.g., selecting the word with the highest predicted probability) and append it to the output from a previous pass, so that each word or token is generated sequentially.The 955 generation mechanism can repeat the process, triggering successive decoder inputs and corresponding predictions, until a character or token is selected or sampled that represents the end of the response, at which point the 955 generation mechanism can output the generated response.

[0238] Fig. 9C is a block diagram of an exemplary implementation where the generative LM 930 incorporates a decoder-transformer-only architecture. The Decoder 960 of the Fig. 9C decoders, for example, can work similarly to the 945 decoder(s). Fig. 9B, with the exception that each decoder 960 of the Fig. 9C omits the encoder-decoder self-attention layer (since there is no encoder in this implementation). The Decoder 960(s) therefore form a decoder stack, in which each decoder includes a self-attention layer and a feedforward network. Furthermore, instead of encoding the input sequence, a character or token representing the end of the input sequence (or the beginning of the output sequence) can be appended to the input sequence, and the resulting sequence (which corresponds, for example, to embeddings with positional encodings) can be applied to the Decoder 960(s). As with the Decoder 945(s), Fig. In 9B, each token (e.g., each word) can flow through a separate path to the decoder(s) 960, and the decoder(s), a classifier 965, and a generation mechanism 970 can use autoregression to sequentially generate one token at a time until a character or token is predicted that represents the end of the response. The classifier 965 and the generation mechanism 970 can perform similar work to the classifier 950 and the generation mechanism 955 of the Fig. 9B, wherein the generation mechanism 970 selects or samples each successive output token based on a corresponding predicted probability and appends it to the output from a previous pass, each token being generated sequentially until a character or token representing the end of the response is selected or sampled. These and other architectures described herein are merely examples, and other suitable architectures may be implemented within the scope of this disclosure. EXEMPLARY NETWORK ENVIRONMENTS

[0239] Network environments suitable for use in implementing embodiments of the disclosure may include one or more client devices, servers, network attached storage (NAS), other backend devices, and / or device types. The client devices, servers, and / or other device types (e.g., each device) may be connected to one or more instances of the one or more computing devices. Fig. 7. Each device can include similar components, features, and / or functionality to one or more computing devices 700. If backend devices (e.g., servers, NAS, etc.) are implemented, the backend devices can additionally be included as part of a data center 800, an example of which is described in more detail herein with reference to Fig. 8 is described.

[0240] The components of a network environment can communicate with each other over one or more networks, which can be wired, wireless, or both. The network can include multiple networks or a network of networks. For example, the network can include one or more Wide Area Networks (WANs), one or more Local Area Networks (LANs), one or more public networks, such as the Internet, and / or public networks, such as a Public Switched Telephone Network (PSTN), and / or one or more private networks. If the network includes a wireless telecommunications network, components such as a base station, a communications tower, or even access points (as well as other components) can provide wireless connectivity.

[0241] Compatible network environments can include one or more peer-to-peer network environments, where a server may not be included in a network environment, and one or more client-server network environments, where one or more servers may be included in a network environment. In peer-to-peer network environments, the functionality described herein can be implemented with reference to one or more servers on any number of client devices.

[0242] In at least one embodiment, a network environment can include one or more cloud-based network environments, a distributed computing environment, a combination thereof, etc. A cloud-based network environment can include a framework layer, a job scheduler, a resource manager, and a distributed file system implemented on one or more servers, which may include one or more core network servers and / or edge servers. A framework layer can include a framework for supporting software of a software layer and / or one or more applications of the application layer. The software or application(s) can each include web-based service software or applications. In embodiments, one or more of the client devices can use the web-based service software or applications (e.g.,by accessing the service software and / or applications via one or more application programming interfaces (APIs). The framework layer can be a type of free and open-source software web application framework that uses, for example, a distributed file system for processing large amounts of data (e.g., "Big Data"), but is not limited to this.

[0243] A cloud-based network environment can provide cloud computing and / or cloud storage, performing any combination (or one or more parts) of the computing and / or data storage functions described herein. Any of these various functions can be distributed across multiple locations from central or core servers (e.g., from one or more data centers distributed across a state, region, country, the Earth, etc.). If a connection to a user (e.g., a client device) is relatively close to one or more edge servers, one or more core servers can allocate at least one part of the functionality to the edge server(s). A cloud-based network environment can be private (e.g., restricted to a single organization), public (e.g., available to many organizations), and / or a combination thereof (e.g., a hybrid cloud environment).

[0244] The client device(s) may incorporate at least some of the components, features, and functionality of the exemplary computing device(s) 700 described herein with reference to Fig.7. For example, and without limitation, a client device may be a personal computer (PC), a laptop computer, a mobile device, a smartphone, a tablet computer, a portable computer, a personal digital assistant (PDA), an MP3 player, a virtual reality headset, a global positioning system (GPS) or device, a video player, a video camera, a surveillance device or system, a vehicle, a boat, a flying vehicle, a virtual machine, a drone, a robot, a handheld communication device, a hospital device, a gaming device or system, an entertainment device, a vehicle computing system, an embedded system controller, a household appliance, a consumer electronics system, a workstation, an edge device, any combination of these devices, or any other suitable device.

[0245] 1. In some embodiments, a computer-implemented method comprises obtaining, using one or more computing devices associated with a robotics system, one or more sensor values ​​associated with one or more sensors of the robotics system; determining, using a machine learning model and based on at least one or more sensor values, that the robotics system requires intervention; obtaining, using the one or more computing devices, one or more sensor values ​​associated with the intervention; generating an intervention event based on at least one or more first sensor values ​​and one or more second sensor values; and modifying an intervention event database based on at least the generated intervention event.

[0246] 2. Computer-implemented method according to clause 1, further comprising generating a warning, wherein the warning includes a request for intervention.

[0247] 3. Computer-implemented method according to clause 1 or 2, wherein the one or more sensor values ​​include one or more of sensor data, LiDAR sensor data, RADAR sensor data, SONAR sensor data, ultrasonic sensor data, IMU sensor data or infrared sensor data.

[0248] 4. Computer-implemented method according to any one of clauses 1-3, wherein the one or more additional sensor values ​​include one or more of a position, movement, orientation, load, stress or torque as indicated or experienced by one or more motor-controlled arrangements included in the robot system.

[0249] 5. A computer-implemented method according to any of clauses 1-4, further comprising the generation, using a machine-learning model, of a natural language description associated with the intervention.

[0250] 6. Computer-implemented method according to any of clauses 1-5, further comprising classification using a machine learning model of intervention into one or more categories or one or more failure modes.

[0251] 7. Computer-implemented method according to any of clauses 1-6, further comprising the generation of one or more entries into a training or validation dataset based at least on the intervention.

[0252] 8. Computer-implemented method according to any of clauses 1-7, further comprising initiating automatic retraining of one or more functions of an autonomy software stack based on at least the training or validation data set, wherein the autonomous software stack includes one or more autonomous or semi-autonomous control routines associated with the robotics system.

[0253] 9. Computer-implemented method according to any one of clauses 1-8, wherein the method is performed by at least: a control system for an autonomous or semi-autonomous machine, a perception system for an autonomous or semi-autonomous machine, a system for performing simulation operations, a system for performing digital twinning operations, a system for performing light transport simulation, a system for performing collaborative content creation for 3D assets, a system for performing deep learning operations, a system for performing remote operations, a system for performing real-time streaming, a system for generating or presenting one or more augmented reality content, virtual reality content, or mixed reality content, a system implemented using an edge device, or a system implemented using a robot.a system for performing operations with conversational AI, a system for implementing one or more multimodal language models, a system for implementing one or more large language models (LLMs), a system for implementing one or more vision language models (VLMs), a system for generating synthetic data; a system for generating synthetic data using AI, a system that includes one or more virtual machines (VMs), a system that is at least partially implemented in a data center, or a system that is at least partially implemented using cloud computing resources.

[0254] 10. In some embodiments, one or more processors comprising processing circuit arrangements determine, using one or more computing devices associated with a robot, one or more sensor values ​​associated with one or more sensors of the robot; determine, using a machine learning model and based on the one or more first sensor values, that the robot requires intervention; using the one or more computing devices, obtain one or more second sensor values ​​associated with the intervention; generate an intervention event based on at least the one or more first sensor values ​​and the one or more second sensor values; and modify an intervention event database based on at least the generated intervention event.

[0255] 11. One or more processors according to clause 10, wherein the one or more sensor values ​​include one or more of image sensor data, LiDAR sensor data, RADAR sensor data, SONAR sensor data, ultrasonic sensor data, IMU sensor data or infrared sensor data.

[0256] 12. One or more processors according to clauses 10 or 11, wherein the one or more additional sensor values ​​include one or more of a position, movement, orientation, load, stress or torque as indicated or experienced by one or more motor-controlled arrangements of the robot.

[0257] 13. One or more processors according to any of clauses 10-12, wherein the processing circuit arrangements further generate a natural language description associated with the intervention using a machine learning model.

[0258] 14. One or more processors according to any of clauses 10-13, wherein the processing circuit arrangements further classify the interference into one or more categories or one or more failure modes using a machine learning model.

[0259] 15. One or more processors according to any of clauses 10-14, wherein the processing circuit arrangements further generate one or more entries into a training or validation data set based at least on the intervention.

[0260] 16. One or more processors according to any of clauses 10-15, wherein the processing circuit arrangements further initiate automatic retraining of at least one section of an autonomy software stack based on at least the training or validation data set, and wherein the autonomy software stack includes one or more autonomous or semi-autonomous control routines associated with the robot.

[0261] 17. One or more processors according to any one of clauses 10 to 16, wherein the one or more processors comprise at least the following: a control system for an autonomous or semi-autonomous machine, a perception system for an autonomous or semi-autonomous machine, a system for performing simulation operations, a system for performing digital twinning operations, a system for performing light transport simulation, a system for performing collaborative content creation for 3D assets, a system for performing deep learning operations, a system for performing remote operations, a system for performing real-time streaming;a system for generating or presenting one or more augmented reality content, virtual reality content, or mixed reality content; a system implemented using an edge device; a system implemented using a robot; a system for performing operations with conversational AI; a system for implementing one or more multimodal language models; a system for implementing one or more large language models (LLMs); a system for implementing one or more vision language models (VLMs); a system for generating synthetic data; a system for generating synthetic data using AI; a system that includes one or more virtual machines (VMs); a system that is at least partially implemented in a data center; or a system that is at least partially implemented using cloud computing resources.

[0262] 18. In some embodiments, a teleoperation system comprises one or more processors for sending one or more messages to a deployed robot to cause the deployed robot to perform one or more operations associated with a particular intervention event, wherein the intervention event is determined at least on the basis of one or more multimodal language models that process sensor data obtained using one or more sensors of the deployed robot, wherein the one or more signals are generated on the basis of at least one or more inputs to the teleoperation system, wherein the one or more inputs are received in response to the presentation of event information determined using the one or more multimodal language models and corresponding to the intervention event.

[0263] 19. System according to claim 18, wherein the teleoperation system comprises at least one of the following: a control system for an autonomous or semi-autonomous machine, a perception system for an autonomous or semi-autonomous machine, a system for performing simulation operations, a system for performing digital twinning operations, a system for performing light transport simulation, a system for performing collaborative content creation for 3D assets, a system for performing deep learning operations, a system for performing remote operations, a system for performing real-time streaming, a system for generating or presenting one or more augmented reality content, virtual reality content, or mixed reality content, a system implemented using an edge device, a system implemented using a robot, a system for performing conversational AI operations,a system for implementing one or more multimodal language models, a system for implementing one or more large language models (LLMs), a system for implementing one or more vision language models (VLMs), a system for generating synthetic data, a system for generating synthetic data using AI, a system that includes one or more virtual machines (VMs), a system that is at least partially implemented in a data center, or a system that is at least partially implemented using cloud computing resources.

[0264] 20. Teleoperation system according to clauses 18 or 19, wherein the one or more operations include at least one of navigating to a location specified in the one or more messages, switching off one or more components of the deployed robot, changing an operating state of the deployed robot, or causing the presentation of an indication of a changed operating state of the deployed robot.

[0265] The disclosure can be described in the general context of computer code or machine-usable instructions, including computer-executable instructions such as program modules that are executed by a computer or other machine, such as a personal data assistant or other handheld device. In general, program modules, which include routines, programs, objects, components, data structures, etc., refer to code that performs specific tasks or implements particular abstract data types. The disclosure can be implemented in a variety of system configurations, including handheld devices, consumer electronics, general-purpose computers, more specialized computing devices, etc. The disclosure can also be implemented in distributed computing environments, in which tasks are performed by remote processing devices linked by a communication network.

[0266] As used herein, any mention of "and / or" in relation to two or more elements should be interpreted as referring to only one element or a combination of elements. For example, "element A, element B and / or element C" can mean only element A, only element B, only element C, element A and element B, element A and element C, element B and element C, or elements A, B, and C. Additionally, "at least one of element A or element B" can include at least one element A, at least one element B, or at least one element A and at least one element B. Furthermore, "at least one element A and element B" can include at least one element A, at least one element B, or at least one element A and at least one element B.

[0267] The subject matter of the present disclosure is described herein with a certain degree of accuracy in order to comply with legal requirements. However, the description itself is not intended to limit the scope of protection of this disclosure. Instead, the inventors have considered that the claimed subject matter could also be embodied in other types, so that it includes different steps or combinations of steps similar to those described in this patent specification, in conjunction with other present or future technologies.Furthermore, and although the terms “step”, “operation” and / or “block” may be used herein to denote different elements of procedures used, the terms should not be construed as a particular sequence among or between the various steps disclosed herein, except and exclusively where the sequence of individual steps is explicitly described.

[0268] It is understood that the aspects and embodiments described above are purely exemplary, and that modifications of details may be made within the scope of protection of the claims.

[0269] Each device, each method and each feature disclosed in the description, and (where applicable) the claims and drawings, may be provided independently or in any suitable combination.

[0270] Reference numerals appearing in the claims are for illustrative purposes only and do not restrict the scope of protection of the claims. QUOTES INCLUDED IN THE DESCRIPTION

[0000] This list of documents cited by the applicant was automatically generated and is included solely for the reader's convenience. The list is not part of the German patent or utility model application. The DPMA accepts no liability for any errors or omissions. Cited patent literature

[0000] US 16 / 101,232

[0124] Cited non-patent literature

[0000] Society of Automotive Engineers, SAE) (Standard No. J3016-201806

[0082] Standard No. J3016-201609, published on 30 September 2016

[0082]

Claims

A computer-implemented method comprising: obtaining, using one or more computing devices associated with a robotics system, one or more sensor values ​​associated with one or more sensors of the robotics system; determining, using a machine learning model and based on at least one or more sensor values, that the robotics system requires intervention; obtaining, using the one or more computing devices, one or more sensor values ​​associated with the intervention; generating an intervention event based on at least the one or more first sensor values ​​and the one or more second sensor values; and modifying an intervention event database based on at least the generated intervention event. Computer-implemented method according to claim 1, further comprising generating a warning, wherein the warning includes an intervention request. Computer-implemented method according to claim 1 or 2, wherein the one or more sensor values ​​include one or more of sensor data, LiDAR sensor data, RADAR sensor data, SONAR sensor data, ultrasonic sensor data, IMU sensor data or infrared sensor data. A computer-implemented method according to any preceding claim, wherein the one or more second sensor values ​​include one or more of a position, movement, orientation, load, stress or torque, as indicated or experienced by one or more motor-controlled arrangements included in the robot system. A computer-implemented method according to any preceding claim, further comprising generating, using a machine-learning model, a natural language description associated with the intervention. A computer-implemented method according to any preceding claim, further comprising classifying the intervention into one or more categories or one or more failure modes using a machine-learning model. A computer-implemented method according to any of the preceding claims, further comprising generating one or more entries into a training or validation data set based at least on the intervention. A computer-implemented method according to claim 7, further comprising initiating an automatic retraining of one or more functions of an autonomy software stack based at least on the training or validation data set, wherein the autonomous software stack includes one or more autonomous or semi-autonomous control routines associated with the robotics system. A computer-implemented method according to any of the preceding claims, wherein the method is performed by at least one of the following: a control system for an autonomous or semi-autonomous machine; a perception system for an autonomous or semi-autonomous machine; a system for performing simulation operations; a system for performing digital twinning operations; a system for performing light transport simulation; a system for performing collaborative content creation for 3D assets; a system for performing deep learning operations; a system for performing remote operations; a system for performing real-time streaming; a system for generating or presenting one or more augmented reality content, virtual reality content, or mixed reality content; a system implemented using an edge device; a system implemented using a robot;a system for performing operations with conversational AI; a system for implementing one or more multimodal language models; a system for implementing one or more large language models (LLMs); a system for implementing one or more vision language models (VLMs); a system for generating synthetic data; a system for generating synthetic data using AI; a system that includes one or more virtual machines (VMs); a system that is at least partially implemented in a data center; or a system that is at least partially implemented using cloud computing resources. One or more processors comprising processing circuit arrangements for: obtaining, using one or more computing devices associated with a robot, one or more sensor values ​​associated with one or more sensors of the robot; determining, using a machine learning model and based on the one or more first sensor values, that the robot requires intervention; obtaining, using the one or more computing devices, one or more second sensor values ​​associated with the intervention; generating an intervention event based on at least the one or more first sensor values ​​and the one or more second sensor values; and modifying an intervention event database based on at least the generated intervention event. One or more processors according to claim 10, wherein the one or more sensor values ​​include one or more of image sensor data, LiDAR sensor data, RADAR sensor data, SONAR sensor data, ultrasonic sensor data, IMU sensor data or infrared sensor data. One or more processors according to claim 10 or 11, wherein the one or more second sensor values ​​include one or more of a position, movement, orientation, load, stress or torque, as indicated or experienced by one or more motor-controlled arrangements of the robot. One or more processors according to one of claims 10 to 12, wherein the processing circuit arrangements further generate a natural language description associated with the intervention using a machine learning model. One or more processors according to one of claims 10 to 13, wherein the processing circuit arrangements further classify the interference into one or more categories or one or more failure modes using a machine learning model. One or more processors according to one of claims 10 to 14, wherein the processing circuit arrangements further generate one or more entries in a training or validation data set based at least on the intervention. One or more processors according to claim 15, wherein the processing circuit arrangements further initiate an automatic retraining of at least one section of an autonomy software stack based on at least the training or validation data set, and wherein the autonomy software stack includes one or more autonomous or semi-autonomous control routines assigned to the robot. One or more processors according to any one of claims 10 to 16, wherein the one or more processors consist of at least one of the following: a control system for an autonomous or semi-autonomous machine; a perception system for an autonomous or semi-autonomous machine; a system for performing simulation operations; a system for performing digital twinning operations; a system for performing light transport simulation; a system for performing collaborative content creation for 3D assets; a system for performing deep learning operations; a system for performing remote operations; a system for performing real-time streaming; a system for generating or presenting one or more augmented reality content, virtual reality content, or mixed reality content; a system implemented using an edge device; a system implemented using a robot;a system for performing operations with conversational AI; a system for implementing one or more multimodal language models; a system for implementing one or more large language models (LLMs); a system for implementing one or more vision language models (VLMs); a system for generating synthetic data; a system for generating synthetic data using AI; a system that includes one or more virtual machines (VMs); a system that is at least partially implemented in a data center; or a system that is at least partially implemented using cloud computing resources. A teleoperation system comprising: one or more processors for sending one or more messages to a deployed robot to cause the deployed robot to perform one or more operations associated with a specific intervention event, wherein the intervention event is determined at least on the basis of one or more multimodal language models that process sensor data obtained using one or more sensors of the deployed robot, wherein the one or more messages are generated on the basis of at least one or more inputs to the teleoperation system, wherein the one or more inputs are received in response to the presentation of event information determined using the one or more multimodal language models and corresponding to the intervention event. System according to claim 18, wherein the teleoperation system comprises at least one of the following: a control system for an autonomous or semi-autonomous machine; a perception system for an autonomous or semi-autonomous machine; a system for performing simulation operations; a system for performing digital twinning operations; a system for performing light transport simulation; a system for performing collaborative content creation for 3D assets; a system for performing deep learning operations; a system for performing remote operations; a system for performing real-time streaming; a system for generating or presenting one or more augmented reality content, virtual reality content, or mixed reality content; a system implemented using an edge device; a system implemented using a robot; a system for performing conversational AI operations;a system for implementing one or more multimodal language models; a system for implementing one or more large language models (LLMs); a system for implementing one or more vision language models (VLMs); a system for generating synthetic data; a system for generating synthetic data using AI; a system that includes one or more virtual machines (VMs); a system that is at least partially implemented in a data center; or a system that is at least partially implemented using cloud computing resources. Teleoperation system according to claim 18 or 19, wherein the one or more operations include at least one of navigating to a location specified in the one or more messages, switching off one or more components of the deployed robot, changing an operating state of the deployed robot, or causing the presentation of an indication of a changed operating state of the deployed robot.