Automatic intervention interpreter for robotic systems and applications

By using an automated robot intervention interpreter and machine learning model to detect and analyze intervention or disengagement events of autonomous or semi-autonomous robot systems online and offline, the problem of high costs associated with manual analysis is solved, and automatic updates and efficiency improvements are achieved.

CN122072476APending Publication Date: 2026-05-22NVIDIA CORP
1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
NVIDIA CORP
Filing Date
2025-11-20
Publication Date
2026-05-22

Smart Images

  • Figure CN122072476A_ABST
    Figure CN122072476A_ABST
Patent Text Reader

Abstract

In various examples, one or more systems detect an intervention or disengagement event associated with one or more autonomous or semi-autonomous robotic systems included in a cluster of robotic systems via received sensor data. Based on the detected intervention or disengagement event, the one or more systems may generate an alert requesting human assistance or guidance. The one or more systems may also analyze the detected intervention or detachment event and generate a natural language description of the event, classify the event as one or more categories and / or failure modes, generate a new entry in a training dataset based on the detected event, and generate a new entry in the training dataset based on the new entry. Or initiating automatic retraining of one or more autonomous or semi-autonomous control routines based on the detected event.
Need to check novelty before this filing date? Find Prior Art

Description

Background Technology

[0001] Organizations operating clusters of autonomous or semi-autonomous robotic systems (e.g., autonomous mobile robots, humanoid robots, forklifts, vehicles, boats, drones, etc.) typically encounter numerous intervention or disengagement events. Intervention or disengagement events can include situations requiring human or other (e.g., local and / or remote) operators to assist or guide robotic systems that are stuck or unable to continue operating in autonomous or semi-autonomous mode. Human operators may be present within the robotic system or may monitor its operation locally or remotely. Intervention or disengagement events provide valuable data points for analyzing the operation of robotic systems and for identifying necessary improvements in autonomous or semi-autonomous control routines. Data collected in association with intervention or disengagement events can also be suitable as supplementary training data when retraining or otherwise refining autonomous or semi-autonomous control routines.

[0002] Typically, the review and analysis of data associated with intervention or disengagement events is performed manually through manual classification. One or more human analysts must categorize events (e.g., by selecting from a pre-listed list of failure categories or failure modes). The human analyst must also manually generate natural language descriptions of the events and perform root cause analysis to determine the root causes of the intervention or disengagement events. Manual review is costly, time-consuming, susceptible to human error, and cannot scale to large volumes of event data. Furthermore, any insights gained from manual review are not automatically incorporated into updates or other refinements of autonomous or semi-autonomous control routines.

[0003] Therefore, more effective technologies are needed to analyze recorded data associated with intervention or disengagement events recorded by autonomous or semi-autonomous robotic systems, and to automatically improve autonomous or semi-autonomous control routines based on the recorded event data. Summary of the Invention

[0004] Embodiments of this disclosure relate to the active learning of autonomous robot intervention interpreters and machine learning models (e.g., base models). Systems and methods are disclosed for providing image data, other sensor data (e.g., LiDAR, RADAR, ultrasound, motor control sensors, inertial measurement units (IMUs), self-motion, etc.) to be collected and analyzed. Based on the image data or other sensor data, the disclosed systems and methods can perform online monitoring of a cluster of robotic systems, detect intervention or disengagement events, and generate alerts requesting assistance or guidance from a human or robot operator or monitor, for example, when the robot or machine experiencing the intervention or disengagement request is located at the same or a remote location. The disclosed systems and methods can also classify intervention or disengagement events into one of several categories and / or failure modes, generate natural language descriptions of the events, or select one or more event data items to include in an updated training or validation dataset. The disclosed systems and methods can further initiate automated retraining of one or more autonomous or semi-autonomous control routines based on the updated training or validation dataset. The disclosed systems and methods can also be used to identify specific characteristics of autonomous or semi-autonomous control routines for improvement based on the relative frequency of various categories of intervention or disengagement events or failure modes.

[0005] Compared to conventional methods, the systems and methods disclosed herein deploy trained (e.g., “basic”) machine learning models (MLMs) such as Visual Language Models (VLMs), Large Language Models (LLMs), or Multimodal Language Models (MMLMs) to monitor clusters of autonomous or semi-autonomous robotic systems. In online mode, the MLM receives sensor data associated with the operation of the robotic system and analyzes the sensor data to detect intervention or disengagement events. The MLM can generate alerts based on the detected events and transmit requests for assistance or intervention to a human operator or monitor. In offline mode, the MLM, or one or more additional MLMs, analyzes event data associated with intervention or disengagement events, generates natural language descriptions of the events, and assigns the events to clusters of similar intervention types. In offline mode, one or more MLMs can also identify one or more events to include in updated training or validation datasets and can automatically trigger retraining of one or more autonomous or semi-autonomous control routines based on the updated training or validation datasets. The disclosed techniques can also select a subset of event data for permanent storage, reducing storage requirements by discarding redundant event data. Attached Figure Description

[0006] The following detailed description, with reference to the accompanying drawings, describes the system and method for an autonomous robot intervention interpreter and active learning via a base model, wherein:

[0007] Figure 1This is a block diagram of an example computing device applicable to implementing some embodiments of the present disclosure;

[0008] Figure 2 According to various embodiments Figure 1 A more detailed diagram of the monitoring engine;

[0009] Figure 3 A flowchart of a method for monitoring a group of robot systems according to various embodiments is shown;

[0010] Figure 4 According to various embodiments Figure 1 A more detailed illustration of the inference engine;

[0011] Figure 5 A flowchart of a method for analyzing event data according to various embodiments is shown;

[0012] Figure 6A These are illustrations of exemplary autonomous vehicles according to some embodiments of the present disclosure;

[0013] Figure 6B According to some embodiments of this disclosure Figure 6A An example of camera position and field of view for an exemplary autonomous vehicle;

[0014] Figure 6C According to some embodiments of this disclosure Figure 6A A block diagram of an exemplary system architecture for an exemplary autonomous vehicle;

[0015] Figure 6D It is one or more cloud-based servers according to some embodiments of this disclosure and Figure 6A An exemplary system diagram of communication between autonomous vehicles;

[0016] Figure 7 This is a block diagram of an example computing device suitable for implementing some embodiments of the present disclosure; and

[0017] Figure 8 This is a block diagram of an example data center applicable to implementing some embodiments of this disclosure.

[0018] Figure 9A This is a block diagram of an example generative language model system suitable for implementing at least some embodiments of the present disclosure;

[0019] Figure 9B It is a block diagram of an example generative language model including a converter encoder-decoder, suitable for implementing at least some embodiments of this disclosure;

[0020] Figure 9CThis is a block diagram of an example generative language model, including a decoder-only converter architecture, suitable for implementing at least some embodiments of this disclosure. Detailed Implementation

[0021] Systems and methods related to autonomous robot intervention interpretation and active learning are disclosed, having a foundational model for autonomous and semi-autonomous systems and applications. Although this disclosure relates to exemplary autonomous or semi-autonomous robots, vehicles, or machines 600 (also referred to herein as "vehicle 600," "robot 600," "machine 600," or "self-machine 600"), examples are provided. Figures 6A to 6D The description herein is provided, but is not intended to be limiting. For example, the systems and methods described herein can be used by, but are not limited to, robotic systems, including non-autonomous vehicles or machines, semi-autonomous vehicles or machines (e.g., in one or more adaptive driver assistance systems (ADAS)), autonomous vehicles or machines, manned and unmanned robots, warehouse vehicles, off-road vehicles, vehicles coupled to one or more trailers, aircraft, boats, shuttles, emergency response vehicles, motorcycles, electric or motorized bicycles, aircraft, engineering vehicles, underwater vehicles, drones, and / or other vehicle types. Furthermore, although this disclosure may describe autonomous or semi-autonomous operation of robotic systems, it is not intended to be limiting, and the systems and methods described herein can be used in augmented reality, virtual reality, mixed reality, security and surveillance, autonomous or semi-autonomous machine applications, and / or any other technical field applicable to the detection and analysis of intervention or disengagement events. Further, although this disclosure primarily uses examples of sensors in the form of cameras and motor-controlled sensors for description, the disclosed techniques can be used with any suitable form of sensor (e.g., audio, LiDAR, RADAR, ultrasound, etc.).

[0022] As described in this paper, conventional techniques may require manual analysis or interpretation of intervention or disengagement events. Intervention or disengagement events can include situations where an autonomous or semi-autonomous machine has become stuck or otherwise requires human assistance or guidance to return to autonomous or semi-autonomous operation. Manual analysis may include classifying the event into one or more predefined event types. Interpretation may include generating a natural language description of the event or performing root cause analysis to determine one or more causes of the intervention or disengagement event. These techniques require costly and time-consuming human interaction and are subject to human variability. Furthermore, manual review and analysis techniques are insufficient for handling large volumes of intervention or disengagement event data. Furthermore, conventional manual review techniques do not allow for automated updates to test or validation datasets or automated retraining of autonomous or semi-autonomous control routines based on newly captured event data.

[0023] To improve the monitoring of clusters of autonomous or semi-autonomous robotic systems, the disclosed technology receives real-time sensor data from robots or other autonomous or semi-autonomous systems and analyzes the sensor data via one or more trained machine learning models to identify intervention or disengagement events. The disclosed technology can then generate alerts associated with the events and transmit requests for intervention from human or other (e.g., robot, computing, etc.) operators or monitors. In offline mode, the disclosed technology can analyze event data associated with intervention or disengagement events, including event data associated with human intervention or other actions taken to resolve the events. The disclosed technology can then classify events into one or more event types and / or failure modes and generate natural language descriptions of the events. The disclosed technology can selectively include event-associated data in updated test or validation datasets and automatically trigger retraining of one or more autonomous or semi-autonomous control routines based on the updated test and validation datasets. The disclosed technology can further select a subset of relevant event data for permanent or other long-term storage. Selecting a subset of event data for storage reduces computational and storage costs because the disclosed technology can selectively discard event data associated with common or redundant events.

[0024] In online mode, the monitoring engine receives real-time sensor data from the autonomous or semi-autonomous robotic system. For example, the robotic system may be one of multiple robotic systems included in a cluster operated by a single organization or enterprise. Sensor data may include, but is not limited to, image data, audio sensor data, LiDAR, RADAR, SONAR, or ultrasonic sensor data. Sensor data may also include motor control sensor data, such as the position or motion of motorized components included in the robotic system. Motor control sensor data may also include feedback data, such as measured position, motion, weight, stress, strain, or torque associated with components included in the robotic system. For example, one or more IMUs may be used in conjunction with wheels or other motion information to track the machine's self-movement over time and understand the robot system's attitude and position.

[0025] The monitoring engine includes (e.g., “foundation”) machine learning models such as Large Language Models (LLM), Visual Language Models (VLM), and / or Multimodal Language Models (MMLM). The machine learning models analyze real-time sensor data and detect intervention or disengagement events, where the robotic system has become stuck or otherwise requires human or other (e.g., robotic, computational, etc.) assistance or guidance to resume normal autonomous or semi-autonomous operation. Based on the detection of intervention or disengagement events, the monitoring engine can generate alerts and transmit them to the system, which can be operated by humans, robots, computer-based systems, and / or other types of operators or monitors. This operator or monitor can be located within the robotic system, locally within the robotic system, or geographically distant from the robotic system. The monitoring engine also records any sensor data received from the robotic system that is caused by human input required to resolve the intervention or disengagement event. For example, the monitoring engine can record motor control sensor data generated by human input to the robotic system. The monitoring engine can also record audio data, such as a human narration of the intervention or disengagement event or any corrective steps taken to resolve the event. The monitoring engine records any received sensor data associated with the event or human corrective action and generates entries in the event database for later offline analysis by the inference engine.

[0026] The monitoring engine can be implemented locally within an organization's enterprise computing environment. Alternatively, the monitoring engine can be implemented as a remote cloud service. When implemented as a remote cloud service, the monitoring engine can perform monitoring and inference as a microservice, which can be used by multiple different organizations, as described in more detail in this paper as the inference microservice.

[0027] In offline mode, the inference engine retrieves entries from an event database, which includes sensor data associated with a specific historical intervention or exit event, as well as sensor data associated with any human corrective actions taken to resolve that event. The machine learning model included in the inference engine can perform one or more actions in offline mode based on these event database entries. As described herein, the machine learning model can be the same as or similar to the underlying machine learning model implemented in the monitoring engine.

[0028] The base model can generate natural language descriptions of events associated with retrieved event database entries. It can also cluster events into one or more predefined intervention types and / or failure modes, such as navigation failures, sensor failures, obstacle avoidance failures, or motor control failures.

[0029] The base model can update the training and validation datasets based on retrieved event database entries. The base model can compare one or more existing entries included in the training and validation datasets with the retrieved event database entries and determine whether the retrieved entries represent unique or uncommon intervention or exit events. The inference engine can then update the training and validation datasets with unique or uncommon events, providing a wider variety of event data in the training and validation datasets.

[0030] The inference engine can trigger automatic retraining of the autonomous stack, which comprises one or more autonomous or semi-autonomous control routines for controlling a cluster of robotic systems. The inference engine can trigger automatic retraining upon any update to the training and validation datasets, upon performing a predetermined number of updates to the training and validation datasets, or based on a predetermined schedule.

[0031] The inference engine can evaluate received event database entries for either persistent or long-term storage. The base model compares a received event database entry to one or more previously stored events and selects or rejects entries for storage based on the similarity between the received entry and those previously stored events. If the inference engine determines that an entry is unique, uncommon, or underrepresented compared to previously stored events, it can select the entry for storage. Conversely, if the entry is an unnecessary duplication considering previously stored events, it can reject it. Selectively evaluating events for long-term storage reduces computational and storage requirements in an organization's computing environment.

[0032] Figure 1 This is a block diagram of an example computing device suitable for implementing some embodiments of the present disclosure. It should be understood that such and other arrangements described herein are merely illustrative examples. Other arrangements and elements (e.g., machines, interfaces, functions, sequences, functional groups, etc.) may be used in addition to or instead of the arrangements and elements shown, and some elements may be omitted entirely. Furthermore, many of the elements described herein are functional entities that can be implemented as discrete or distributed components or combined with other components, and in any suitable combination and location. The various functions described herein for performance by entities can be performed by hardware, firmware, and / or software. For example, various functions can be implemented by a processor executing instructions stored in memory. In some embodiments, the systems, methods, and processes described herein can use... Figures 6A to 6D Exemplary autonomous vehicles 600, Figure 7 Exemplary computing device 700, Figure 8 Exemplary data center 800 and / or Figures 9A to 9CMachine learning models use similar components, features, and / or functions to perform the task.

[0033] In one embodiment, computing device 100 includes a desktop computer, laptop computer, smartphone, personal digital assistant (PDA), tablet computer, in-vehicle or robotic computing device, or any other type of computing device configured to receive input, process data, and optionally display images, and is suitable for implementing one or more embodiments. Computing device 100 is configured to run a monitoring engine 122 and an inference engine 124 residing in memory 116.

[0034] It should be noted that the computing device described herein is illustrative, and any other technically feasible configuration falls within the scope of this disclosure. For example, multiple instances of monitoring engine 122 and inference engine 124 may execute on a set of nodes in a distributed and / or cloud computing system to implement the functionality of computing device 100. In another example, monitoring engine 122 and inference engine 124 may execute on various hardware, device types, or environments to adapt monitoring engine 122 or inference engine 124 to different use cases or applications. In a third example, monitoring engine 122 and inference engine 124 may execute on different computing devices and / or different sets of computing devices.

[0035] In one embodiment, computing device 100 includes, but is not limited to, interconnect (bus) 112 connecting one or more processors 102, input / output device interface 104 coupled to one or more input / output (I / O) devices 108, memory 116, storage device 114, and network interface 106. The one or more processors 102 can be any suitable processor implemented as a central processing unit (CPU), graphics processing unit (GPU), application-specific integrated circuit (ASIC), field-programmable gate array (FPGA), artificial intelligence (AI) accelerator, any other type of processing unit, or combination of different processing units, such as a CPU configured to operate in conjunction with a GPU. Generally, the one or more processors 102 can be any technically feasible hardware unit capable of processing data and / or executing software applications. Further, in the context of this disclosure, the computing elements shown in computing device 100 can correspond to a physical computing system (e.g., a system in a data center) or can be virtual computing instances executing within a computing cloud.

[0036] I / O device 108 includes devices capable of providing input (such as a keyboard, mouse, touchscreen, etc.) and devices capable of providing output (such as a display device). Additionally, I / O device 108 may include devices capable of both receiving input and providing output (such as a touchscreen, Universal Serial Bus (USB) port, etc.). I / O device 108 may be configured to receive various types of input from end users of computing device 100 (e.g., designers) and also to provide various types of output to end users of computing device 100, such as displayed digital images or digital video or text. In some embodiments, one or more I / O devices 108 are configured to couple computing device 100 to network 110.

[0037] Network 110 is any technically feasible type of communication network that allows data to be exchanged between computing device 100 and external entities or devices, such as a web server or another networked computing device. For example, network 110 may include a wide area network (WAN), a local area network (LAN), a wireless (WiFi) network, and / or the Internet.

[0038] Storage device 114 includes non-volatile memory for applications and data, and may include fixed disk drives or removable disk drives, flash memory devices, and CD-ROM, DVD-ROM, Blu-ray, HD-DVD, or other magnetic, optical, or solid-state storage devices. Monitoring engine 122 and inference engine 124 may be stored in storage device 114 and loaded into memory 116 during execution.

[0039] Memory 116 includes random access memory (RAM) modules, flash memory cells, or any other type of storage cell or combinations thereof. One or more processors 102, I / O device interfaces 104, and network interfaces 106 are configured to read data from and write data to memory 116. Memory 116 includes various software programs executable by one or more processors 102 and application data associated with those software programs, including a monitoring engine 122 and an inference engine 124.

[0040] Figure 2 According to various embodiments Figure 1A more detailed illustration of the monitoring engine 122 is provided. It should be understood that the arrangements and other arrangements described herein are merely illustrative examples. Other arrangements and elements (e.g., machines, interfaces, functions, sequences, functional groups, etc.) may be used in addition to or instead of the arrangements and elements shown, and some elements may be omitted entirely. Furthermore, many of the elements described herein are functional entities that can be implemented as discrete or distributed components or combined with other components, and in any suitable combination and location. The various functions described herein as being performed by entities can be performed by hardware, firmware, and / or software. For example, various functions can be implemented by a processor executing instructions stored in memory.

[0041] The monitoring engine 122 can receive sensor data 210, detect intervention or disengagement events associated with the autonomous or semi-autonomous robotic system based on the received sensor data, and generate alarms based on the detected intervention or disengagement events. The monitoring engine can also transmit the generated alarms to one or more human or other (e.g., computing systems, robots, etc.) operators and / or monitors. The monitoring engine 122 can also generate entries in the event database 240 based on the detected intervention or disengagement events and any human or other guidance or assistance associated with the detected events. Among other components, the monitoring engine 122 includes a base model 220 (or typically a machine learning model) and an alarm generator 230.

[0042] As a summary, training engine 122 can be configured to generate, process, preprocess, enhance, and / or otherwise prepare sensor data 210 for analysis by a trained base model 220. In embodiments that include an alarm generator 230, the alarm generator 230 can be configured to generate an alarm when it is determined that sensor data 210 indicates an intervention or disengagement event associated with an autonomous or semi-autonomous robotic system.

[0043] In various embodiments, sensor data 210 may include sensor data generated by the robotic system using any number of sensors (physical and / or virtual or simulated), such as one or more LiDAR sensors 664, one or more RADAR sensors 660, one or more ultrasonic sensors 662, one or more microphones 696, and / or other sensor types. Sensor data may represent the field of view and / or sensing field of the sensors (e.g., one or more LiDAR sensors 664, one or more RADAR sensors 660, etc.), and / or may represent the perception of the environment by one or more sensors (e.g., one or more microphones 696). Sensors such as image sensors (e.g., cameras), LiDAR sensors, RADAR sensors, SONAR sensors, ultrasonic sensors, etc., may be referred to herein as sensing sensors or sensing sensor devices, and sensor data generated by sensing sensors may also be referred to herein as sensing sensor data. In some examples, instances or representations of sensor data may be represented by images captured by image sensors (e.g., image data), depth maps generated by LiDAR sensors, etc. Audio data, LiDAR data, SONAR data, RADAR data, and / or other sensor data types can be related to or associated with image data generated using one or more image sensors. For example, image data representing one or more images can be updated to include data related to LiDAR sensors, SONAR sensors, RADAR sensors, etc., so that the sensor data used as input to the base model 220 can be more informative or detailed than individual image data.

[0044] In embodiments using sensor data 210, the sensor can be calibrated such that the sensor data is associated with pixel coordinates in the image data. In some embodiments, such as where the sensor data indicates depth (e.g., RADAR data, LiDAR data, etc.), depth values ​​can be associated with pixel coordinates in the image data and then used as additional (or alternative, in some examples,) inputs to the base model 220. For example, one or more pixels can have associated additional values ​​representing depth, as determined from the sensor data.

[0045] In various embodiments, the base model 220 includes a trained machine learning model, including but not limited to a visual language model (VLM), a multimodal language model (MMLM), or the one described herein. Figures 9A to 9C The description refers to any machine learning model discussed therein. The base model 220 analyzes sensor data 210 and detects intervention or disengagement events associated with the autonomous or semi-autonomous robotic system. (As described in this paper...) Figure 4As discussed in the description, the base model 220 can also analyze sensor data associated with intervention or disengagement events to generate natural language descriptions of the events, classify the events into one or more predefined categories and / or failure modes, generate new entries in the training dataset, or initiate automatic retraining of an autonomous stack that includes one or more autonomous or semi-autonomous control routines.

[0046] In various embodiments, the base model 220 can be trained "from scratch" to detect intervention or disengagement events based on domain-specific training data generated by an organization from the historical operation of one or more robotic systems included in a cluster of robotic systems. For example, domain-specific training data comprising thousands or tens of thousands of hours of annotated historical sensor data may be sufficient to fully train the base model 220 to detect and analyze intervention or disengagement events in autonomous or semi-autonomous robotic systems.

[0047] In other embodiments, base model 220 may include a pre-trained machine learning model, which is then further trained or refined with relatively little annotated domain-specific training data. For example, domain-specific training data including tens or hundreds of hours of annotated historical sensor data may be sufficient to refine the pre-trained machine learning model included in base model 220, enabling the refined machine learning model to analyze sensor data 210 and detect and analyze intervention or disengagement events in autonomous or semi-autonomous robotic systems.

[0048] In embodiments including alarm generator 230, alarm generator 230 may be configured to generate an alarm when base model 220 determines that sensor data 210 indicates an intervention or disengagement event associated with the robotic system. The generated alarm may include a subset of sensor data 210, such as one or more images and / or audio recordings. The generated alarm may also include an identifier associated with a particular robotic system and / or a location associated with that robotic system.

[0049] The monitoring engine 122 can transmit generated alarms to one or more human operators and / or monitors. These operators and / or monitors may be paired with a robotic system, such as a human operator of an autonomous or semi-autonomous forklift. They may also include remote operators located at approximately the same location as or away from the robotic system. The generated alarms may prompt the operators and / or monitors to exercise local or remote control over the robotic system and provide guidance or assistance to return the robotic system to its normal autonomous or semi-autonomous operating mode. The monitoring engine 122 can transmit the generated alarms via any suitable means, including but not limited to email, text messages, organizational messaging applications, or a dedicated receiver paired with an operator or monitor. In various embodiments, the dedicated receiver is operable to generate audible and / or visual alarms, as well as graphical and / or text displays associated with the alarms.

[0050] In embodiments that include event database 240, monitoring engine 122 may generate entries in event database 240 associated with detected intervention or disengagement events. The generated entries may include sensor data 210, which, when processed by base model 220, causes base model 220 to detect the event. The generated entries may also include sensor data 210 associated with assistance or guidance provided to the robotic system by a human operator or monitor in response to the detected event. In various embodiments, sensor data 210 associated with human assistance or guidance may include sensing data associated with motor controllers included in the robotic system. Assistive human control inputs provided to the robotic system may be recorded as position, motion, orientation, stress, strain, or torque manifested or experienced by one or more motor control components included in the robotic system and included in sensor data 210. Sensor data 210 may also include audio recordings, such as a description of the intervention or disengagement event by a human operator or monitor, or a narration of assistive actions taken to return the robotic system to normal autonomous or semi-autonomous operation. The monitoring engine 122 transmits the event database 240 to the inference engine 124 described herein.

[0051] While this paper describes examples of using language models, and particularly visual language models or multimodal language models as base models 220, this is not intended to be limiting. For example, but not limited to, the base model 220 described herein may include one or more of any type of machine learning model, such as using linear regression, logistic regression, decision trees, support vector machines (SVM), Naive Bayes, k-nearest neighbors (Knn), k-means clustering, random forests, dimensionality reduction algorithms, gradient boosting algorithms, neural networks (e.g., autoencoders, convolutions, recurrent neural networks, perceptrons, long / short-term memory (LSTM), Hopfield, Boltzmann, deep belief, deconvolution, generative adversarial, liquid machines, etc.) and / or other types of machine learning models.

[0052] As an example and not a limitation, the monitoring engine 122 is described with respect to one or more MLMs trained for computer vision and / or perception operations to monitor the operation of a robotic system. However, aspects of this disclosure are more broadly applicable to any form of MLM trained and / or deployed to make predictions based on sensor data. In some examples, the base model 220 may be trained to predict trajectory points, vehicle orientation (e.g., regarding environmental features such as lane markings) and / or vehicle states (e.g., regarding object manipulation such as lane changes, turning, merging, etc.), which can be used to control an autonomous vehicle. However, this is not intended to be limiting.

[0053] Additionally, monitoring engine 122 is an example of an engine that can be used in at least one embodiment, such as for performing MLM for computer vision and / or perception operations to navigate a vehicle, or for other purposes. However, monitoring engine 122 can vary to include more than Figure 2 The more, fewer, and / or different components and / or processing paths shown.

[0054] The monitoring engine 122 can be implemented in a cloud computing environment and is available to one or more clients as a monitoring microservice for an autonomous or semi-autonomous robot system cluster. In various embodiments where the monitoring engine executes as a monitoring microservice, each of one or more instances of the base model 220 can be trained or refined based on historical cluster operation data from different clients. In other embodiments, one or more instances of the base model can be trained or refined on the same historical cluster operation data.

[0055] Now for reference Figure 3 and Figure 5Each operation of methods 300 and 500 described herein includes a computational process that can be performed using any combination of hardware, firmware, and / or software. For example, various functions can be performed by a processor executing instructions stored in memory. These methods can also be embodied as computer-usable instructions stored on a computer storage medium. These methods can be provided by standalone applications, services, or managed services (standalone or in combination with another managed service) or plug-ins to another product, to name a few. Furthermore, methods 300 and 500, by way of example, relate to... Figures 1 to 2 and Figure 4 The system described herein. However, these methods may be additionally or alternatively implemented by any system or any combination of systems, including but not limited to the systems described herein.

[0056] Figure 3 Flowcharts of methods for detecting intervention or disengagement events according to various embodiments are shown. Figure 3 As shown, method 300 begins with operation 302, where monitoring engine 122 receives sensor data 210 from the robot system, such as those described herein. Figures 6A to 6C The vehicle discussed in the description. In various embodiments, sensor data 210 includes, but is not limited to, image data, video data, audio recordings, and data from one or more LiDAR, RADAR, SONAR, ultrasonic, or infrared sensors.

[0057] At operation 304, method 300 includes detecting an intervention or disengagement event associated with the robotic system via a trained base model 220 and based on sensor data 210, wherein the robotic system has become stuck or otherwise requires human intervention to return to a normal autonomous or semi-autonomous operating mode. In various embodiments, the base model 220 includes a trained machine learning model, including but not limited to a visual language model, a multimodal language model, or otherwise described herein. Figures 9A to 9C Any machine learning model discussed in the description.

[0058] In various embodiments, the base model 220 may be fully trained on a relatively large, annotated, domain-specific training dataset generated by the organization from historical operations of one or more robotic systems included in a cluster of robotic systems. In other embodiments, the base model 220 may include a pre-trained machine learning model, which is then further trained or refined using relatively little annotated, domain-specific training data.

[0059] At operation 306, method 300 includes generating and transmitting an alarm based on the detected intervention or disengagement event. The generated alarm indicates that human intervention or assistance is required to return the robotic system to a normal autonomous or semi-autonomous operating mode. Monitoring engine 122 may transmit the generated alarm to a human operator associated with the robotic system. Alternatively or additionally, monitoring engine 122 may transmit the generated alarm to a human monitor located in the same geographical area as the robotic system or at a distance from the robotic system.

[0060] At operation 308, method 300 includes receiving additional sensor data associated with human intervention or assistance provided to the robot system in response to a generated alarm. In various embodiments, the additional sensor data 210 associated with human assistance or guidance may include sensor data associated with motor controllers included in the robot system. Assistive human control inputs provided to the robot system may be recorded as position, motion, orientation, stress, strain, or torque manifested or experienced by one or more motor control components included in the robot system and included in the sensor data 210. The sensor data 210 may also include audio recordings, such as a description of the intervention or disengagement event by a human operator or monitor, or a narration of assistive actions taken to return the robot system to normal autonomous or semi-autonomous operation.

[0061] At operation 310, method 300 includes generating entries in event database 240 based on detected intervention or disengagement events, sensor data 210, and additional sensor data 210. The generated entries may include sensor data 210, which, when processed by the base model 220, causes the base model 220 to detect intervention or disengagement events. The generated entries may also include additional sensor data 210 associated with assistance or guidance provided to the robotic system by a human operator or monitor in response to the detected events. Monitoring engine 122 transmits event database 240 to inference engine 124, which is discussed herein.

[0062] Figure 4 According to various embodiments Figure 1A more detailed illustration of the inference engine 124 is provided. It should be understood that the arrangements and other arrangements described herein are merely illustrative examples. Other arrangements and elements (e.g., machines, interfaces, functions, sequences, functional groupings, etc.) may be used in addition to or instead of the arrangements and elements shown, and some elements may be omitted entirely. Furthermore, many of the elements described herein are functional entities that can be implemented as discrete or distributed components or combined with other components, and in any suitable combination and location. The various functions described herein as being performed by entities can be performed by hardware, firmware, and / or software. For example, various functions can be implemented by a processor executing instructions stored in memory.

[0063] In summary, inference engine 124 may receive entries from event database 240 and perform one or more actions, including but not limited to generating natural language descriptions of intervention or disengagement events, classifying intervention or disengagement events into one or more categories and / or failure modes, and selecting entries received from event database 24 for permanent or long-term storage in storage device 114. Inference engine 124 may also select one or more event data items to include in training dataset 400. Inference engine 124 may also initiate automatic retraining of an autonomous stack comprising one or more autonomous or semi-autonomous control routines based on an updated training or validation dataset. Inference engine 124 may also identify specific characteristics of autonomous or semi-autonomous control routines for improvement based on the relative frequency of various categories and / or failure modes of intervention or disengagement events. Among other elements, inference engine 124 may include base model 220, natural language generator 410, event classifier 420, training data generator 430, autonomous stack trainer 440, and event selection module 450. Each of the natural language generator 410, event classifier 420, training data generator 430, and event selection module 450 can represent different input / output mechanisms included in the base model 220.

[0064] Inference engine 124 receives event entries included in event database 240 and transmitted from monitoring engine 122. As described herein, entries included in event database 240 may include sensor data associated with an intervention or disengagement event detected by monitoring engine 122. In embodiments including natural language generator 410, inference engine 124 may generate a natural language description of the intervention or disengagement event associated with the received event entries via base model 220 and natural language generator 410. Natural language generator 410 may provide text prompts to base model 220, instructing base model 220 to generate a natural language description based on sensor data included in the received entries, which is associated with the intervention or disengagement event itself, or with the sensor data of the received entries and with human assistance input provided to the robotic system in response to the intervention or disengagement event. Natural language generator 410 may receive natural language descriptions from base model 220 and transmit natural language descriptions to inference engine 124.

[0065] In embodiments including event classifier 420, inference engine 124 can classify event entries into one or more predefined categories of intervention type and / or fault mode via base model 220 and event classifier 420. In various embodiments, one or more predefined categories of intervention type may include, but are not limited to, navigation fault, sensor fault, obstacle avoidance fault, or motor control fault. Event classifier 420 can generate text prompts instructing base model 220 to classify event entries and transfer classifications from base model 220 to inference engine 124. In various embodiments, event classifier 420 can also generate prompts instructing base model 220 to determine a fault mode associated with an event entry, where the fault mode provides more specific characteristics than a general category of intervention type. For example, base model 220 may classify an event entry as a navigation fault and further determine a fault mode specifying a fault or degradation of a particular sensor included in the robotic system.

[0066] In embodiments including a training data generator 430, the inference engine 124 may generate new entries in the training dataset 400 via the training data generator 430. The training dataset 400 includes one or more entries, each including sensor data associated with an intervention or disengagement event and sensor data associated with human assistance and / or guidance input provided in response to the intervention or disengagement event. Entries included in the training dataset 400 include test and validation data for training an autonomous stack, which includes one or more autonomous or semi-autonomous control routines for controlling a cluster of robotic systems. The training data generator 430 calculates the similarity between event entries received from the event database 240 and one or more entries included in the training dataset 400. Based on the calculated similarity, the training data generator 430 determines whether a received event entry is unusual or uncommon compared to one or more entries included in the training dataset 400. Based on this determination, the training data generator 430 may prompt the base model 220 to generate and format new entries in the training dataset 400 based on the received event entries. The inference engine 124 may record new entries in the training dataset 400. If the training data generator 430 determines, based on one or more existing entries included in the training dataset 400, that the received event entry will be duplicated or redundant, the training data generator 430 may stop analyzing the received event entry.

[0067] In embodiments including an autonomous stack trainer 440, inference engine 124 may initiate automatic retraining of one or more autonomous or semi-autonomous control routines included in the autonomous stack via autonomous stack trainer 440 based on one or more new entries generated by training data generator 430 and added to training dataset 400. In various embodiments, inference engine 124 may initiate automatic retraining of the autonomous stack immediately after new entries are generated in training dataset 400. In other embodiments, inference engine 124 may initiate automatic retraining when a predetermined number of new entries are generated in training dataset 400, or according to a predetermined schedule.

[0068] In embodiments including the event selection module 450, the event selection module 450 may be configured to record event entries received from the event database 240 in a permanent storage device or other long-term storage device, such as storage device 114. The event selection module 450 may compare event entries with existing entries included in storage device 114 based on sensor data 210 included in the event entries, generated natural language descriptions generated by the natural language generator 410 and associated with the event entries, and / or categories to which the event entries have been classified by the event classifier 420. Based on this comparison, if the event entry and its associated natural language description and / or classification are unique, uncommon, or otherwise not adequately represented in storage device 114, the event selection module 450 may select the event entry to record in storage device 114. Alternatively, if the event entry is duplicated or redundant given existing event entries included in storage device 114, the event selection module 450 may reject the event entry for permanent storage or other long-term storage. Evaluating individual event entries for permanent storage or other long-term storage can reduce computational and storage requirements compared to saving all received event entries to storage device 114.

[0069] Inference engine 124 can analyze multiple event entries over time and, based on these entries, suggest specific areas for improvement for autonomous or semi-autonomous control routines included in the autonomy stack. For example, inference engine 124 can determine a threshold percentage of event entries that have been classified as navigation errors by base model 220 and event classifier 420, and can suggest examining one or more navigation routines included in the autonomy stack. In various embodiments, inference engine 124 can also suggest reviewing one or more control or navigation routines or robot / vehicle components based on specific failure modes determined by base model 220, rather than solely on event classification categories. For example, inference engine 124 can suggest reviewing both navigation routines and robot system navigation sensors based on a threshold number of event entries that have been classified as navigation failures with corresponding specific failure modes indicating malfunctions or degradation in one or more navigation sensors (such as LiDAR or RADAR sensors) included in the robot system.

[0070] Figure 5 Flowcharts of methods for analyzing event data according to various embodiments are shown. Figure 5As shown, method 500 begins with operation 502, where inference engine 124 receives event entries from event database 240. These event entries are associated with intervention or disengagement events related to the robot system and are detected by monitoring engine 122. Event entries include sensor data 210 associated with intervention or disengagement events, as well as sensor data 210 associated with human assistance or guidance input provided to the robot system to return it to normal autonomous or semi-autonomous operation.

[0071] At operation 504, method 500 includes generating a natural language description of an intervention or disengagement event associated with the event entry. In various embodiments, the natural language generator 410 may provide text prompts to the base model 220, instructing the base model 220 to generate a natural language description based on sensor data included in the received entry and associated with the intervention or disengagement event itself, or sensor data included in the received entry and associated with human-assisted input provided to the robotic system in response to the intervention or disengagement event. The natural language generator 410 may receive the natural language description from the base model 220 and transmit the natural language description to the inference engine 124.

[0072] At operation 506, method 500 includes classifying intervention or detachment events associated with the event entry into one or more categories and / or determining a failure mode associated with the event entry. Event classifier 420 may generate text prompts instructing base model 220 to classify the event entry into one or more predefined intervention type categories. In various embodiments, one or more predefined intervention type categories may include, but are not limited to, navigation failure, sensor failure, obstacle avoidance failure, or motor control failure. Event classifier 420 may transfer classifications from base model 220 to inference engine 124. In various embodiments, event classifier 420 may also generate prompts instructing base model 220 to determine a failure mode associated with the event entry, where the failure mode provides more specific characteristics than a general category of intervention type. For example, base model 220 may classify the event entry as a navigation failure and further determine a failure mode that specifies a failure or degradation of a LiDAR or RADAR sensor included in the robotic system.

[0073] At operation 508, method 500 includes evaluating the event entries to include them as new entries in training dataset 400. Entries included in training dataset 400 include test and validation data for training an autonomous stack, which comprises one or more autonomous or semi-autonomous control routines for controlling a cluster of robotic systems. Training data generator 430 calculates the similarity between event entries received from event database 240 and one or more entries included in training dataset 400. Based on the calculated similarity, training data generator 430 determines whether the received event entries are unusual or uncommon compared to one or more entries included in training dataset 400. Based on this determination, training data generator 430 may prompt base model 220 to generate and format new entries in training dataset 400 based on the received event entries. Inference engine 124 may record the new entries in training dataset 400. If the training data generator 430 determines, based on one or more existing entries included in the training dataset 400, that the received event entry will be duplicated or redundant, the training data generator 430 may stop analyzing the received event entry.

[0074] At operation 510, method 500 includes initiating automatic retraining of the autonomous stack based on new entries in the training dataset 400. Inference engine 124 may initiate automatic retraining of one or more autonomous or semi-autonomous control routines included in the autonomous stack via autonomous stack trainer 440 based on one or more new entries generated by training data generator 430 and added to training dataset 400. In various embodiments, inference engine 124 may initiate automatic retraining of the autonomous stack immediately after new entries are generated in training dataset 400. In other embodiments, inference engine 124 may initiate automatic retraining when a predetermined number of new entries are generated in training dataset 400, or according to a predetermined schedule.

[0075] At operation 512, method 500 includes evaluating event entries for recording in permanent storage or other long-term storage (such as storage device 114). Event selection module 450 may compare event entries with existing entries included in storage device 114 based on sensor data 210 included in the event entries, generated natural language descriptions generated by natural language generator 410 and associated with the event entries, and / or categories to which event entries have been classified by event classifier 420. Based on this comparison, if the event entry and its associated natural language description and / or classification are unique, uncommon, or otherwise not adequately represented in storage device 114, event selection module 450 may select the event entry for recording in storage device 114. Alternatively, if the event entry is duplicated or redundant given existing event entries included in storage device 114, event selection module 450 may not select the event entry for permanent storage or other long-term storage.

[0076] The systems and methods described herein can be used for a variety of purposes, such as, but not limited to, machine (e.g., robots, vehicles, construction machinery, warehouse vehicles / machines, autonomous, semi-autonomous and / or other machine types) control, machine motion, machine driving, synthetic data generation, model training (e.g., using real data, augmented data and / or synthetic data, such as synthetic data generated using simulation platforms or systems, synthetic data generation techniques, such as, but not limited to, those described herein), perception, augmented reality (AR), virtual reality (VR), mixed reality (MR), robotics, security and surveillance (e.g., in smart city implementations), autonomous or semi-autonomous machine applications, deep learning, environmental simulation, object or participant simulation and / or digital twins, data center processing, conversational AI, optical transport simulation (e.g., ray tracing, path tracing, etc.), distributed or collaborative content creation for 3D assets (e.g., using generic scene descriptor (USD) data, such as OpenUSD and / or other data types), cloud computing, generative artificial intelligence (e.g., using one or more diffusion models, transformer models, etc.), and / or any other suitable application.

[0077] The disclosed embodiments can be included in a variety of different systems, such as automotive systems (e.g., control systems for autonomous or semi-autonomous machines, perception systems for autonomous or semi-autonomous machines), systems implemented using robots or robotic platforms, aviation systems, medical systems, boating systems, smart area monitoring systems, systems for performing deep learning operations, systems for performing simulation operations (e.g., in driving or vehicle simulations, robot simulations, smart city or surveillance simulations, etc.), systems for performing digital twin operations (e.g., in conjunction with collaborative content creation platforms or systems, such as, but not limited to, NVIDIA's OMNIVERSE and / or another platform, system, or service using USD or OpenUSD data types), systems implemented using edge devices, systems containing one or more virtual machines (VMs), and systems for performing deep learning operations. Systems that perform synthetic data generation operations (e.g., using one or more neural rendering fields (NERF), Gaussian scattering techniques, diffusion models, transformer models, etc.), systems implemented at least partially in a data center, systems for performing conversational AI operations, systems that implement one or more language models (such as one or more large language models (LLM), one or more visual language models (VLM), one or more multimodal language models, etc.), systems for performing optical transport simulations, systems for performing collaborative content creation for 3D assets (e.g., using generic scene descriptor (USD) data, such as OpenUSD, computer-aided design (CAD) data, 2D and / or 3D graphics or design data and / or other data types), systems implemented at least partially using cloud computing resources, and / or other types of systems.

[0078] In some embodiments, the systems and methods described herein can be executed within a simulated environment (e.g., NVIDIA DriveSIM, NVIDIA ISAAC GYM, NVIDIA ISAAC SIM, etc.) using simulated data (e.g., simulated sensor data from simulated sensors of virtual or simulated machines). For example, simulated sensor data (e.g., processed using one or more machine learning models, neural networks, etc.) can be used to identify, detect, and / or classify lane lines, obstacles, navigation paths, road boundary lines, other lines, vertical structures / features, etc., using points of curves and / or one or more curve fitting algorithms in the simulated environment, and this information can be used to perform operations associated with virtual machines in the environment (e.g., control, navigation, obstacle avoidance, planning, etc.). These simulated operations can be used to test their performance before deploying the underlying algorithms, systems, and / or processes to the real world. In some cases, simulation can be used to generate synthetic training data, for example, training data including regions of interest and / or subregions of interest from the simulation. In some embodiments, additionally or alternatively, other methods besides simulation can be used to generate synthetic training data. For example, neural rendering fields (NERF), Gaussian scattering techniques, diffusion models, electrostatic models (e.g., Poisson flow generation models (PFGM), etc.) can be used to generate synthetic training data. The synthetic training data (additionally or alternatively derived from real-world data) can then be processed to determine geometry, curvature, semantic information, classification information, and / or other information related to features of interest, such as lines, obstacles, paths, longitudinal features (e.g., poles), and / or other features in a driving environment, warehouse, etc. In any example, such as in a simulated environment used for testing, validation, training, etc., one or more optical propagation algorithms (such as ray tracing and / or path tracing algorithms) can be used to render or otherwise generate the simulated environment and / or associated training data. In some embodiments, the simulated environment and / or one or more of its objects, features, and components can be generated or managed within a 3D content collaboration platform (e.g., NVIDIA's OMNIVERSE) for industrial digitization, generative physics AI, and / or other use cases, applications, or services. For example, a content collaboration platform or system may include a system that uses generic scene descriptors (USD) (e.g., OpenUSD) data to manage objects, features, scenes, etc., in simulated environments, digital environments, etc. The platform may include realistic physics simulations, such as using NVIDIA's PhysX SDK, to simulate real physics and physical interactions with simulations hosted on the platform.This platform can integrate OpenUSD with ray tracing / path tracing / light transport simulations (e.g., NVIDIA's RTX rendering technology) into software tools and simulation workflows for building, training, deploying, or testing AI systems, such as systems for testing, validating, training (e.g., machine learning models, neural networks, etc.) and / or other tasks related to automobiles, robots, machines, or other applications.

[0079] In some embodiments, a remote control or teleoperation system can be used to perform remote operation or remote control of a vehicle or other machine. For example, the systems and methods described herein can be used to identify lane lines, road boundary lines, obstacles, paths, longitudinal features, etc., which can be included in a visualization or mapping of the environment to assist remote operator control or provide waypoints or other control or navigation instructions to enable autonomous or semi-autonomous machines to navigate the environment.

[0080] In some embodiments, the systems and methods described herein can be deployed in robotic applications. For example, a robot or robotic system may include one or more onboard processors (e.g., CPU, GPU, hardware-based deep learning accelerator (DLA), hardware-based programmable vision accelerator (PVA), which may include one or more vector processing units (VPU), direct memory access (DMA) systems and / or pixel processing engines (PPE), hardware-based optical flow accelerators (OFA), SoCs, etc.) and memory and / or storage devices (e.g., for storing control algorithms, sensor data, and one or more machine learning models). The robotic system can use these processors to execute one or more machine learning models (e.g., language models) that allow it to autonomously or semi-autonomously perform complex tasks, such as interacting with and / or manipulating static and / or dynamic objects, or navigating its environment using sensors such as cameras, LiDAR, RADAR, and ultrasonic sensors. The system can use sensor fusion techniques to combine data from multiple sensors (e.g., cameras, infrared, LiDAR, RADAR, accelerometers) to create a comprehensive model of the robot's surrounding environment. This data can be processed locally on the robot or sent to a remote server for more computationally intensive tasks, such as 3D mapping or SLAM (Simultaneous Localization and Mapping). In one or more embodiments, data from individual robots (e.g., sensor data, task status, or environmental conditions) can be uploaded to the cloud, where a centralized AI model can analyze and distribute optimized commands across the cluster. In some embodiments, one or more machine learning models described herein (e.g., language models, VLM, LLM, MMLM, diffusion models, NeRF models, DNN, etc.) can be used to allow the robot to perceive and reason about its environment and / or communicate with one or more other robots and / or humans in the environment. In some embodiments, the robot can communicate with one or more locally hosted servers / computing devices and / or with one or more remotely located servers / computing devices (e.g., in one or more data centers) using one or more network interface cards (NICs) and / or data processing units (DPUs).

[0081] In some examples, one or more machine learning models described herein (e.g., deep neural networks, language models, LLMs, VLMs, multimodal language models, perceptual models, tracking models, fusion models, transformer models, diffusion models, encoder-only models, decoder-only models, encoder-decoder models, neural rendering field (NERF) models, etc.) can be packaged into microservices (such as inference microservices (e.g., NVIDIA NIM)), which may include containers (e.g., operating system (OS) level virtualization packages) that may include an application programming interface (API) layer, a server layer, a runtime layer, and / or at least one model "engine". For example, an inference microservice may include the container itself and one or more models (e.g., weights and biases). In some cases, such as when one or more machine learning models are small enough (e.g., have a sufficiently small number of parameters), the one or more models may be included within the container itself. In other examples, such as when one or more models are large, the one or more models may be hosted / stored in the cloud (e.g., in a data center) and / or may be hosted on-premises and / or at the edge (e.g., on a local server or computing device, but outside the container). In this embodiment, one or more models can be accessed via one or more APIs (such as REST APIs). Therefore, and in some embodiments, the one or more machine learning models described herein can be deployed as inference microservices to accelerate the deployment of one or more models on any cloud, data center, or edge computing system while ensuring data security. For example, an inference microservice may include one or more APIs, pre-configured containers for simplified deployment, optimized inference engines (e.g., execution software built using standardized AI model deployments, such as NVIDIA's Triton Inference Server, and / or one or more APIs for high-performance deep learning inference, which may include inference runtimes and model optimizations for low latency and high throughput for production applications, such as NVIDIA's TensorRT), and / or enterprise management data for telemetry (e.g., including identity, metrics, health checks, and / or monitoring). The one or more machine learning models described herein can be used as part of an acceleration infrastructure with the ability to deploy using a single command and / or orchestrate and autoscale on the acceleration infrastructure using a container orchestration system (e.g., reaching data center scale on a single device). Therefore, inference microservices may include one or more machine learning models (e.g., models that have been optimized for high-performance inference), inference runtime software for executing one or more machine learning models and providing output / response to inputs (e.g., user queries, prompts, etc.), and enterprise management software for providing health checks, identity and / or other monitoring.In some embodiments, the inference microservice may include software for performing in-situ replacements and / or updates to one or more machine learning models. During replacement or update, the software performing the replacement / update may maintain user configurations for the inference runtime software and enterprise management software.

[0082] Example autonomous vehicles

[0083] Figure 6A This is an illustration of an exemplary autonomous or semi-autonomous vehicle 600 according to some embodiments of this disclosure. The autonomous vehicle 600 (or referred to herein as “vehicle 600”) may include, but is not limited to, passenger vehicles such as cars, trucks, buses, first-response vehicles, shuttles, electric or motorized bicycles, motorcycles, fire trucks, police cars, ambulances, boats, construction vehicles, underwater vehicles, robotic vehicles, drones, aircraft, vehicles coupled to trailers (e.g., semi-tractor-trailer trucks for hauling goods) and / or other types of vehicles (e.g., driverless and / or vehicles accommodating one or more passengers). Autonomous vehicles are typically described according to their level of automation, as defined by the National Highway Traffic Safety Administration (NHTSA) of the U.S. Department of Transportation, and by the Society of Automotive Engineers (SAE) in its "Classification and Definition of Terms Related to Driving Automation Systems for Road Motor Vehicles" (Standard No.: J3016-201806, published June 15, 2018; Standard No.: J3016-201609, published September 30, 2016; and previous and future versions of this standard). Vehicle 600 may implement functions according to one or more of Level 3 to Level 5 of autonomous driving. Vehicle 600 may be able to have one or more functionalities that conform to Level 1 to Level 5 of autonomous driving. For example, according to an embodiment, vehicle 600 may have driver assistance (Level 1), partial automation (Level 2), conditional automation (Level 3), high automation (Level 4), and / or full automation (Level 5). As used herein, the term “autonomy” can include any and / or all types of autonomy of the vehicle 600 or other machines, such as full autonomy, high autonomy, conditional autonomy, partial autonomy, provision of auxiliary autonomy, semi-autonomy, primary autonomy or other specified autonomy.

[0084] Vehicle 600 may include a chassis, body, wheels (e.g., 2, 4, 6, 8, 18, etc.), tires, axles, and other vehicle components. Vehicle 600 may include a propulsion system 650, such as an internal combustion engine, a hybrid power plant, an all-electric motor, and / or other propulsion system types. Propulsion system 650 may be connected to the drivetrain of vehicle 600, which may include a transmission, to enable propulsion of vehicle 600. Propulsion system 650 may be controlled in response to receiving signals from throttle valve / accelerator 652.

[0085] When the propulsion system 650 is in operation (e.g., when the vehicle is moving), the steering system 654, including a steering wheel, can be used to guide the vehicle 600 (e.g., along a desired path or route). The steering system 654 can receive signals from the steering actuator 656. For fully automatic (level 5) functionality, the steering wheel may be optional.

[0086] The brake sensor system 646 can be used to operate the vehicle brakes in response to signals received from the brake actuator 648 and / or the brake sensor.

[0087] One or more controllers 636 may include one or more system-on-chip (SoC) 604 ( Figure 6C One or more controllers 636 and / or one or more GPUs may provide signals (e.g., representing commands) to one or more components and / or systems of vehicle 600. For example, one or more controllers 636 may send signals to operate vehicle brakes via one or more brake actuators 648, operate steering system 654 via one or more steering actuators 656, and operate propulsion system 650 via one or more throttles / accelerators 652. One or more controllers 636 may include one or more onboard (e.g., integrated) computing devices (e.g., supercomputers) that process sensor signals and output operating commands (e.g., signals representing commands) to enable autonomous driving and / or assist a human driver in driving vehicle 600. One or more controllers 636 may include a first controller 636 for autonomous driving functions, a second controller 636 for functional safety functions, a third controller 636 for artificial intelligence functions (e.g., computer vision), a fourth controller 636 for infotainment functions, a fifth controller 636 for redundancy in emergency situations, and / or other controllers. In some examples, a single controller 636 may handle two or more of the functions described above, and two or more controllers 636 may handle a single function and / or any combination thereof.

[0088] One or more controllers 636 may provide signals for controlling one or more components and / or systems of vehicle 600 in response to sensor data (e.g., sensor input) received from one or more sensors. Sensor data may be received from, for example, but not limited to, one or more Global Navigation Satellite System (“GNSS”) sensors 658 (e.g., one or more Global Positioning System sensors), one or more radar sensors 660, one or more ultrasonic sensors 662, one or more lidar sensors 664, one or more inertial measurement unit (IMU) sensors 666 (e.g., one or more accelerometers, one or more gyroscopes, one or more magnetometers, magnetometers, etc.), one or more microphones 696, one or more stereo cameras 668, one or more wide-angle cameras 670 (e.g., fisheye cameras), one or more infrared cameras 672, one or more surround cameras 674 (e.g., 360-degree cameras), one or more long-range and / or mid-range cameras 698, one or more speed sensors 644 (e.g., for measuring the speed of vehicle 600), one or more vibration sensors 642, one or more steering sensors 640, one or more braking sensors (e.g., as part of braking sensor system 646) and / or other sensor types.

[0089] One or more of the controllers 636 may receive inputs (e.g., represented by input data) from the instrument cluster 632 of the vehicle 600 and provide outputs (e.g., represented by output data, displayed data, etc.) via a human-machine interface (HMI) display 634, an audio signaler, a speaker, etc., and / or via other components of the vehicle 600. Outputs may include information such as vehicle speed, rate, time, map data (e.g., ...). Figure 6C Information such as a high-definition (“HD”) map 622, location data (e.g., the location of vehicle 600, such as its location on a map), direction, the location of other vehicles (e.g., grid occupancy), and information about objects and their states perceived by one or more controllers 636. For example, the HMI display 634 may display information about the presence of one or more objects (e.g., street signs, warning signs, traffic light changes, etc.), and / or information about driving actions that the vehicle has performed, is performing, or will perform (e.g., changing lanes now, exiting from exit 34B in two miles, etc.).

[0090] Vehicle 600 also includes a network interface 624, which can communicate over one or more networks using one or more wireless antennas 626 and / or a modem. For example, network interface 624 can be used to communicate via Long Term Evolution (“LTE”), Wideband Code Division Multiple Access (“WCDMA”), Universal Mobile Telecommunications System (“UMTS”), Global System for Mobile Communications (“GSM”), IMT-CDMA Multicarrier (“CDMA2000”), etc. One or more wireless antennas 626 can also enable communication between objects in the environment (e.g., vehicles, mobile devices, etc.) using one or more local area networks (e.g., Bluetooth, Bluetooth Low Energy (“LE”), Z-Wave, ZigBee, etc.) and / or Low Power Wide Area Networks (LPWANs, such as LoRaWAN, SigFox, etc.).

[0091] Figure 6B According to some embodiments of this disclosure Figure 6A An example of the camera position and field of view of an exemplary autonomous vehicle 600. The camera and its respective field of view are an example embodiment and are not intended to be limiting. For example, additional and / or alternative cameras may be included and / or the cameras may be located at different positions on the vehicle 600.

[0092] The camera type may include, but is not limited to, a digital camera, which may be suitable for components and / or systems of vehicle 600. One or more cameras may operate at Automotive Safety Integrity Level (ASIL) B and / or other ASILs. According to embodiments, the camera type may have any image capture rate, such as 60 frames per second (fps), 120 fps, 240 fps, etc. The camera may use a rolling shutter, a global shutter, another type of shutter, or a combination thereof. In some examples, the color filter array may include a red transparent (RCCC) color filter array, a red transparent blue (RCCB) color filter array, a red blue green transparent (RBGC) color filter array, a Foveon X3 color filter array, a Bayer sensor (RGGB) color filter array, a monochrome sensor color filter array, and / or other types of color filter arrays. In some embodiments, a sharp-pixel camera, such as a camera with RCCC, RCCB, and / or RBGC color filter arrays, may be used to improve light sensitivity.

[0093] In some examples, one or more cameras can be used to perform advanced driver assistance system (ADAS) functions (e.g., as part of a redundancy or fail-safe design). For example, a multi-functional single camera can be installed to provide functions such as lane departure warning, traffic sign assistance, and intelligent headlight control. One or more cameras (e.g., all cameras) can simultaneously record and provide image data (e.g., video).

[0094] One or more cameras may be mounted in mounting components, such as custom-designed (3D-printed) components, to cut off stray light and interior reflections (e.g., dashboard reflections from the windshield rearview mirror) that could interfere with the camera's ability to capture image data. Regarding wing mirror mounting components, these components may be custom-3D printed so that the camera mounting plate matches the shape of the wing mirror. In some examples, one or more cameras may be integrated into the wing-shaped rearview mirror. For side-view cameras, one or more cameras may also be integrated into the four pillars at each corner of the cab.

[0095] A camera with a field of view including a portion of the environment in front of the vehicle (e.g., a front-facing camera) can be used for surround view to help identify forward paths and obstacles, and, with the assistance of one or more controllers 636 and / or control SOCs, to provide information crucial for generating an occupancy grid and / or determining the preferred vehicle path. The front-facing camera can be used to perform many of the same ADAS functions as lidar, including emergency braking, pedestrian detection, and collision avoidance. The forward-facing camera can also be used in ADAS functions and systems, including lane departure warning (“LDW”), adaptive cruise control (“ACC”), and / or other functions such as traffic sign recognition.

[0096] Various cameras can be used in front-facing configurations, including, for example, monocular camera platforms that include a complementary metal-oxide-semiconductor (“CMOS”) color imager. Another example could be a wide-angle camera 670 that can be used to perceive objects entering the field of view from the periphery (e.g., pedestrians, cross traffic, or bicycles). Although Figure 6B Only one wide-angle camera is shown, but the vehicle 600 may have any number (including zero) of wide-angle cameras 670. Furthermore, any number of remote cameras 698 (e.g., a pair of long-angle stereo cameras) can be used for depth-based object detection, particularly for objects for which neural networks have not yet been trained. One or more remote cameras 698 can also be used for object detection and classification, as well as basic object tracking.

[0097] Any number of stereo cameras 668 may also be included in the front-mounted configuration. In at least one embodiment, one or more stereo cameras 668 may include an integrated control unit comprising a scalable processing unit that may provide programmable logic (“FPGA”) and a multi-core microprocessor with an integrated controller area network (“CAN”) or Ethernet interface on a single chip. This unit can be used to generate a 3D map of the vehicle environment, including distance estimates for all points in the image. One or more alternative stereo cameras 668 may include a compact stereo vision sensor that may include two camera lenses (one on each side) and an image processing chip that measures the distance from the vehicle to a target object and uses the generated information (e.g., metadata) to activate automatic emergency braking and lane departure warning functions. In addition to the stereo cameras described herein, or alternatively, other types of stereo cameras 668 may be used.

[0098] A camera with a field of view including a portion of the vehicle's side environment (e.g., a side-view camera) can be used in the surround view to provide information for creating and updating the occupancy mesh and generating side collision warnings. For example, one or more surround cameras 674 (e.g., such as...) Figure 6B The four surround cameras 674 shown may be positioned on the vehicle 600. One or more surround cameras 674 may include one or more wide-angle cameras 670, one or more fisheye cameras, one or more 360-degree cameras, etc. For example, four fisheye cameras may be located at the front, rear, and sides of the vehicle. In an alternative arrangement, the vehicle may use three surround cameras 674 (e.g., left, right, and rear) and may utilize one or more other cameras (e.g., a front-facing camera) as a fourth surround-view camera.

[0099] Cameras with a view that includes a portion of the environment behind the vehicle 600 (e.g., rear-view cameras) can be used for parking assistance, surround view, rear-end collision warning, and creating and updating occupancy grids. A variety of cameras can be used, including but not limited to cameras that are also suitable as front-facing cameras (e.g., one or more long-range and / or mid-range cameras 698, one or more stereo cameras 668, one or more infrared cameras 672, etc.), as described herein.

[0100] Figure 6C According to some embodiments of this disclosure Figure 6AA block diagram of an exemplary system architecture for an exemplary autonomous vehicle 600 is provided. It should be understood that this and other arrangements described herein are merely illustrative. Other arrangements and elements (e.g., machines, interfaces, functions, sequences, functional groups, etc.) may be used in addition to the arrangements and elements shown, or other arrangements and elements may be used instead of those shown, and some elements may be omitted entirely. Furthermore, many of the elements described herein are functional entities that can be implemented as discrete or distributed components, or in combination with other components, and can be implemented in any suitable combination and location. The various functions described herein as being performed by entities can be performed by hardware, firmware, and / or software. For example, various functions can be performed by a processor executing instructions stored in memory.

[0101] Figure 6C Every component, feature, and system of vehicle 600 is connected via bus 602. Bus 602 may include a Controller Area Network (CAN) data interface (also referred to herein as the "CAN bus"). CAN can be a network within vehicle 600 used to help control various features and functions of vehicle 600, such as the actuation of brakes, acceleration, braking, steering, windshield wipers, etc. The CAN bus can be configured to have dozens or even hundreds of nodes, each with its own unique identifier (e.g., a CAN ID). The CAN bus can be read to locate steering wheel angle, ground speed, engine speed per minute (RPM), button positions, and / or other vehicle status indicators. The CAN bus may conform to the ASIL B standard.

[0102] Although bus 602 is described herein as a CAN bus, this is not intended to be limiting. For example, in addition to or from a CAN bus, FlexRay and / or Ethernet may also be used. Furthermore, although bus 602 is represented by a single line, this is not intended to be limiting. For example, any number of buses 602 may exist, which may include one or more CAN buses, one or more FlexRay buses, one or more Ethernet buses, and / or one or more other types of buses using different protocols. In some examples, two or more buses 602 may be used to perform different functions and / or for redundancy. For example, a first bus 602 may be used for a collision avoidance function, and a second bus 602 may be used for drive control. In any example, each bus 602 may communicate with any component of vehicle 600, and two or more buses 602 may communicate with the same component. In some examples, each SoC 604, each controller 636, and / or each computer within the vehicle may access the same input data (e.g., input from sensors of vehicle 600) and may be connected to a common bus, such as a CAN bus.

[0103] Vehicle 600 may include one or more controllers 636, as described herein. Figure 6A The controller 636 is described above. Controller 636 can be used for various functions. One or more controllers 636 can be coupled to any of the various other components and systems of vehicle 600, and can be used to control vehicle 600, artificial intelligence of vehicle 600, infotainment of vehicle 600, etc.

[0104] Vehicle 600 may include one or more System-on-Chip (SoC) 604. SoC 604 may include one or more CPUs 606, one or more GPUs 608, one or more processors 610, one or more caches 612, one or more accelerators 614, one or more data storage 616, and / or other components and features not shown. One or more SoCs 604 can be used to control vehicle 600 in various platforms and systems. For example, one or more SoCs 604 may be combined with an HD map 622 in a system (e.g., the system of vehicle 600), the HD map 622 being accessible via a network interface 624 from one or more servers (e.g., [server name missing]). Figure 6D Server 678) receives map refresh and / or updates.

[0105] One or more CPUs 606 may include CPU clusters or CPU complexes (or referred to herein as “CCPLEX”). One or more CPUs 606 may include multiple cores and / or a L2 cache. For example, in some embodiments, one or more CPUs 606 may include eight cores in a coherent multiprocessor configuration. In some embodiments, one or more CPUs 606 may include four dual-core clusters, each with a dedicated L2 cache (e.g., 2MB L2 cache). One or more CPUs 606 (e.g., CCPLEX) may be configured to support simultaneous cluster operation, such that any combination of clusters of CPUs 606 is active at any given time.

[0106] One or more CPU 606s can implement power management capabilities including one or more of the following features: automatic clock gating of a single hardware block when idle to conserve dynamic power; clock gating of each core when the core is not actively executing instructions due to executing WFI / WFE instructions; independent power gating of each core; independent clock gating of each core cluster when all cores are clock-gated or power-gated; and / or independent power gating of each core cluster when all cores are power-gated. One or more CPU 606s can further implement enhanced algorithms for managing power states, specifying allowed power states and expected wake-up times, and the hardware / microcode determines the optimal power state for the core, cluster, and CCPLEX to enter. The processing core can support simplified power state input sequences in software and offload the work to the microcode.

[0107] One or more GPUs 608 may include integrated GPUs (or referred to herein as “iGPUs”). GPUs 608 may be programmable and efficient for parallel workloads. In some examples, one or more GPUs 608 may use an enhanced tensor instruction set. One or more GPUs 608 may include one or more streaming microprocessors, wherein each streaming microprocessor may include a Level 1 cache (e.g., a Level 1 cache with at least 96KB of storage), and two or more streaming microprocessors may share a Level 2 cache (e.g., a Level 2 cache with 512KB of storage). In some embodiments, one or more GPUs 608 may include at least eight streaming microprocessors. One or more GPUs 608 may use one or more computation application programming interfaces (APIs). Furthermore, one or more GPUs 608 may use one or more parallel computing platforms and / or programming models (e.g., NVIDIA’s CUDA).

[0108] One or more GPU 608s can be power-optimized for optimal performance in automotive and embedded use cases. For example, one or more GPU 608s can be fabricated on FinFETs. However, this is not intended to limit, and other semiconductor manufacturing processes can be used to fabricate one or more GPU 608s. Each streaming microprocessor can combine multiple mixed-precision processing cores partitioned into multiple blocks. For example, but not limited to, 64 PF32 cores and 32 PF64 cores can be divided into four processing blocks. In such an example, each processing block could be allocated 16 FP32 cores, 8 FP64 cores, 16 INT32 cores, two mixed-precision NVIDIA tensor cores for deep learning matrix algorithms, an L0 instruction cache, a thread bundle scheduler, a dispatch unit, and / or a 64KB register file. Furthermore, the streaming microprocessor can include independent parallel integer and floating-point data paths to provide efficient execution of workloads through mixed computation and addressing computation. The streaming microprocessor can include independent thread scheduling capabilities to enable finer-grained synchronization and cooperation between parallel threads. Streaming microprocessors can include a combination of a level-one data cache and a shared memory unit to improve performance while simplifying programming.

[0109] One or more GPUs 608 may include a high-bandwidth memory (HBM) and / or a 16GB HBM2 memory subsystem to provide a peak memory bandwidth of approximately 900GB / s in some examples. In some examples, in addition to HBM memory, or optionally from HBM memory, synchronous graphics random access memory (SGRAM), such as fifth-generation graphics dual data rate synchronous random access memory (GDDR5), may be used.

[0110] The fifth-generation GPU 608 may include unified memory technology, which includes access counters to allow more accurate migration of memory pages to the processors that access them most frequently, thereby improving the efficiency of shared memory ranges between processors. In some examples, address translation service (ATS) support may be used to allow one or more GPUs 608 to directly access the page tables of one or more CPUs 606. In such examples, when one or more GPUs 608 memory management units (MMUs) experience a miss, an address translation request may be sent to one or more CPUs 606. In response, one or more CPUs 606 may look up the virtual-to-physical mapping of the address in their page tables and send the translation back to one or more GPUs 608. Therefore, unified memory technology allows for a single unified virtual address space for the memory of both one or more CPUs 606 and one or more GPUs 608, thereby simplifying the programming of one or more GPUs 608 and porting applications to one or more GPUs 608.

[0111] In addition, one or more GPUs 608 may include access counters that track the frequency with which one or more GPUs 608 access the memory of other processors. Access counters help ensure that memory pages are moved to the physical memory of the processor that accesses those pages most frequently.

[0112] One or more SoCs 604 may include any number of caches 612, including the caches 612 described herein. For example, one or more caches 612 may include an L3 cache available to one or more CPUs 606 and one or more GPUs 608 (e.g., connecting both one or more CPUs 606 and one or more GPUs 608). Caches 612 may include write-back caches with traceable thread state, for example, by using cache coherence protocols (e.g., MEI, MESI, MSI, etc.). Although a small cache size may be used, according to embodiments, the L3 cache may include 4 MB or more.

[0113] One or more SoCs 604 may include one or more arithmetic logic units (ALUs) that can be used to perform processing related to various tasks or operations of the vehicle 600, such as processing a DNN. Furthermore, one or more SoCs 604 may include one or more floating-point units (FPUs) or other mathematical coprocessors or digital coprocessor types for performing mathematical operations within the system. For example, one or more SoCs 604 may include one or more FPUs integrated as execution units within a CPU 606 and / or a GPU 608.

[0114] One or more SoCs 604 may include one or more accelerators 614 (e.g., hardware accelerators, software accelerators, or a combination thereof). For example, one or more SoCs 604 may include a hardware acceleration cluster that may include optimized hardware accelerators and / or large on-chip memory. Large on-chip memory (e.g., 4MB of SRAM) can enable the hardware acceleration cluster to accelerate neural networks and other computations. The hardware acceleration cluster can be used to supplement one or more GPUs 608 and offload some tasks from one or more GPUs 608 (e.g., freeing up more cycles from one or more GPUs 608 to perform other tasks). For example, one or more accelerators 614 may be sufficiently stable to be suitable for accelerating target workloads (e.g., perception, convolutional neural networks (CNNs), etc.). The term "CNN" as used herein may include all types of CNNs, including region-based or region convolutional neural networks (RCNNs) and fast RCNNs (e.g., for object detection).

[0115] One or more accelerators 614 (e.g., hardware acceleration clusters) may include one or more deep learning accelerators (DLAs). One or more DLAs may include one or more tensor processing units (TPUs) configured to provide an additional trillion operations per second for deep learning applications and inference. TPUs may be accelerators configured to perform image processing functions (e.g., for CNNs, RCNNs, etc.) and optimized for them. One or more DLAs may also be optimized for specific neural network types and floating-point operations and inference. One or more DLAs are designed to provide higher performance per millimeter than general-purpose GPUs and significantly outperform CPUs. One or more TPUs may perform multiple functions, including single-instance convolution functions, such as supporting INT8, INT16, and FP16 data types for features and weights, and post-processor functions.

[0116] One or more DLAs can execute neural networks, especially CNNs, quickly and efficiently on processed or unprocessed data for any of a variety of functions, including but not limited to: CNNs for object recognition and detection using data from camera sensors; CNNs for distance estimation using data from camera sensors; CNNs for emergency vehicle detection and recognition using microphone data; CNNs for facial recognition and vehicle owner recognition using data from camera sensors; and / or CNNs for safety and / or safety-related events.

[0117] One or more DLAs can perform any function of one or more GPUs 608. For example, by using inference accelerators, designers can perform any function for one or more DLAs or one or more GPUs 608. For example, designers can centralize the processing of CNNs and floating-point operations on one or more DLAs and leave other functions to one or more GPUs 608 and / or one or more other accelerators 614.

[0118] One or more accelerators 614 (e.g., hardware acceleration clusters) may include programmable vision accelerators (PVAs), which may also be referred to herein as computer vision accelerators. One or more PVAs may be designed and configured to accelerate computer vision algorithms for advanced driver assistance systems (ADAS), autonomous driving, and / or augmented reality (AR) and / or virtual reality (VR) applications. One or more PVAs may provide a balance between performance and flexibility. For example, each PVA may include, for example, but not limited to, any number of reduced instruction set computer (RISC) cores, direct memory access (DMA), and / or any number of vector processors.

[0119] The RISC core can interact with an image sensor (e.g., the image sensor of any camera described herein), one or more image signal processors, etc. Each RISC core may include any amount of memory. The RISC core can use any of a variety of protocols, depending on the implementation. In some examples, the RISC core can run a real-time operating system (RTOS). The RISC core can be implemented using one or more integrated circuit devices, application-specific integrated circuits (ASICs), and / or memory devices. For example, the RISC core may include an instruction cache and / or tightly coupled RAM.

[0120] DMA enables PVA components to access system memory independently of one or more CPUs. DMA can support any features used to optimize PVA, including but not limited to support for multidimensional addressing and / or circular addressing. In some examples, DMA can support addressing in up to six or more dimensions, which may include block width, block height, block depth, horizontal block step, vertical block step, and / or depth step.

[0121] Vector processors can be programmable processors designed to efficiently and flexibly execute computer vision algorithms and provide signal processing capabilities. In some examples, a PVA may include a PVA core and two vector processing subsystem partitions. The PVA core may include a processor subsystem, one or more DMA engines (e.g., two DMA engines), and / or other peripherals. The vector processing subsystem may operate as the main processing engine of the PVA and may include a vector processing unit (VPU), an instruction cache, and / or a vector memory (e.g., a VMEM). The VPU core may include a digital signal processor, such as a single-instruction, multiple-data (SIMD) or very long instruction word (VLIW) digital signal processor. The combination of SIMD and VLIW can improve throughput and speed.

[0122] Each vector processor may include an instruction cache and may be coupled to dedicated memory. As a result, in some examples, each vector processor may be configured to execute independently of other vector processors. In other examples, vector processors included in a particular PVA may be configured to employ data parallelism. For example, in some embodiments, multiple vector processors included in a single PVA may execute the same computer vision algorithm on different regions of an image. In other examples, vector processors included in a particular PVA may execute different computer vision algorithms simultaneously on the same image, or even different algorithms on a sequence of images or portions of an image. Among other things, any number of PVAs may be included in a hardware-accelerated cluster, and any number of vector processors may be included in each PVA. Furthermore, PVAs may include additional error-correcting code (ECC) memory to enhance overall system security.

[0123] One or more accelerators 614 (e.g., a hardware acceleration cluster) may include on-chip computer vision network and SRAM for providing high-bandwidth, low-latency SRAM for one or more accelerators 614. In some examples, the on-chip memory may include at least 4 MB of SRAM, including but not limited to eight field-configurable memory blocks accessible by the PVA and DLA. Each pair of memory blocks may include an Advanced Peripheral Bus (APB) interface, configuration circuitry, a controller, and a multiplexer. Any type of memory can be used. The PVA and DLA access memory via a backbone that provides high-speed memory access for the PVA and DLA. The backbone may include an on-chip computer vision network that interconnects the PVA and DLA to memory (e.g., using an APB).

[0124] An on-chip computer vision network may include an interface that determines that both the PVA and DLA have provided ready and valid signals before transmitting any control signals / addresses / data. Such an interface can provide independent phases and independent channels for transmitting control signals / addresses / data, as well as burst communication for continuous data transmission. This type of interface may conform to ISO 26262 or IEC 61508 standards, but other standards and protocols may also be used.

[0125] In some examples, one or more SoCs 604 may include a real-time ray tracing hardware accelerator, as described in U.S. Patent Application No. 16 / 101232, filed August 10, 2018. The real-time ray tracing hardware accelerator can be used to rapidly and efficiently determine the location and extent of an object (e.g., within a world model), generate real-time visualization simulations for radar signal interpretation, sound propagation synthesis and / or analysis, sonar system simulation, general wave propagation simulation, comparison with lidar data for localization and / or other functions and / or uses. In some embodiments, one or more tree traversal units (TTUs) may be used to perform one or more ray tracing-related operations.

[0126] One or more accelerators (e.g., hardware accelerator clusters) have wide applications in autonomous driving. A PVA (Programmable Vision accelerator) could be a programmable vision accelerator used in critical processing stages of ADA (Advanced Driver Assistance Systems) and autonomous vehicles. The capabilities of a PVA are well-suited for algorithmic domains requiring predictable processing with low power consumption and low latency. In other words, PVAs perform well on semi-intensive or conventionally intensive computations, even on small datasets that require predictable runtimes with low latency and low power consumption. Therefore, in the context of autonomous vehicle platforms, PVAs are designed to run classic computer vision algorithms, as they are highly efficient in object detection and integer mathematical operations.

[0127] For example, according to one embodiment of this technology, a PVA is used to perform computer stereo vision. In some examples, a semi-global matching-based algorithm can be used, although this is not intended to limit it. Many applications of Level 3-5 autonomous driving require instantaneous motion estimation / stereo matching (e.g., motion structures, pedestrian recognition, lane detection, etc.). A PVA can perform computer stereo vision functions on input from two monocular cameras.

[0128] In some examples, PVA can be used to perform dense optical flow, providing processed radar data based on the raw radar data (e.g., using 4D Fast Fourier Transform). In other examples, PVA is used for time-of-flight depth processing, for example, to provide processed time-of-flight data by processing the raw time-of-flight data.

[0129] DLA can be used to run any type of network to enhance control and driving safety, including neural networks that output a confidence measure for each object detection. Such a confidence value can be interpreted as a probability or to provide a relative “weight” for each detection relative to other detections. This confidence value allows the system to further determine which detections should be considered true positives rather than false positives. For example, the system can set a confidence threshold and only consider detections exceeding the threshold as true positives. In an Automatic Emergency Braking (AEB) system, false positives would cause the vehicle to automatically perform emergency braking, which is clearly undesirable. Therefore, only the most reliable detections should be considered as triggers for AEB. DLA can run a neural network to regress the confidence value. The neural network can take at least a subset of parameters as its input, such as bounding box dimensions, obtained ground plane estimates (e.g., from another subsystem), inertial measurement unit (IMU) sensor 666 outputs related to the vehicle's 600-degree orientation and distance, and three-dimensional position estimates of objects obtained from the neural network and / or other sensors (e.g., lidar sensor 664 or radar sensor 660).

[0130] One or more SoCs 604 may include one or more data storage units 616 (e.g., memory). The data storage unit 616 may be on-chip memory of the SoC 604, which may store neural networks to be executed on the GPU and / or DLA. In some examples, the capacity of the data storage unit 616 may be large enough to store multiple instances of neural networks for redundancy and security. The data storage unit 612 may include a level 2 or level 3 cache 612. As described herein, references to one or more data storage units 616 may include references to memory associated with the PVA, DLA, and / or one or more other accelerators 614.

[0131] One or more SoCs 604 may include one or more processors 610 (e.g., embedded processors). Processor 610 may include a boot and power management processor, which may be a dedicated processor and subsystem for handling boot power and management functions, as well as related security implementations. The boot and power management processor may be part of a boot sequence for one or more SoCs 604s and may provide runtime power management services. The boot power and management processor may provide clock and voltage programming, system low-power state transition assistance, management of SoC 604 thermal sensors and temperature sensors, and / or management of SoC 604 power states. Each temperature sensor may be implemented as a ring oscillator with an output frequency proportional to temperature, and one or more SoCs 604s may use the ring oscillator to detect the temperature of one or more CPUs 606s, one or more GPUs 608s, and / or one or more accelerators 614s. If a temperature is determined to exceed a threshold, the boot and power management processor may enter a temperature fault routine and place one or more SoCs 604s into a low-power state and / or place vehicle 600 into a driver-safe parking mode (e.g., safely parking vehicle 600).

[0132] One or more processors 610 may also include a set of embedded processors that can serve as an audio processing engine. The audio processing engine may be an audio subsystem capable of providing full hardware support for multi-channel audio through multiple interfaces and a wide range of flexible audio I / O interfaces. In some examples, the audio processing engine is a dedicated processor core of a digital signal processor with dedicated RAM.

[0133] One or more processors 610 may also include a normally-on processor engine that provides the necessary hardware functionality to support low-power sensor management and wake-up use cases. The normally-on processor engine may include a processor core, tightly coupled RAM, peripheral support (e.g., timers and interrupt controllers), various I / O controller peripherals, and routing logic.

[0134] One or more processors 610 may also include a secure cluster engine, which includes a dedicated processor subsystem for handling security management for automotive applications. The secure cluster engine may include two or more processor cores, tightly coupled RAM, support for peripheral devices (e.g., timers, interrupt controllers, etc.), and / or routing logic. In secure mode, the two or more cores may operate in lockstep mode and function as a single core with comparison logic to detect any differences in their operations.

[0135] One or more processors 610 may also include a real-time camera engine, which may include a dedicated processor subsystem for handling real-time camera management.

[0136] One or more processors 610 may also include a high dynamic range signal processor, which may include an image signal processor, which is a hardware engine as part of the camera processing pipeline.

[0137] One or more processors 610 may include a video image synthesizer, which may be a processing block (e.g., implemented on a microprocessor) that implements the video post-processing functions required by the video playback application to generate the final image of the player window. The video image synthesizer may perform lens distortion correction on one or more wide-angle cameras 670, one or more surround cameras 674, and / or in-cabin monitoring camera sensors. The in-cabin monitoring camera sensors are preferably monitored by a neural network running on another instance of an advanced SoC, configured to recognize in-cabin events and respond accordingly. The in-cabin system may perform lip reading to activate cellular service and make phone calls, dictate emails, change vehicle destination, activate or change the vehicle's infotainment system and settings, or provide voice-activated web browsing. Some functions are only available to the driver when the vehicle is running in automatic mode; otherwise, they are disabled.

[0138] Video image synthesizers may include enhanced temporal noise reduction for both spatial and temporal noise reduction. For example, in the case of motion in a video, noise reduction appropriately weights spatial information, thereby reducing the weight of information provided by adjacent frames. In cases where an image or part of an image does not contain motion, temporal noise reduction performed by the video image synthesizer can use information from the previous image to reduce noise in the current image.

[0139] The video image compositor can also be configured to perform stereoscopic correction on input stereoscopic shot frames. When the operating system desktop is in use, the video image compositor can also be used for user interface compositing, without requiring the GPU 608 to continuously render new surfaces. Even when one or more GPUs 608 are powered on and active during 3D rendering, the video image compositor can be used to offload one or more GPUs 608 to improve performance and responsiveness.

[0140] One or more SoCs 604 may also include a Mobile Industrial Processor Interface (MIPI) camera serial interface, a high-speed interface, and / or a video input block that can be used for receiving video and input from a camera and associated pixel input functions. One or more SoCs 604 may also include one or more input / output controllers that may be software-controlled and can be used to receive I / O signals that are not assigned a specific role.

[0141] One or more SoCs 604 may also include a wide range of peripheral interfaces for communication with peripheral devices, audio codecs, power management and / or other devices. One or more SoCs 604 may be used to process data from cameras (e.g., via gigabit multimedia serial links and Ethernet connections), sensors (e.g., one or more LiDAR sensors 664, one or more radar sensors 660, etc., connected via Ethernet), from bus 602 (e.g., vehicle 600 speed, steering wheel position, etc.), and from one or more GNSS sensors 658 (e.g., via Ethernet or CAN bus connections). One or more SoCs 604 may also include a dedicated high-performance, high-capacity memory controller, which may include its own DMA engine and may be used to free one or more CPUs 606 from routine data management tasks.

[0142] One or more SoCs 604 can form an end-to-end platform with a flexible architecture spanning automation levels 3-5, providing a comprehensive functional safety architecture that leverages and effectively utilizes computer vision and ADAS technologies to achieve diversity and redundancy, providing a platform for a flexible and reliable driver software stack and deep learning tools. Compared to traditional systems, one or more SoCs 604 can be faster, more reliable, and even more energy-efficient and space-saving. For example, when one or more accelerators 614 are combined with one or more CPUs 606, one or more GPUs 608, and one or more data storage units 616, a fast and efficient platform can be provided for Level 3-5 autonomous vehicles.

[0143] Therefore, this technology offers capabilities and functionalities that are unavailable in traditional systems. For example, computer vision algorithms can be executed on a CPU, which can be configured using a high-level programming language (such as C) to execute a wide variety of processing algorithms on diverse visual data. However, CPUs often cannot meet the performance requirements of many computer vision applications, such as those related to execution time and power consumption. In particular, many CPUs cannot execute complex object detection algorithms in real time, which is a requirement for automotive ADAS applications and practical Level 3-5 autonomous vehicles.

[0144] Compared to traditional systems, the techniques described in this paper, by providing CPU complexes, GPU complexes, and hardware acceleration clusters, allow for the simultaneous and / or sequential execution of multiple neural networks and the combination of results to achieve Level 3–5 autonomous driving capabilities. For example, a CNN executing on a DLA or dGPU (e.g., one or more GPU 620s) can include text and character recognition, allowing a supercomputer to read and understand traffic signs, including signs for which the neural network has not yet been specifically trained. The DLA can also include a neural network capable of recognizing, interpreting, and providing semantic understanding of symbols, and passing this semantic understanding to a path planning module running on the CPU complex.

[0145] Another example is the ability to run multiple neural networks simultaneously, as required for Level 3, 4, or 5 driving. For instance, a warning sign consisting of "Warning: Flashing lights indicate icing conditions" and a light can be interpreted independently or jointly by multiple neural networks. The sign itself can be recognized as a traffic sign by a first deployed neural network (e.g., a trained neural network), while the text "Flashing lights indicate icing conditions" can be interpreted by a second deployed neural network, which, when a flashing light is detected, notifies the vehicle routing software (preferably executed on a CPU complex) of the presence of icing conditions. A third deployed neural network can be used to identify the flashing light and notify the vehicle routing software of its presence by operating it across multiple frames. All three neural networks can run simultaneously, for example, within a DLA and / or on one or more GPUs 608.

[0146] In some examples, the CNN used for facial recognition and owner identification can use data from camera sensors to identify the presence of an authorized driver and / or the owner of the vehicle 600. The engine can be unlocked using a normally open sensor when the owner approaches the driver's door and turns on the lights, and the vehicle can be disabled in safe mode when the owner leaves. In this way, one or more SoCs 604 provide anti-theft and / or carjacking protection.

[0147] In another example, the CNN for emergency vehicle detection and identification can use data from microphone 696 to detect and identify emergency vehicle sirens. Unlike conventional systems that use a general classifier to detect sirens and manually extract features, one or more SoCs(s) 604 use the CNN to classify environmental and urban sounds, as well as visual data. In a preferred embodiment, the CNN running on the DLA is trained to identify the relative closing speed of emergency vehicles (e.g., by using the Doppler effect). The CNN can also be trained to identify emergency vehicles specific to the vehicle's operating area, as identified by one or more GNSS sensors 658. Thus, for example, when operating in Europe, the CNN will seek to detect European sirens, and when operating in the United States, the CNN will seek to identify only North American sirens. Once an emergency vehicle is detected, an emergency vehicle safety routine can be executed using a control program with the aid of ultrasonic sensors 662, causing the vehicle to slow down, pull over, stop, and / or idle until one or more emergency vehicles pass.

[0148] The vehicle may include one or more CPUs 618 (e.g., one or more discrete CPUs or one or more dCPUs) coupled to one or more SoCs 604 via high-speed interconnects (e.g., PCIe). For example, one or more CPUs 618 may include x86 processors. The CPUs 618 can be used to perform any of a variety of functions, including arbitrating potentially inconsistent results between ADAS sensors and SoCs 604, and / or monitoring the status and health of one or more controllers 636 and / or infotainment SoCs 630.

[0149] Vehicle 600 may include one or more GPUs 620 (e.g., one or more discrete GPUs or one or more dGPUs) coupled to SoC 604 via high-speed interconnects (e.g., NVIDIA's NVLINK). One or more GPUs 620 may provide additional artificial intelligence capabilities, such as by executing redundant and / or different neural networks, and may be used to train and / or update neural networks based on inputs from sensors of vehicle 600 (e.g., sensor data).

[0150] Vehicle 600 may also include a network interface 624, which may include one or more wireless antennas 626 (e.g., one or more wireless antennas for different communication protocols, such as cellular antennas, Bluetooth antennas, etc.). Network interface 624 can be used to enable wireless connectivity via the Internet to the cloud (e.g., with one or more servers 678 and / or other network devices), other vehicles, and / or computing devices (e.g., a passenger's client device). For communication with other vehicles, direct and / or indirect links can be established between the two vehicles (e.g., via a network and the Internet). A vehicle-to-vehicle communication link can provide a direct link. A vehicle-to-vehicle communication link can provide vehicle 600 with information about vehicles nearby (e.g., vehicles in front, to the side, and / or behind vehicle 600). This functionality may be part of vehicle 600's cooperative adaptive cruise control function.

[0151] Network interface 624 may include a SoC that provides modulation and demodulation functions and enables one or more controllers 636 to communicate over a wireless network. Network interface 624 may include an RF front-end for up-conversion from baseband to RF and down-conversion from RF to baseband. Frequency conversion can be performed using well-known processes and / or using superheterodyne processes. In some examples, the RF front-end functionality may be provided by a separate chip. The network interface may include wireless functions for communication via LTE, WCDMA, UMTS, GSM, CDMA2000, Bluetooth, Bluetooth LE, Wi-Fi, Z-Wave, ZigBee, LoRaWAN, and / or other wireless protocols.

[0152] The vehicle 600 may further include one or more data storage units 628, which may be off-chip (e.g., off-SoC). The data storage unit 628 may include one or more storage elements, including RAM, SRAM, DRAM, VRAM, flash memory, hard disk, and / or other components and / or devices capable of storing at least one bit of data.

[0153] The vehicle 600 may also include one or more GNSS sensors 658. One or more GNSS sensors 658 (e.g., GPS, assisted GPS sensors, differential GPS (DGPS) sensors, etc.) are used to assist in mapping, sensing, occupancy grid generation, and / or path planning functions. Any number of GNSS sensors 658 can be used, including, for example, but not limited to, GPS with a USB connector having an Ethernet-to-serial (RS-232) bridge.

[0154] Vehicle 600 may also include one or more radar sensors 660. Even in dark and / or inclement weather conditions, vehicle 600 can use one or more radar sensors 660 for remote vehicle detection. The radar functional safety level can be ASIL B. One or more radar sensors 660 can use CAN and / or bus 602 (e.g., to transmit data generated by one or more radar sensors 660) for control and access to target tracking data; in some examples, raw data is accessed via Ethernet. Various radar sensor types can be used. For example, but not limited to, one or more radar sensors 660 can be used for front, rear, and side radar applications. In some examples, pulse-Doppler radar sensors are used.

[0155] One or more radar sensors 660 may include different configurations, such as long-range with a narrow field of view, short-range with a wide field of view, short-range side coverage, etc. In some examples, long-range radar can be used for adaptive cruise control functions. Long-range radar systems can provide a wide field of view, for example, within a 250-meter range, achieved through two or more independent scans. One or more radar sensors 660 can help distinguish between static and moving objects and can be used by ADAS systems for emergency braking assistance and forward collision warning. Long-range radar sensors may include monostatic multimode radars with multiple (e.g., six or more) fixed radar antennas and high-speed CAN and FlexRay interfaces. In an example with six antennas, the four central antennas can create a focused beam pattern designed to record 600 elements of the environment around the vehicle at high speed with minimal traffic interference from adjacent lanes. The other two antennas can expand the field of view, enabling rapid detection of vehicles entering or leaving the vehicle within 600 lanes.

[0156] For example, a mid-range radar system may include a range of up to 660 meters (front) or 80 meters (rear), and a field of view of up to 42 degrees (front) or 650 degrees (rear). Short-range radar systems may include, but are not limited to, radar sensors designed to be mounted at both ends of the rear bumper. When mounted at both ends of the rear bumper, such radar sensor systems may generate two beams, continuously monitoring the blind spots behind and beside the vehicle.

[0157] ADAS systems can use short-range radar systems for blind spot detection and / or lane change assistance.

[0158] Vehicle 600 may also include one or more ultrasonic sensors 662. One or more ultrasonic sensors 662 may be located at the front, rear, and / or sides of vehicle 600 and may be used for parking assistance and / or creating and updating occupancy grids. Various ultrasonic sensors 662 may be used, and different ultrasonic sensors 662 may be used for different detection ranges (e.g., 2.5m, 4m). One or more ultrasonic sensors 662 may operate at the ASIL B functional safety level.

[0159] Vehicle 600 may include one or more lidar sensors 664. The one or more lidar sensors 664 may be used for object and pedestrian detection, emergency braking, collision avoidance, and / or other functions. The functional safety level of the one or more lidar sensors 664 may be ASIL B. In some examples, vehicle 600 may include multiple lidar sensors 664 (e.g., two, four, six, etc.) that can use Ethernet (e.g., to provide data to a Gigabit Ethernet switch).

[0160] In some examples, one or more LiDAR sensors 664 can provide a 360-degree field of view of objects and a list of their distances. One or more commercially available LiDAR sensors 664 have an advertised range of approximately 600m, an accuracy of 2cm-3cm, and support, for example, a 600Mbps Ethernet connection. In some examples, one or more non-protruding LiDAR sensors 664 can be used. In the examples of this type, one or more LiDAR sensors 664 can be implemented as a small device that can be embedded in the front, rear, sides, and / or corners of a vehicle 600. In the examples of this type, one or more LiDAR sensors 664 can provide a horizontal field of view of up to 120 degrees and a vertical field of view of 35 degrees, with a range of 200m even for low-reflectivity objects. A front-mounted one or more LiDAR sensors 664 can be configured with a horizontal field of view between 45 and 135 degrees.

[0161] In some examples, lidar technology, such as 3D flash lidar, can also be used. 3D flash lidar uses a laser flash as a transmission source, illuminating approximately 200 meters around the vehicle. A flash lidar device includes a receiver that records the laser pulse transmission time and reflected light on each pixel, with each pixel corresponding to the range from the vehicle to the object. Flash lidar can utilize each laser flash to generate a highly accurate, distortion-free environmental image. In some examples, four flash lidar sensors can be deployed, one on each side of the vehicle 600. Available 3D flash lidar systems include solid-state 3D staring array lidar cameras with no moving parts other than a fan (e.g., a non-scanning lidar device). The flash lidar device can use 5 nanosecond Class I (eye-safe) laser pulses per frame and can capture reflected laser light in the form of a 3D distance point cloud and co-registered intensity data. By using flash lidar, and because flash lidar is a solid-state device with no moving parts, one or more lidar sensors 664 can be less susceptible to motion blur, vibration, and / or shock.

[0162] The vehicle may also include one or more IMU sensors 666. In some examples, one or more IMU sensors 666 may be located at the center of the rear axle of the vehicle 600. One or more IMU sensors 666 may include, for example, but not limited to, one or more accelerometers, one or more magnetometers, one or more gyroscopes, one or more magnetic compasses, and / or other sensor types. In some examples, such as in a six-axis application, one or more IMU sensors 666 may include accelerometers and gyroscopes, while in a nine-axis application, one or more IMU sensors 666 may include accelerometers, gyroscopes, and magnetometers.

[0163] In some embodiments, one or more IMU sensors 666 can be implemented as a miniaturized, high-performance GPS-assisted inertial navigation system (GPS / INS) that combines a microelectromechanical system (MEMS) inertial sensor, a high-sensitivity GPS receiver, and an advanced Kalman filtering algorithm to provide estimates of position, velocity, and attitude. Thus, in some examples, one or more IMU sensors 666 can enable vehicle 600 to estimate heading by directly observing and correlating velocity changes from GPS to one or more IMU sensors 666, without requiring input from magnetic sensors. In some examples, one or more IMU sensors 666 and one or more GNSS sensors 658 can be combined in a single integrated unit.

[0164] The vehicle may include one or more microphones 696 placed inside and / or around the vehicle 600. One or more microphones 696 may be used for emergency vehicle detection and identification, etc.

[0165] The vehicle may also include any number of camera types, including one or more stereo cameras 668, one or more wide-angle cameras 670, one or more infrared cameras 672, one or more surround cameras 674, one or more long-range and / or mid-range cameras 698, and / or other camera types. The cameras can be used to capture image data of the entire perimeter of the vehicle 600. The types of cameras used depend on the embodiment and requirements of the vehicle 600, and any combination of camera types can be used to provide the necessary coverage around the vehicle 600. Furthermore, the number of cameras can vary depending on the embodiment. For example, the vehicle may include six cameras, seven cameras, ten cameras, twelve cameras, and / or other numbers of cameras. As an example, the cameras may support, but are not limited to, Gigabit Multimedia Serial Link (GMSL) and / or Gigabit Ethernet. This document will refer to... Figure 6A and Figure 6B Describe each camera in more detail.

[0166] Vehicle 600 may also include one or more vibration sensors 642. One or more vibration sensors 642 can measure vibrations of vehicle components, such as axles. For example, changes in vibration may indicate changes in road surface. In another example, when two or more vibration sensors 642 are used, differences between vibrations can be used to determine friction or slippage on the road surface (e.g., when the vibration difference is between a driven shaft and a freely rotating shaft).

[0167] Vehicle 600 may include ADAS system 638. In some examples, ADAS system 638 may include SoC. ADAS system 638 may include automatic / adaptive / automatic cruise control (ACC), cooperative adaptive cruise control (CACC), forward collision warning (FCW), automatic emergency braking (AEB), lane departure warning (LDW), lane keeping assist (LKA), blind spot warning (BSW), rear cross traffic warning (RCTW), collision warning system (CWS), lane centering (LC) and / or other features and functions.

[0168] The ACC system may use one or more radar sensors 660, one or more lidar sensors 664, and / or one or more cameras. The ACC system may include longitudinal ACC and / or lateral ACC. Longitudinal ACC monitors and controls the distance to the vehicle directly in front of vehicle 600 and automatically adjusts the vehicle speed to maintain a safe distance. Lateral ACC performs distance keeping and suggests lane changes to vehicle 600 if necessary. Lateral ACC is associated with other ADAS applications such as LCA and CWS.

[0169] CACC uses information from other vehicles, which can be received indirectly from other vehicles via a wireless link or a network connection (e.g., via the Internet) through network interface 624 and / or one or more wireless antennas 626. The direct link can be provided by a vehicle-to-vehicle (V2V) communication link, while the indirect link can be an infrastructure-to-vehicle (I2V) communication link. Generally, the V2V communication concept provides information about vehicles ahead (e.g., vehicles directly in front of vehicle 600 and in the same lane), while the I2V communication concept provides information about traffic ahead. A CACC system may include one or both I2V and V2V information sources. By taking into account information about vehicles ahead of vehicle 600, CACC may be more reliable and has the potential to improve traffic flow smoothness and reduce congestion on the road.

[0170] The Forward-Facing Warning (FCW) system is designed to alert the driver to hazards so that the driver can take corrective action. The FCW system uses a forward-facing camera and / or one or more radar sensors 660 coupled to a dedicated processor, DSP, FPGA, and / or ASIC, which are electrically coupled to driver feedback, such as a display, speaker, and / or vibration components. The FCW system can provide warnings such as audible, visual, haptic, and / or rapid braking pulses.

[0171] The AEB system detects an impending forward collision with another vehicle or other object. If the driver does not take corrective action within a specified time or distance parameter, the AEB system may automatically apply the brakes. The AEB system may use one or more front-facing cameras and / or one or more radar sensors 660 coupled to a dedicated processor, DSP, FPGA, and / or ASIC. When the AEB system detects a hazard, it typically first alerts the driver to take corrective action to avoid a collision. If the driver does not take corrective action, the AEB system may automatically apply the brakes to prevent or at least mitigate the effects of the anticipated collision. The AEB system may include technologies such as dynamic brake support and / or collision proximity braking.

[0172] The Lane Departure Warning (LDW) system provides visual, auditory, and / or tactile warnings, such as steering wheel or seat vibrations, to alert the driver when the vehicle crosses lane markings. The LDW system does not activate when the driver indicates intentional lane departure by activating a turn signal. The LDW system may use a front-facing camera coupled to a dedicated processor, DSP, FPGA, and / or ASIC, which is electrically coupled to actuator feedback, such as a display, speaker, and / or vibration assembly.

[0173] The LKA system is a variant of the LDW system. If the vehicle begins to leave the lane at 60°, the LKA system provides steering input or braking to correct the vehicle's 60° deviation.

[0174] The BSW system detects and warns drivers of vehicles within the blind spot. The BSW system can provide visual, auditory, and / or tactile alerts to indicate unsafe merging or lane changing. Additional warnings may be provided when the driver uses a turn signal. The BSW system may use a rear-facing camera and / or radar sensor 660 coupled to a dedicated processor, DSP, FPGA, and / or ASIC, which is electrically coupled to driver feedback, such as a display, speaker, and / or vibration assembly.

[0175] When the vehicle 600 is reversing and detects an object outside the range of the rear camera, the RCTW system can provide visual, auditory, and / or tactile notifications. Some RCTW systems include AEB (Autonomous Emergency Braking) to ensure the application of the vehicle's brakes to avoid a collision. The RCTW system may use one or more rear-facing radar sensors 660, which are coupled to a dedicated processor, DSP, FPGA, and / or ASIC, electrically coupled to actuator feedback, such as a display, speaker, and / or vibration assembly.

[0176] Traditional ADAS systems can be prone to false positives, which can be frustrating and distracting for drivers, but usually do not lead to catastrophic consequences because the ADAS system alerts the driver and allows them to determine whether a safe situation truly exists and take appropriate action. However, in an autonomous vehicle 600, in the event of conflicting results, the vehicle 600 itself must decide whether to listen to the results from the main computer or the auxiliary computer (e.g., the first controller 636 or the second controller 636). For example, in some embodiments, the ADAS system 638 may be a backup and / or auxiliary computer for providing perception information to a backup computer module. A backup computer rationality monitor may run redundant different software on hardware components to detect faults in perception and dynamic driving tasks. The output from the ADAS system 638 may be provided to a monitoring MCU. If the outputs of the main computer and the auxiliary computer conflict, the monitoring MCU must determine how to reconcile the conflict to ensure safe operation.

[0177] In some examples, the master computer can be configured to provide a confidence score to the monitoring MCU, indicating the master computer's confidence level in a selected result. If the confidence score exceeds a threshold, the monitoring MCU can follow the master computer's direction regardless of whether the auxiliary computer provides conflicting or inconsistent results. If the confidence score does not meet the threshold, and the master computer and the auxiliary computer indicate different results (e.g., conflict), the monitoring MCU can arbitrate between the computers to determine the appropriate result.

[0178] The monitoring MCU can be configured to run one or more trained and configured neural networks to determine the conditions under which the auxiliary computer provides a false alarm based on the outputs of the main computer and the auxiliary computer. Thus, one or more neural networks in the monitoring MCU can learn when the output of the auxiliary computer is reliable and when it is not. For example, when the auxiliary computer is a radar-based FCW system, one or more neural networks in the monitoring MCU can learn when the FCW system identifies a metallic object that is not actually dangerous, such as a drain grille or manhole cover that triggers an alarm. Similarly, when the auxiliary computer is a camera-based LDW system, the neural network in the monitoring MCU can learn to override the LDW when a bicycle or pedestrian is present and lane departure is actually the safest maneuver. In embodiments that include one or more neural networks running on the monitoring MCU, the monitoring MCU may include at least one of a DLA or GPU adapted to run one or more neural networks with associated memory. In a preferred embodiment, the monitoring MCU may include and / or include components that are SoC 604.

[0179] In other examples, ADAS system 638 may include an auxiliary computer that performs ADAS functions using conventional computer vision rules. Therefore, the auxiliary computer can use classic computer vision rules (if-then), and the presence of one or more neural networks in the monitoring MCU can improve reliability, safety, and performance. For example, diverse implementations and intentional non-identification make the entire system more fault-tolerant, especially to failures caused by software (or hardware / software interface) functionality. For instance, if a software defect or error exists in the software running on the main computer, and different software code running on the auxiliary computer provides the same overall result, the monitoring MCU can have greater confidence in the correctness of the overall result, and the software or hardware defect on the main computer did not lead to a major error.

[0180] In some examples, the output of ADAS system 638 may be fed into the perception block and / or the dynamic drive task block of the host computer. For example, if ADAS system 638 indicates a forward collision warning due to an object directly in front, the perception block may use this information when identifying the object. In other examples, as described herein, the secondary computer may have its own trained neural network, thereby reducing the risk of false positives.

[0181] Vehicle 600 may also include an infotainment SoC 630 (e.g., an in-vehicle infotainment system (IVI)). Although shown and described as an SoC, the infotainment system may not be an SoC and may include two or more discrete components. The infotainment SoC 630 may include a combination of hardware and software that can be used to provide audio (e.g., music, personal digital assistant, navigation instructions, news, radio, etc.), video (e.g., television, movies, streaming media, etc.), telephone (e.g., hands-free calling), network connectivity (e.g., LTE, Wi-Fi, etc.), and / or provide information services to vehicle 600 (e.g., navigation system, rear parking assist, radio data system, vehicle-related information such as fuel level, total distance traveled, brake fuel level, engine oil level, door opening / closing, air filter information, etc.). For example, the infotainment SoC 630 may be a radio, disk player, navigation system, video player, USB and Bluetooth connectivity, in-vehicle computer, in-vehicle entertainment, Wi-Fi, steering wheel audio controls, hands-free voice control, head-up display (HUD), HMI display 634, telecom device, control panel (e.g., for controlling and / or interacting with various components, features, and / or systems), and / or other components. The infotainment SoC 630 may also be used to provide information to vehicle users (e.g., visual and / or auditory), such as information from ADAS system 638, autonomous driving information such as planned vehicle maneuvers, trajectories, surrounding environment information (e.g., intersection information, vehicle information, road information, etc.), and / or other information.

[0182] The infotainment SoC 630 may include GPU functionality. The infotainment SoC 630 can communicate with other devices, systems, and / or components of the vehicle 600 via bus 602 (e.g., CAN bus, Ethernet, etc.). In some examples, the infotainment SoC 630 may be coupled to a monitoring MCU, allowing the GPU of the infotainment system to perform self-driving functions in the event of a failure of one or more main controllers 636 (e.g., the main computer and / or backup computer of the vehicle 600). In such an example, the infotainment SoC 630 may place the vehicle 600 into a driver-to-safe parking mode, as described herein.

[0183] Vehicle 600 may also include an instrument cluster 632 (e.g., a digital instrument panel, electronic instrument cluster, etc.). The instrument cluster 632 may include a controller and / or a supercomputer (e.g., a discrete controller or supercomputer). The instrument cluster 632 may include a set of instruments such as a speedometer, fuel level, oil pressure, tachometer, odometer, turn signals, shift position indicator, one or more seatbelt warning lights, one or more parking brake warning lights, one or more engine malfunction lights, airbag (SRS) system information, lighting controls, safety system controls, navigation information, etc. In some examples, information may be displayed and / or shared between the infotainment SoC 630 and the instrument cluster 632. In other words, the instrument cluster 632 may be included as part of the infotainment SoC 630, and vice versa.

[0184] Figure 6D It is one or more cloud-based servers according to some embodiments of this disclosure and Figure 6A A system diagram illustrating communication between exemplary autonomous vehicles 600 is provided. System 676 may include one or more servers 678, one or more networks 690, and vehicles, including vehicle 600. One or more servers 678 may include multiple GPUs 684(a)-684(H) (collectively referred to as GPU 684), PCIe switches 682(a)-682(H) (collectively referred to as PCIe switch 682), and / or CPUs 680(A)-680(B) (collectively referred to herein as CPU 680). GPUs 684, CPUs 680, and PCIe switches may be interconnected via high-speed interconnects, such as, but not limited to, NVIDIA-developed NVLink interface 688 and / or PCIe connection 686. In some examples, GPUs 684 are connected via NVLink and / or NV switch SoCs, and GPUs 684 and PCIe switches 682 are connected via PCIe interconnects. Although eight GPUs 684, two CPUs 680, and two PCIe switches are illustrated, this is not intended to be limiting. According to an embodiment, each of one or more servers 678 may include any number of GPUs 684, CPUs 680, and / or PCIe switches. For example, one or more servers 678 may each include eight, sixteen, thirty-two, and / or more GPUs 684.

[0185] One or more servers 678 may receive image data representing images from vehicles via one or more networks 690, showing unexpected or changed road conditions, such as recently started roadwork. One or more servers 678 may send neural network 692, updated neural network 692, and / or map information 694, including information about traffic and road conditions, to vehicles via one or more networks 690. Updates to map information 694 may include updates to HD map 622, such as information about construction sites, potholes, detours, floods, and / or other obstacles. In some examples, neural network 692, updated neural network 692, and / or map information 694 may originate from new training and / or experience represented in data received from any number of vehicles in the environment, and / or based on training performed in a data center (e.g., using one or more servers 678 and / or other servers).

[0186] One or more servers 678 can be used to train machine learning models (e.g., neural networks) based on training data. Training data may be generated by the vehicle and / or generated in a simulation (e.g., using a game engine). In some examples, the training data is labeled (e.g., the neural network benefits from supervised learning) and / or undergoes other preprocessing, while in other examples, the training data is unlabeled and / or preprocessed (e.g., the neural network does not require supervised learning). Training can be performed according to any one or more classes of machine learning techniques, including but not limited to: supervised training, semi-supervised training, unsupervised training, self-learning, reinforcement learning, joint learning, transfer learning, feature learning (including principal component analysis and cluster analysis), multilinear subspace learning, manifold learning, representation learning (including alternative dictionary learning), rule-based machine learning, anomaly detection, and any variations or combinations thereof. Once the machine learning model is trained, it can be used by the vehicle (e.g., transmitted to vehicle 690 via one or more networks, and / or the machine learning model can be used by one or more servers 678 for remote monitoring of the vehicle).

[0187] In some examples, one or more servers 678 may receive data from a vehicle and apply the data to state-of-the-art real-time neural networks for real-time intelligent inference. One or more servers 678 may include a deep learning supercomputer and / or a dedicated AI computer powered by a GPU 684, such as the DGX and DGX Station machines developed by NVIDIA. However, in some examples, one or more servers 678 may include a deep learning infrastructure in a data center using only CPU power.

[0188] The deep learning infrastructure of one or more servers 678 can perform rapid real-time inference and use this capability to assess and verify the health status of the processors, software, and / or associated hardware in vehicle 600. For example, the deep learning infrastructure can receive periodic updates from vehicle 600, such as sequences of images and / or objects located by vehicle 600 in the image sequence (e.g., through computer vision and / or other machine learning object classification techniques). The deep learning infrastructure can run its own neural network to identify objects and compare them with objects identified by vehicle 600. If the results do not match and the infrastructure concludes that the AI ​​in vehicle 600 has malfunctioned, one or more servers 678 can send a signal to vehicle 600 instructing the fail-safe computer of vehicle 600 to take control, notify passengers, and complete a safe stopping operation.

[0189] For inference, one or more servers 678 may include one or more GPUs 684 and one or more programmable inference accelerators (such as NVIDIA's TensorRT). The combination of GPU-powered servers and inference acceleration can enable real-time responses. In other examples, such as where performance is less critical, inference can be performed using servers powered by CPUs, FPGAs, and other processors.

[0190] Example computing device

[0191] Figure 7 This is a block diagram of an example computing device 700 suitable for implementing some embodiments of the present disclosure. The computing device 700 may include an interconnect system 702 directly or indirectly coupled to: a memory 704, one or more central processing units (CPUs) 706, one or more graphics processing units (GPUs) 708, a communication interface 710, input / output (I / O) ports 712, input / output components 714, a power supply 716, one or more presentation components 718 (e.g., displays), and one or more logic units 720. In at least one embodiment, one or more computing devices 700 may include one or more virtual machines (VMs), and / or any component thereof may include virtual components (e.g., virtual hardware components). For a non-limiting example, one or more GPUs 708 may include one or more vGPUs, one or more CPUs 706 may include one or more vCPUs, and / or one or more logic units 720 may include one or more virtual logic units. Therefore, one or more computing devices 700 may include discrete components (e.g., a complete GPU dedicated to computing device 700), virtual components (e.g., a portion of the GPU dedicated to computing device 700), or a combination thereof.

[0192] although Figure 7The various modules are shown as being connected to lines via interconnect system 702, but this is not intended to be limiting, but merely for clarity. For example, in some embodiments, a presentation component 718 such as a display device can be considered as I / O component 714 (e.g., if the display is a touchscreen). As another example, CPU 706 and / or GPU 708 may include memory (e.g., memory 704 may also represent a storage device in addition to the memory of GPU 708, CPU 706, and / or other components). In other words, Figure 7 The computing devices mentioned are merely illustrative. No distinction is made between "workstation," "server," "laptop," "desktop," "tablet," "client device," "mobile device," "handheld device," "game console," "electronic control unit (ECU)," "virtual reality system," and / or other device or system types, as is the case in [the context of the previous sentence]. Figure 7 As envisioned within the scope of computing devices.

[0193] Interconnect system 702 may represent one or more links or buses, such as address buses, data buses, control buses, or combinations thereof. Interconnect system 702 may include one or more bus or link types, such as Industry Standard Architecture (ISA) buses, Extended Industry Standard Architecture (EISA) buses, Video Electronics Standards Association (VESA) buses, Peripheral Component Interconnect (PCI) buses, Peripheral Component Interconnect Through (PCIE) buses, and / or other types of buses or links. In some embodiments, there is a direct connection between components. For example, CPU 706 may be directly connected to memory 704. Furthermore, CPU 706 may be directly connected to GPU 708. In cases where there is a direct connection or point-to-point connection between components, interconnect system 702 may include a PCIe link for performing the connection. In these examples, a PCI bus is not required in computing device 700.

[0194] The memory 704 may include any of a variety of computer-readable media. Computer-readable media can be any available medium accessible by the computing device 700. Computer-readable media may include volatile and non-volatile media, as well as removable and non-removable media. By way of example and not limitation, computer-readable media may include computer storage media and communication media.

[0195] Computer storage media may include volatile and non-volatile media and / or removable and non-removable media implemented in any method or technology for storing information, such as computer-readable instructions, data structures, program modules, and / or other data types. For example, memory 704 may store computer-readable instructions (e.g., instructions representing one or more programs and / or one or more program elements), such as an operating system. Computer storage media may include, but is not limited to, RAM, ROM, EEPROM, flash memory or other storage technologies, CD-ROM, digital versatile disk (DVD) or other optical disc storage, cassette tape, magnetic tape, disk storage or other magnetic storage devices, or any other medium that can be used to store desired information and is accessible by computing device 700. As used herein, computer storage media itself does not include signals.

[0196] Computer storage media can contain computer-readable instructions, data structures, program modules, and / or other data types contained in modulated data signals (such as carrier waves or other transmission mechanisms), and include any information delivery medium. The term "modulated data signal" can refer to a signal whose one or more characteristics are set or altered in a manner that encodes information in the signal. By way of example and not limitation, computer storage media can include wired media, such as wired networks or direct wired connections, and wireless media, such as acoustic, RF, infrared, and other wireless media. Any combination of the foregoing should also be included within the scope of computer-readable media.

[0197] One or more CPUs 706 may be configured to execute at least some computer-readable instructions to control one or more components of computing device 700 to perform one or more methods and / or processes described herein. Each of the one or more CPUs 706 may include one or more cores capable of processing multiple software threads simultaneously (e.g., one, two, four, eight, twenty-eight, seventy-two, etc.). The one or more CPUs 706 may include any type of processor and may include different types of processors depending on the type of computing device 700 implemented (e.g., processors with fewer cores for mobile devices and processors with more cores for servers). For example, depending on the type of computing device 700, the processor may be an advanced RISC machine (ARM) processor implemented using Reduced Instruction Set Computing (RISC) or an x86 processor implemented using Complex Instruction Set Computing (CISC). In addition to one or more microprocessors or supplementary coprocessors (e.g., math coprocessors), computing device 700 may also include one or more CPUs 706.

[0198] In addition to, or selected from, one or more CPUs 706, one or more GPUs 708 may be configured to execute at least some computer-readable instructions to control one or more components of computing device 700 to perform one or more methods and / or processes described herein. One or more GPUs 708 may be integrated GPUs (e.g., having one or more CPUs 706 and / or one or more GPUs 708 may be discrete GPUs). In embodiments, one or more GPUs 708 may be coprocessors of one or more CPUs 706. Computing device 700 may use one or more GPUs 708 to render graphics (e.g., 3D graphics) or perform general-purpose computing. For example, one or more GPUs 708 may be used for general-purpose computing on a GPU (GPGPU). One or more GPUs 708 may include hundreds or thousands of cores capable of processing hundreds or thousands of software threads simultaneously. One or more GPUs 708 may generate pixel data for an output image in response to rendering commands (e.g., rendering commands received via a host interface from one or more CPUs 706). One or more GPUs 708 may include graphics memory, such as display memory, for storing pixel data or any other suitable data, such as GPGPU data. Display memory may be included as part of memory 704. One or more GPUs 708 may include two or more GPUs operating in parallel (e.g., via a link). The link may directly connect the GPUs (e.g., using NVLINK) or connect the GPUs via a switch (e.g., using NVSwitch). When combined, each GPU 708 may generate pixel data or GPGPU data for different portions of the output or different outputs (e.g., a first GPU for a first image, a second GPU for a second image). Each GPU may include its own memory or may share memory with other GPUs.

[0199] In addition to one or more CPUs 706 and / or one or more GPUs 708, or selected from one or more CPUs 706 and / or one or more GPUs 708, one or more logic units 720 may be configured to execute at least some computer-readable instructions to control one or more components of computing device 700 to perform one or more methods and / or processes described herein. In embodiments, one or more CPUs 706, one or more GPUs 708, and / or one or more logic units 720 may execute any combination of methods, processes, and / or portions thereof discretely or jointly. One or more logic units 720 may be part of and / or integrated into one or more CPUs 706 and / or GPUs 708, and / or one or more logic units 720 may be discrete components or otherwise located external to one or more CPUs 706 and / or one or more GPUs 708. One or more logic units 720 may be coprocessors of one or more CPUs 706 and / or one or more GPUs 708.

[0200] Examples of one or more logic units 720 include one or more processing cores and / or components thereof, such as a data processing unit (DPU), a tensor core (TC), a tensor processing unit (TPU), a pixel vision core (PVC), a vision processing unit (VPU), a graphics processing cluster (GPC), a texture processing cluster (TPC), a streaming multiprocessor (SM), a tree traversal unit (TTU), an artificial intelligence accelerator (AIA), a deep learning accelerator (DLA), an arithmetic logic unit (ALU), an application-specific integrated circuit (ASIC), a floating-point unit (FPU), input / output (I / O) elements, peripheral component interconnect (PCI) or peripheral component interconnect pass-through (PCIe) elements, and / or the like.

[0201] The communication interface 710 may include one or more receivers, transmitters, and / or transceivers that enable the computing device 700 to communicate with other computing devices via an electronic communication network including wired and / or wireless communications. The communication interface 710 may include components and functions to support communication over any of a variety of different networks, such as wireless networks (e.g., Wi-Fi, Z-Wave, Bluetooth, Bluetooth LE, ZigBee, etc.), wired networks (e.g., communication over Ethernet or InfiniBand), low-power wide-area networks (e.g., LoRaWAN, SigFox, etc.), and / or the Internet. In one or more embodiments, one or more logic units 720 and / or the communication interface 710 may include one or more data processing units (DPUs) for directly transmitting data received via a network and / or via interconnect system 702 to one or more GPUs 708 (e.g., their memory).

[0202] I / O port 712 enables the computing device 700 to be logically coupled to other devices, including I / O components 714, one or more presentation components 718, and / or other components, some of which may be built into (e.g., integrated into) the computing device 700. Illustrative I / O components 714 include microphones, mice, keyboards, joysticks, game pads, game controllers, satellite antennas, scanners, printers, wireless devices, etc. I / O components 714 can provide a natural user interface (NUI) that processes user-generated air gestures, voice, or other physiological input. In some cases, the input can be transmitted to appropriate network elements for further processing. The NUI can implement any combination of voice recognition, stylus recognition, facial recognition, biometrics, on-screen and near-screen gesture recognition, air gestures, head and eye tracking, and touch recognition associated with the display of the computing device 700 (described in more detail below). The computing device 700 may include depth cameras, such as stereo camera systems, infrared camera systems, RGB camera systems, touchscreen technology, and combinations thereof, for gesture detection and recognition. In addition, computing device 700 may include an accelerometer or gyroscope capable of detecting motion (e.g., as part of an inertial measurement unit (IMU)). In some examples, the output of the accelerometer or gyroscope may be used by computing device 700 to render immersive augmented reality or virtual reality.

[0203] Power supply 716 may include hard-wired power supply, battery power supply, or a combination thereof. Power supply 716 may supply power to computing device 700 so that components of computing device 700 can operate.

[0204] One or more presentation components 718 may include displays (e.g., monitors, touchscreens, television screens, head-up displays (HUDs), other display types, or combinations thereof), speakers, and / or other presentation components. One or more presentation components 718 may receive data from other components (e.g., one or more GPUs 708, one or more CPUs 706, DPUs, etc.) and output data (e.g., as images, videos, sounds, etc.).

[0205] Example Data Center

[0206] Figure 8 An example data center 800 that can be used in at least one embodiment of this disclosure is shown. The data center 800 may include a data center infrastructure layer 810, a framework layer 820, a software layer 830, and / or an application layer 840.

[0207] like Figure 8 As shown, the data center infrastructure layer 810 may include a resource coordinator 812, packet computing resources 814, and node computing resources (“nodes CRs”) 816(1)-816(N), where “N” represents any integer. In at least one embodiment, nodes CRs 816(1)-816(N) may include, but are not limited to, any number of central processing units (CPUs) or other processors (including DPUs, accelerators, field-programmable gate arrays (FPGAs), graphics processors or graphics processing units (GPUs), etc.), storage devices (e.g., dynamic read-only memory), and in some embodiments, storage devices (e.g., solid-state or disk drives), network input / output (NW I / O) devices, network switches, virtual machines (VMs), power modules and / or cooling modules, etc. One or more nodes CRs 816(1)-816(N) may correspond to a server having one or more of the aforementioned computing resources. In addition, in some embodiments, nodes CRs 816(1)-8161(N) may include one or more virtual components, such as vGPU, vCPU and / or similar components, and / or one or more nodes CRs 816(1)-816(N) may correspond to virtual machines (VMs).

[0208] In at least one embodiment, the packet computing resource 814 may include individual packets of node CRs 816 located within one or more racks (not shown), or multiple racks located within data centers in different geographical locations (also not shown). Individual packets of node CRs 816 within the packet computing resource 814 may include packet computing, networking, memory, or storage resources, which may be configured or allocated to support one or more workloads. In at least one embodiment, multiple node CRs 816, including CPUs, GPUs, DPUs, and / or other processors, may be grouped within one or more racks to provide computing resources to support one or more workloads. One or more racks may also include any combination of any number of power modules, cooling modules, and / or network switches.

[0209] Resource coordinator 812 may be configured or otherwise control one or more nodes CRs816(1)-816(N) and / or grouped computing resources 814. In at least one embodiment, resource coordinator 812 may include a Software Design Infrastructure (SDI) management entity for data center 800. Resource coordinator 812 may include hardware, software, or some combination thereof.

[0210] In at least one embodiment, such as Figure 8 As shown, framework layer 820 may include job scheduler 833, configuration manager 834, resource manager 836, and / or distributed file system 838. Framework layer 820 may include frameworks for software 832 supporting software layer 830 and / or one or more applications 842 supporting application layer 840. Software 832 or one or more applications 842 may respectively include web-based service software or applications, such as services provided by Amazon Web Services, Google Cloud, and Microsoft Azure. Framework layer 820 may be, but is not limited to, a free and open-source software web application framework, such as Apache Spark™ (hereinafter referred to as "Spark"), which can utilize distributed file system 838 for large-scale data processing (e.g., "big data"). In at least one embodiment, job scheduler 833 may include Spark drivers to facilitate the scheduling of workloads supported by the various layers of data center 800. Configuration manager 834 may configure different layers, such as software layer 830 and framework layer 820, including Spark and distributed file system 838, to support large-scale data processing. Resource manager 836 may be able to manage cluster or group computing resources mapped to or allocated to support distributed file system 838 and job scheduler 833. In at least one embodiment, the cluster or group computing resources may include group computing resources 814 at data center infrastructure layer 810. Resource manager 836 may coordinate with resource coordinator 812 to manage these mapped or allocated computing resources.

[0211] In at least one embodiment, the software 832 included in the software layer 830 may include software used in at least a portion of the distributed file system 838 of the nodes CRs 816(1)-816(N), the grouped computing resources 814, and / or the framework layer 820. One or more types of software may include, but are not limited to, internet webpage search software, email virus scanning software, database software, and streaming video content software.

[0212] In at least one embodiment, the application 842 included in the application layer 840 may include one or more types of applications used by at least a portion of the nodes CRs 816(1)-816(N), the grouped computing resources 814 and / or the distributed file system 838 of the framework layer 820, but is not limited to any number of genomics applications, perceptual computing and machine learning applications, including training or inference software, machine learning framework software (e.g., PyTorch, TensorFlow, Caffe, etc.) and / or other machine learning applications used in combination with one or more embodiments.

[0213] In at least one embodiment, any of the configuration manager 834, resource manager 836, and resource coordinator 812 can implement any number and type of self-modification actions based on any amount and type of data acquired in any technically feasible manner. Self-modification actions can protect data center operators of data center 800 from making potentially erroneous configuration decisions and may prevent underutilized and / or poorly performing portions of the data center.

[0214] According to one or more embodiments described herein, data center 800 may include tools, services, software, or other resources for training one or more machine learning models or using one or more machine learning models to predict or infer information. For example, one or more machine learning models may be trained by calculating weight parameters based on a neural network architecture using the software and / or computing resources described above for data center 800. In at least one embodiment, a trained or deployed machine learning model corresponding to one or more neural networks may be used to infer or predict information using the resources described above for data center 800 by using weight parameters calculated through one or more training techniques (e.g., but not limited to the training techniques described herein).

[0215] In at least one embodiment, the data center 800 may use a CPU, application-specific integrated circuit (ASIC), GPU, FPGA, and / or other hardware (or corresponding virtual computing resources) to perform training and / or inference using the aforementioned resources. Furthermore, one or more of the aforementioned software and / or hardware resources may be configured as services to allow a user to train or perform information inference, such as image recognition, speech recognition, or other artificial intelligence services.

[0216] Example language model

[0217] In at least some embodiments, language models such as Large Language Models (LLMs), Visual Language Models (VLMs), Multimodal Language Models (MMLMs), and / or other types of generative artificial intelligence (AI) can be implemented. These models may be able to understand, summarize, translate, and / or otherwise generate text (e.g., natural language text, code, etc.), images, videos, computer-aided design (CAD) assets, OMNIVERSE and / or METAVERSE file information (e.g., USD formats such as OpenUSD), and / or the like based on context provided in input prompts or queries. In embodiments, these language models may be considered “large” because they are trained on massive datasets and have architectures with a large number of learnable network parameters (weights and biases)—e.g., millions or billions of parameters. LLMs / VLMs / MMLMs / etc. can be implemented for summarizing textual data, analyzing data (e.g., text, images, videos, etc.), extracting insights from data (e.g., text, images, videos, etc.), and generating new text / images / videos / etc. in a user-specified style, tone, and / or format. In some embodiments, the LLM / VLM / MMLM / etc. disclosed herein may be specifically designed for text processing, while in others, a multimodal LLM may be implemented to accept, understand, and / or generate text and / or other types of content, such as images, audio, 2D and / or 3D data (e.g., USD format) and / or video. For example, a Visual Language Model (VLM) or more specifically a Multimodal Language Model (MMLM) may be implemented to accept images, video, audio, text, 3D designs (e.g., CAD) and / or other input data types and / or generate or output images, video, audio, text, 3D designs and / or other output data types.

[0218] Various types of LLM / VLM / MMLM / etc. architectures can be implemented in various embodiments. For example, different architectures can be implemented using different techniques to understand and generate outputs (e.g., text, audio, video, images, 2D and / or 3D design or asset data, etc.). In some embodiments, LLM / VLM / MMLM / etc. architectures (e.g., recurrent neural networks (RNNs) or long short-term memory networks (LSTMs)) can be used, while in other embodiments, converter architectures (e.g., architectures relying on self-attention and / or cross-attention (e.g., between contextual data and textual data) mechanisms) can be used to understand and recognize relationships between words or tokens and / or contextual data (e.g., other text, video, images, design data, USD, etc.). One or more generative processing pipelines including LLM / VLM / MMLM / etc. may also include one or more diffusion blocks (e.g., noise reduction blocks). The LLM / VLM / MMLM / etc. of this disclosure may include encoder and / or decoder blocks. For example, discriminative or encoder-only models (e.g., BERT (Bidirectional Encoder Representations from Transformers)) can be implemented for tasks involving language understanding (e.g., classification, sentiment analysis, question answering, and named entity recognition). As another example, generative or decoder-only models (e.g., GPT (Generative Pretrained Transformer)) can be implemented for tasks involving language and content generation (e.g., text completion, story generation, and dialogue generation). LLM / VLM / MMLM / etc., including encoder and decoder components (e.g., T5 (Text-to-Text Transformer)), can be implemented to understand and generate content, such as for translation and summarization. These examples are not intended to be limiting and any architecture type (including, but not limited to, those described herein) can be implemented depending on the specific implementation and the task performed using LLM / VLM / MMLM / etc.

[0219] In various embodiments, unsupervised learning can be used to train LLM / VLM / MMLM / etc., where LLM / VLM / MMLM / etc. learns patterns from a large amount of unlabeled text / audio / video / image / design / USD / etc. data. Due to extensive training, in these embodiments, the model may not require task-specific or domain-specific training. An LLM / VLM / MMLM / etc. extensively pre-trained on a large amount of unlabeled data can be considered a base model and can excel at various tasks, such as question answering, summarizing, filling in missing information, translation, and image / video / design / USD / data generation. Some LLM / VLM / MMLM / etc. can be customized for specific use cases using techniques such as cue tuning, fine-tuning, retrieval augmentation generation (RAG), adding adapters (e.g., custom neural networks and / or neural network layers to tune or adjust cues or labels to bias the language model towards a specific task or domain), and / or using optimization models for specific tasks and / or other fine-tuning or customization techniques within a specific domain.

[0220] In some embodiments, the LLM / VLM / MMLM / etc. disclosed herein can be implemented using various model alignment techniques. For example, in some embodiments, guardrails can be implemented to identify incorrect or unwanted inputs (e.g., prompts) and / or outputs of the model. In this process, the system can use guardrails and / or other model alignment techniques to prevent the processing of specific unwanted inputs using LLM / VLM / MMLM / etc., and / or to prevent the output or presentation of information generated by LLM / VLM / MMLM / etc. (e.g., displays, audio outputs, etc.). In some embodiments, one or more additional models (or layers thereof) can be implemented to identify problems with the model's inputs and / or outputs. For example, these "protective" models can be trained to identify "safe" or otherwise okay or desired inputs and / or outputs and / or "unsafe" or otherwise unwanted inputs and / or outputs for a particular application / implementation. Therefore, the LLM / VLM / MMLM / etc. disclosed herein are unlikely to output language / text / audio / video / design data / USD data / etc. that may be offensive, vulgar, inappropriate, insecure, out of scope, and / or unwanted for a particular application / implementation.

[0221] In some embodiments, an LLM / VLM / etc. can be configured or able to access or use one or more plugins, application programming interfaces (APIs), databases, data stores, repositories, etc. For example, for certain tasks or operations where the model is not ideally suited, the model may have instructions for accessing one or more plugins (e.g., third-party plugins) to help process the current input (e.g., as a result of training, and / or based on instructions in a given prompt). In such an example, when at least part of the prompt relates to restaurants or weather, the model can access one or more restaurant or weather plugins (e.g., via one or more APIs) to retrieve relevant information. Another example is that if at least part of the response requires mathematical computation, the model can access one or more mathematical plugins or APIs to help solve the problem, and then the response from the plugins and / or APIs can be used in the model's output. This process can be repeated (e.g., recursively) an arbitrary number of iterations, using any number of plugins and / or APIs, until a response to each query / question / request / process / action / etc. can be generated in response to the input prompt. Therefore, models can rely not only on their own knowledge gained from training on large datasets, but also on the expertise or optimized properties of one or more external resources (such as APIs, plugins, etc.).

[0222] In some embodiments, multiple language models (e.g., LLM / VLM / MMLM / etc., multiple instances of the same language model, and / or multiple hints provided to the same language model or instances of the same language model) can be implemented, executed, or accessed (e.g., using one or more plugins, user interfaces, APIs, databases, data stores, repositories, etc.) to provide output in response to the same query or in response to separate parts of a query. In at least one embodiment, the same input query and hints (e.g., a set of constraints, condition generators, etc.) can be provided to multiple language models (e.g., language models with different architectures, language models trained on different (e.g., updated) data corpora). In one or more embodiments, the language models can be different versions of the same underlying model. In one or more embodiments, at least one language model can be instantiated as multiple agents, for example, providing more than one hint to constrain, guide, or otherwise influence the style, content, or character of the provided output. In one or more exemplary non-limiting embodiments, the same language model can be required to provide output corresponding to different roles, perspectives, characters, or different knowledge bases, as defined by the provided hints.

[0223] In any such embodiment, the outputs of two or more (e.g., each) language models, two or more versions of at least one language model, two or more instantiated proxies of at least one language model, and / or provided to two or more prompts for at least one language model can be further processed, such as aggregated, compared, or filtered, or used to determine (and provide) a consensus response. In one or more embodiments, the output from one language model (or version, instance, or proxy) can be provided as input to another language model for further processing and / or validation. In one or more embodiments, the language model can be required to generate or otherwise obtain output about the input source material, wherein the output is associated with the input source material. This association may include, for example, generating captions or text portions embedded (e.g., as metadata) within the input source text or image. In one or more embodiments, the output of the language model can be used to determine the validity of the input source material for further processing or inclusion in a dataset. For example, the language model can be used to evaluate the presence (or absence) of a target word in a text portion or the presence (or absence) of an object in an image, wherein the text or image is annotated to indicate such presence (or absence). Alternatively, the determination from the language model can be used to determine whether the source material should be included in the curatorial dataset, for example, but not limited to this.

[0224] Figure 9A This is a block diagram of an example generative language model system 900 applicable to implementing at least some embodiments of this disclosure. Figure 9A In the example shown, the generative language model system 900 includes a retrieval augmentation (RAG) component 992, an input processor 905, a tokenizer 910, an embedding component 920, a plugin / API 995, and a generative language model (LM) 930 (which may include LLM, VLM, multimodal LM, etc.).

[0225] At a high level, the input processor 905 can receive input 901, which includes text and / or other types of input data (e.g., audio data, video data, image data, sensor data (e.g., LiDAR, RADAR, ultrasound, etc.), 3D design data, CAD data, generic scene descriptor (USD) data (e.g., OpenUSD, etc.), depending on the architecture of the generative LM 930 (e.g., LLM / VLM / MMLM, etc.). In some embodiments, input 901 includes plain text in the form of one or more sentences, paragraphs, and / or documents. Additionally or alternatively, input 901 may include numerical sequences, pre-computed embeddings (e.g., word or sentence embeddings), and / or structured data (e.g., tabular format, JSON, or XML). In the generative LM In some implementations of 930 capable of handling multimodal input, input 901 can combine text (or text that may be omitted) with image data, audio data, video data, design data, USD data, and / or other types of input data (e.g., but not limited to the data described herein). Taking raw input text as an example, input processor 905 can prepare the raw input text in various ways. For example, input processor 905 can perform various types of text filtering to remove noise from relevant text content (e.g., special characters, punctuation marks, HTML tags, stop words, portions of images, portions of audio, etc.). In examples involving stop words (common words that often have little semantic meaning), input processor 905 can remove stop words to reduce noise and allow the generative LM 930 to focus on more meaningful content. Input processor 905 can apply text normalization, for example, by converting all characters to lowercase, removing accent marks, and / or handling special cases (such as abbreviations or abbreviations) to ensure consistency. These are just a few examples; other types of input processing can be applied.

[0226] In some embodiments, RAG component 992 (which may include one or more RAG models, and / or may be performed using generative LM 930 itself) may be used to retrieve additional information to be used as part of input 901 or a prompt. RAGs can be used to enhance input to LLM / VLM / MMLM / etc. with external knowledge to make the answer to a specific question or query or request more relevant, such as where specific knowledge is required. RAG component 992 may obtain this additional information from one or more external sources (e.g., basic information such as basic text / images / videos / audio / USD / CAD / etc.), which can then be fed along with the prompt to LLM / VLM / MMLM / etc. to improve the accuracy of the model's response or output.

[0227] For example, in some embodiments, in addition to the data retrieved using RAG component 992, input 901 may also be generated using query or model input (e.g., questions, requests, etc.). In some embodiments, input processor 905 may analyze input 901 and communicate with RAG component 992 (or in embodiments, RAG component 992 may be part of input processor 905) to identify relevant text and / or other data to provide to generative LM 930 as additional context or information source, typically from which to identify responses, answers, or outputs 990. For example, when the input indicates that a user is interested in the required tire pressure for a particular brand and model of vehicle, RAG component 992 may use a RAG model, for example, to perform a vector search in the embedding space to retrieve tire pressure information or its corresponding text from a digital (embedded) version of the owner's manual for that particular vehicle brand and model. Similarly, when a user revisits the chatbot related to a specific product sale or service, the RAG component 992 can retrieve previously stored conversation history (or at least its summary) and provide the previous conversation history, along with the current inquiry / request, as part of the generative LM 930 as input 901.

[0228] RAG component 992 can use various RAG techniques. For example, it can use naive RAG ( The document is indexed, chunked, and applied to an embedding model to generate embeddings corresponding to chunks. User queries can also be applied to this embedding model and / or another embedding model of the RAG component 992, and the embeddings of the chunks can be compared with the embeddings of the query to identify the most similar / relevant embeddings to the query. These most similar / relevant embeddings can be provided to the generative LM 930 to generate output.

[0229] In some embodiments, more advanced RAG techniques can be used. For example, chunks can undergo pre-retrieval processes (e.g., routing, rewriting, metadata analysis, expansion, etc.) before being passed to the embedding model. Furthermore, post-retrieval processes (e.g., re-ranking, hint compression, etc.) can be performed on the output of the embedding model before generating the final embedding, which is then used for comparison with the input query.

[0230] As a further example, modular RAG techniques can be used, such as those similar to Naive RAG and / or Advanced RAG, but also including features such as hybrid search, recursive retrieval and query engines, StepBack methods, subqueries and hypothetical document embeddings.

[0231] As another example, Graph RAG can use a knowledge graph as a source of context or factual information. Graph RAG can be implemented using a graph database as a source of contextual information sent to LLM / VLM / MMLM / etc. Instead of providing the model with data chunks extracted from larger documents (which may result in a lack of context, factual accuracy, linguistic accuracy, etc.) (or anything other than providing the model with data chunks extracted from larger documents), Graph RAG can also provide structured entity information to LLM / VLM / MMLM / etc. by combining structured entity text descriptions with their many attributes and relationships, thus giving the model deeper insights. In implementing Graph RAG, the systems and methods described herein use graphs as content stores and extract relevant document chunks, requiring LLM / VLM / MMLM / etc. to use them to answer questions. In such embodiments, the knowledge graph may contain relevant textual content and metadata about the knowledge graph, or it may be integrated with a vector database. In some embodiments, Graph RAG can use the graph as a subject matter expert, where descriptions of concepts and entities relevant to the query / hint can be extracted and passed to the model as semantic context. These descriptions may include relationships between concepts. In other examples, the graph can be used as a database where a portion of a query / hint can be mapped to a graph query, the graph query can be executed, and LLM / VLM / MMLM / etc. can aggregate the results. In such examples, the graph can store relevant factual information and can be used for queries (natural language queries) and entity links to graph query tools (NL to graph query tools). In some embodiments, the graph RAG (e.g., using a graph database) can be combined with standard (e.g., vector database) RAGs and / or other RAG types to benefit from a variety of approaches.

[0232] In any embodiment, the RAG component 992 can implement plugins, APIs, user interfaces, and / or other functions to perform RAG. For example, LLM / VLM / MMLM / etc. can use graph RAG plugins to run queries on knowledge graphs to extract relevant information to feed into the model, and can use standard or vector RAG plugins to run queries on vector databases. For example, the graph database can interact with the plugin's REST interface, thus decoupling the graph database from the vector database and / or the embedded model.

[0233] The tokenizer 910 can segment (e.g., processed) text data into smaller units (tags) for subsequent analysis and processing. Depending on the implementation, the tags can represent individual words, sub-words, characters, audio / video / images, etc. Word-based tokenization divides the text into individual words, treating each word as a separate tag. Sub-word tokenization breaks words down into smaller meaningful units (e.g., prefixes, suffixes, stems), enabling the generative LM 930 to understand morphological changes and process words outside the vocabulary more effectively. Character-based tokenization represents each character as a separate tag, enabling the generative LM 930 to process text at a fine-grained level. The choice of tokenization strategy can depend on factors such as the language being processed, the task at hand, and / or the characteristics of the training dataset. Therefore, the tokenizer 910 can transform (e.g., processed) text into a structured format according to the tokenization scheme implemented in a particular embodiment.

[0234] Embedding component 920 can use any known embedding technique to transform discrete tokens into semantically meaningful (e.g., dense, continuous vector) representations. For example, embedding component 920 can use pre-trained word embeddings (e.g., Word2Vec, GloVe, or FastText), one-hot encoding, Term Frequency-Inverse Document Frequency (TF-IDF) encoding, one or more embedding layers of a neural network, and / or others.

[0235] In some implementations where input 901 includes image data / video data, etc., input processor 901 may resize the data to a standard size compatible with the format of the corresponding input channel and / or normalize pixel values ​​to a common range (e.g., 0 to 1) to ensure consistent representation, and embedding component 920 may encode the image data using any known technique (e.g., using one or more convolutional neural networks (CNNs) to extract visual features). In some implementations where input 901 includes audio data, input processor 901 may resample the audio file to a consistent sampling rate for uniform processing, and embedding component 920 may use any known technique to extract and encode audio features, such as in the form of a spectrogram (e.g., a Mel spectrogram). In some implementations where input 901 includes video data, input processor 901 may extract frames or apply resizing to extracted frames, and embedding component 920 may extract features such as optical flow embedding or video embedding and / or encode temporal information or frame sequences. In some implementations where input 901 includes multimodal data, the embedded component 920 can use techniques such as early fusion (concatenation), late fusion (sequential processing), and attention-based fusion (e.g., self-attention, cross-attention) to fuse representations of different types of data (e.g., text, images, audio, USD, video, design, etc.).

[0236] Other components of the generative LM 930 and / or generative LM system 900 may use different types of neural network architectures depending on the implementation scheme. For example, a transducer-based architecture (e.g., the architecture used in models such as GPT) may be implemented, and it may include a self-attention mechanism that weights the importance of different words or tokens in the input sequence and / or a feedforward network that processes the output of the self-attention layer, applying a nonlinear transformation to the input representation and extracting higher-level features. Some non-limiting example architectures include transducers (e.g., encoder-decoder, decoder-only, multimodal), RNNs, LSTMs, fusion models, diffusion models, cross-modal embedding models that learn a joint embedding space, graph neural networks (GNNs), hybrid architectures that combine different types of adversarial networks (such as generative adversarial networks or GANs or adversarial autoencoders (AAEs) for joint distribution learning), etc. Therefore, depending on the implementation scheme and architecture, the embedded component 920 can apply the encoded representation of the input 901 to the generative LM 930, and the generative LM 930 can process the encoded representation of the input 901 to generate an output 990, which may include response text and / or other types of data.

[0237] As described herein, in some embodiments, the generative LM 930 may be configured to access or use (or be able to access or use) plugins / APIs 995 (which may include one or more plugins, application programming interfaces (APIs), databases, data stores, repositories, etc.). For example, for certain tasks or operations where the generative LM 930 is not ideally suited, the model may have instructions (e.g., as a result of training, and / or based on instructions in a given prompt, such as instructions retrieved using RAG component 992) to access one or more plugins / APIs 995 (e.g., third-party plugins) to help process the current input. In such an example, when at least part of the prompt is related to a restaurant or weather, the model may access one or more restaurant or weather plugins (e.g., via one or more APIs), sending at least part of the prompt related to a particular plugin / API 995 to the plugin / API 995, which can process the information and return an answer to the generative LM 930, which can then use the response to generate output 990. This process can be repeated (e.g., recursively) an arbitrary number of iterations and repeated using any number of plugins / APIs 995 until an output 990 that resolves each query / question / request / process / action / etc. from input 901 is generated. Therefore, the model can rely not only on its own knowledge acquired from training on a large dataset and / or from data retrieved using the RAG component 992, but also on the expertise or optimized properties of one or more external resources (e.g., plugins / APIs 995).

[0238] Figure 9B This is a block diagram of an example implementation scheme, where the generative LM 930 includes a converter encoder-decoder. For example, suppose the input text (e.g., “Who discovered gravity”) is tokenized (e.g., by...) Figure 9A The tokenizer 910) is used for tokens such as words, and each token is encoded (e.g., by...). Figure 9A The embedding component 920 is a corresponding embedding (e.g., of size 512). Since these token embeddings do not typically represent the position of the tokens in the input sequence, positional encoding can be added to each token embedding using any known technique to encode the order relation and context of the tokens in the input sequence. Thus, (e.g., the resulting) embeddings can be applied to one or more encoders 935 of the generative LM 930.

[0239] In the example implementation, encoder 935 forms an encoder stack, where each encoder includes a self-attention layer and a feedforward network. In the example converter architecture, each token (e.g., a word) flows through a separate path. Therefore, each encoder can accept a sequence of vectors, pass each vector through the self-attention layer, then through the feedforward network, and then up to the next encoder in the stack. Any known self-attention technique can be used. For example, to compute a self-attention score for each token (word), a query vector, a key vector, and a value vector can be created for each token. The self-attention score for a token pair can be computed by taking the dot product of the query vector and the corresponding key vector, normalizing the resulting score, multiplying by the corresponding value vector, and summing the weighted value vectors. The encoder can apply multi-head attention, where the attention mechanism is applied multiple times in parallel with different learned weight matrices. Any number of encoders can be cascaded to generate a context vector encoding the input. Attention projection layer 940 can transform the context vector into attention vectors (keys and values) for decoder 945.

[0240] In the example implementation, decoder 945 forms a decoder stack, where each decoder includes a self-attention layer, an encoder-decoder self-attention layer that uses attention vectors (keys and values) from the encoder to focus on relevant parts of the input sequence, and a feedforward network. Similar to encoder 935, in the example converter architecture, each token (e.g., a word) flows through a separate path in decoder 945. During the first pass, decoder 945, classifier 950, and generation mechanism 955 can generate a first token, and generation mechanism 955 can apply the generated token as input during a second pass. This process can be repeated cyclically, generating tokens (e.g., words) and adding them to the output of the previous pass, and in subsequent passes applying token embeddings of positionally encoded composite sequences as input to decoder 945, generating one token at a time (called autoregression) until a symbol or token indicating the end of the response is predicted. In each decoder, the self-attention layer is typically restricted to focusing only on earlier positions in the output sequence by applying a masking technique (e.g., setting future positions to negative infinity) before the softmax operation. In the example implementation, the encoder-decoder attention layer operates similarly to the (e.g., multi-head) self-attention operation in encoder 935, except that it creates its queries from the layers below it and obtains keys and values ​​(e.g., matrices) from the output of encoder 935.

[0241] Therefore, decoder 945 can output some decoded (e.g., vector) representation of the input applied during a particular pass. Classifier 950 can include a multi-class classifier comprising one or more neural network layers and a softmax operation that transforms logit probabilities into probabilities, the neural network layers projecting the decoded (e.g., vector) representation onto corresponding dimensions (e.g., one dimension for each supported word or token in the output vocabulary). Thus, generation mechanism 955 can select or sample words or tokens based on corresponding predicted probabilities (e.g., selecting the word with the highest predicted probability) and append it to the output of the previous pass, thereby generating each word or token sequentially. Generation mechanism 955 can repeat this process, triggering successive decoder inputs and corresponding predictions until a symbol or token representing the end of the response is selected or sampled, at which point generation mechanism 955 can output the generated response.

[0242] Figure 9C This is a block diagram of an example implementation where the generative LM 930 includes a decoder-only converter architecture. For example, Figure 9C The decoder 960 can be used with Figure 9B The decoder 945 operates similarly, except... Figure 9C Each decoder 960 omits the encoder-decoder self-attention layer (because there is no encoder in this implementation). Therefore, decoders 960 can form a decoder stack, where each decoder includes a self-attention layer and a feedforward network. Furthermore, instead of encoding the input sequence, a symbol or tag indicating the end of the input sequence (or the beginning of the output sequence) can be appended to the input sequence, and the resulting sequence (e.g., a corresponding embedding with positional encoding) can be applied to decoder 960. Figure 9B Similar to decoder 945, each tag (e.g., a word) can flow through a separate path in decoder 960, and decoder 960, classifier 965, and generation mechanism 970 can use autoregression to generate one tag at a time sequentially until a symbol or tag indicating the end of the response is predicted. Classifier 965 and generation mechanism 970 can be combined with... Figure 9B The classifier 950 and the generation mechanism 955 operate similarly, wherein the generation mechanism 970 selects or samples each consecutive output label based on the corresponding predicted probability and appends it to the output of the previous iteration, generating each label sequentially until a symbol or label representing the end of the response is selected or sampled. The architectures described herein, and others, are merely examples, and other suitable architectures may be implemented within the scope of this disclosure.

[0243] Example network environment

[0244] A network environment suitable for implementing embodiments of this disclosure may include one or more client devices, servers, network attached storage (NAS), other backend devices, and / or other device types. Client devices, servers, and / or other device types (e.g., each device) may... Figure 7 Implemented on one or more instances of computing devices 700—for example, each device may include similar components, features, and / or functions of computing device 700. Additionally, backend devices (servers, NAS, etc.) may be included as part of data center 800, examples of which are referred to herein. Figure 8 To describe in more detail.

[0245] Components of a network environment can communicate with each other through one or more networks, whether wired, wireless, or a combination of both. This network can include multiple networks, or networks of networks. For example, the network can include one or more wide area networks (WANs), one or more local area networks (LANs), one or more public networks (such as the Internet and / or the Public Switched Telephone Network (PSTN)), and / or one or more private networks. In cases where the network includes a wireless telecommunications network, components such as base stations, communication towers, or even access points (along with other components) can provide wireless connectivity.

[0246] A compatible network environment may include one or more peer-to-peer network environments—in which case the server may not be included in the network environment—and one or more client-server network environments—in which case one or more servers may be included in the network environment. In a peer-to-peer network environment, the server functionality described herein can be implemented on any number of client devices.

[0247] In at least one embodiment, the network environment may include one or more cloud-based network environments, distributed computing environments, and combinations thereof. The cloud-based network environment may include a framework layer, a job scheduler, a resource manager, and a distributed file system implemented on one or more servers, which may include one or more core network servers and / or edge servers. The framework layer may include a framework supporting software at the software layer and / or application at the application layer. The software or one or more applications may respectively include web-based service software or applications. In embodiments, one or more client devices may use web-based service software or applications (e.g., by accessing the service software and / or applications via one or more application programming interfaces (APIs)). The framework layer may be, but is not limited to, a free and open-source software web application framework, for example, one that can use a distributed file system for large-scale data processing (e.g., "big data").

[0248] A cloud-based network environment can provide cloud computing and / or cloud storage for any combination of the computing and / or data storage functions (or one or more of them) described herein. Any of these various functions can be distributed from a central or core server (e.g., servers in one or more data centers) to multiple locations, which may be located in a state, region, country, globally, etc. If the connection to a user (e.g., a client device) is relatively close to one or more edge servers, one or more core servers may assign at least a portion of the functionality to one or more edge servers. A cloud-based network environment can be private (e.g., limited to a single organization), public (e.g., available to many organizations), and / or a combination thereof (e.g., a hybrid cloud environment).

[0249] One or more client devices may include the information described in this article. Figure 7 The client device may be, by way of example and not limitation, at least some of the components, features and functions of one or more example computing devices 700. As an example and not a limitation, the client device may be a personal computer (PC), laptop computer, mobile device, smartphone, tablet computer, smartwatch, wearable computer, personal digital assistant (PDA), MP3 player, virtual reality headset, global positioning system (GPS) or device, video player, camera, surveillance equipment or system, vehicle, ship, aircraft, virtual machine, drone, robot, handheld communication device, hospital equipment, gaming device or system, entertainment system, vehicle computer system, embedded system controller, remote control, device, consumer electronics device, workstation, edge device, any combination of these depicted devices or any other suitable device.

[0250] 1. A computer-implemented method, the method comprising: obtaining one or more first sensor values ​​associated with one or more sensors of a robot system using one or more computing devices associated with the robot system; determining, using a machine learning model and based at least on the one or more sensor values, that the robot system requires intervention; obtaining one or more second sensor values ​​associated with the intervention using the one or more computing devices; generating an intervention event based at least on the one or more first sensor values ​​and the one or more second sensor values; and modifying an intervention event database based at least on the generated intervention event.

[0251] 2. The computer-implemented method of claim 1, further comprising generating an alarm, wherein the alarm includes a request for intervention.

[0252] 3. The computer-implemented method of claim 1, wherein the one or more sensor values ​​include one or more of image sensor data, LiDAR sensor data, RADAR sensor data, SONAR sensor data, ultrasonic sensor data, IMU sensor data, or infrared sensor data.

[0253] 4. The computer-implemented method of claim 1, wherein the one or more second sensor values ​​include one or more of position, motion, orientation, stress, strain, or torque exhibited or experienced by one or more motor control components included in the robot system.

[0254] 5. The computer-implemented method of claim 1, further comprising using a machine learning model to generate a natural language description associated with the intervention.

[0255] 6. The computer-implemented method of claim 1 further includes using a machine learning model to classify the intervention into one or more categories or one or more failure modes.

[0256] 7. The computer-implemented method of claim 1, further comprising generating one or more entries in the training or validation dataset based at least on the intervention.

[0257] 8. The computer-implemented method of claim 7, further comprising initiating automatic retraining of at least one or more functions of the autonomous software stack based on the training or validation dataset, wherein the autonomous software stack includes one or more autonomous or semi-autonomous control routines associated with the robotic system.

[0258] 9. The computer-implemented method of claim 1, wherein the method is performed by at least one of the following: a control system for an autonomous or semi-autonomous machine; a perception system for an autonomous or semi-autonomous machine; a system for performing simulated operations; a system for performing digital twin operations; a system for performing optical transmission simulation; a system for performing collaborative content creation for 3D assets; a system for performing deep learning operations; a system for performing remote operations; a system for performing real-time streaming; a system for generating or presenting one or more of augmented reality content, virtual reality content, or mixed reality content; a system implemented using an edge device; a system implemented using a robot; a system for performing conversational AI operations; a system for implementing one or more multimodal language models; a system for implementing one or more large language models (LLMs); a system for implementing one or more visual language models (VLMs); a system for generating synthetic data; a system for generating synthetic data using AI; a system comprising one or more virtual machines (VMs); a system implemented at least partially in a data center; or a system implemented at least partially using cloud computing resources.

[0259] 10. One or more processors, including processing circuitry, the processing circuitry being configured to: obtain one or more first sensor values ​​associated with one or more sensors of the robot using one or more computing devices associated with the robot; determine, using a machine learning model and based on the one or more first sensor values, that the robot requires intervention; obtain one or more second sensor values ​​associated with the intervention using the one or more computing devices; generate an intervention event based at least on the one or more first sensor values ​​and the one or more second sensor values; and modify an intervention event database based at least on the generated intervention event.

[0260] 11. The processor of claim 10 or more, wherein the one or more sensor values ​​include one or more of image sensor data, LiDAR sensor data, RADAR sensor data, SONAR sensor data, ultrasonic sensor data, IMU sensor data, or infrared sensor data.

[0261] 12. The processor of claim 10, wherein the one or more second sensor values ​​include one or more of position, motion, orientation, stress, strain, or torque exhibited or experienced by one or more motor control components of the robot.

[0262] 13. One or more processors as claimed in claim 10, wherein the processing circuitry further uses a machine learning model to generate a natural language description associated with the intervention.

[0263] 14. The processor of claim 10 or more, wherein the processing circuitry further uses a machine learning model to classify the intervention into one or more categories or one or more failure modes.

[0264] 15. One or more processors as claimed in claim 10, wherein the processing circuitry further generates one or more entries in the training or validation dataset based at least on the intervention.

[0265] 16. The processor of claim 15, wherein the processing circuitry further initiates automatic retraining of at least a portion of the autonomous software stack based on the training or validation dataset, and wherein the autonomous software stack includes one or more autonomous or semi-autonomous control routines associated with the robot.

[0266] 17. The processor of claim 10 or more, wherein the processor is included in at least one of: a control system for an autonomous or semi-autonomous machine; a perception system for an autonomous or semi-autonomous machine; a system for performing analog operations; a system for performing digital twin operations; a system for performing optical transmission simulation; a system for performing collaborative content creation for 3D assets; a system for performing deep learning operations; a system for performing remote operations; a system for performing real-time streaming; a system for generating or presenting one or more of augmented reality content, virtual reality content, or mixed reality content; a system implemented using an edge device; a system implemented using a robot; a system for performing conversational AI operations; a system for implementing one or more multimodal language models; a system for implementing one or more large language models (LLMs); a system for implementing one or more visual language models (VLMs); a system for generating synthetic data; a system for generating synthetic data using AI; a system comprising one or more virtual machines (VMs); a system implemented at least partially in a data center; or a system implemented at least partially using cloud computing resources.

[0267] 18. A teleoperation system comprising: one or more processors for sending one or more messages to a deployed robot to cause the deployed robot to perform one or more operations associated with a determined intervention event, the intervention event being determined based at least on sensor data obtained using one or more sensors of the deployed robot processed by one or more multimodal language models, wherein the one or more messages are generated based at least on one or more inputs to the teleoperation system, the one or more inputs being received in response to the presentation of event information determined using the one or more multimodal language models and corresponding to the intervention event.

[0268] 19. The teleoperation system of claim 18, wherein the teleoperation system comprises at least one of: a control system for an autonomous or semi-autonomous machine; a sensing system for an autonomous or semi-autonomous machine; a system for performing simulation operations; a system for performing digital twin operations; a system for performing optical transmission simulation; a system for performing collaborative content creation for 3D assets; a system for performing deep learning operations; a system for performing remote operations; a system for performing real-time streaming; a system for generating or presenting one or more of augmented reality content, virtual reality content, or mixed reality content; a system implemented using an edge device; a system implemented using a robot; a system for performing conversational AI operations; a system for implementing one or more multimodal language models; a system for implementing one or more large language models (LLMs); a system for implementing one or more visual language models (VLMs); a system for generating synthetic data; a system for generating synthetic data using AI; a system comprising one or more virtual machines (VMs); a system implemented at least partially in a data center; or a system implemented at least partially using cloud computing resources.

[0269] 20. The teleoperation system of claim 18, wherein the one or more operations include at least one of: navigating to a location indicated in the one or more messages, shutting down one or more components of the deployed robot, changing the operational state of the deployed robot, or causing an indication of a changed operational state of the deployed robot to be presented. This disclosure can be described in the general context of computer code or machine-usable instructions, including computer-executable instructions, such as program modules, that are executed by a computer or other machine (such as a personal data assistant or other handheld device). Typically, program modules include routines, programs, objects, components, data structures, etc., and refer to code that performs a specific task or implements a specific abstract data type. This disclosure can be implemented in a variety of system configurations, including handheld devices, consumer electronics, general-purpose computers, and more specialized computing devices. This disclosure can also be implemented in a distributed computing environment, where tasks are performed by remote processing devices linked via a communication network.

[0270] This disclosure can be described in the general context of computer code or machine-usable instructions, including computer-executable instructions, such as program modules, that are executed by a computer or other machine (such as a personal data assistant or other handheld device). Typically, program modules include routines, programs, objects, components, data structures, etc., and refer to code that performs a specific task or implements a specific abstract data type. This disclosure can be implemented in a variety of system configurations, including handheld devices, consumer electronics, general-purpose computers, and more specialized computing devices. This disclosure can also be implemented in distributed computing environments, where tasks are performed by remote processing devices linked via a communication network.

[0271] As used herein, the phrase “and / or” relating to two or more elements should be interpreted as referring to only one element or a combination of elements. For example, “element A, element B, and / or element C” may include only element A, element B, element C, element A and element B, element A and element C, element B and element C, or element A, element B, and element C. Furthermore, “at least one of element A or element B” may include at least one of element A, at least one of element B, or at least one of element A and at least one of element B. Additionally, “at least one of element A and element B” may include at least one of element A, at least one of element B, or at least one of element A and at least one of element B.

[0272] To meet legal requirements, the subject matter of this disclosure has been described in detail herein. However, the description itself is not intended to limit the scope of this disclosure. Rather, the inventors have envisioned that the claimed subject matter may also be embodied in other ways, including different steps or combinations of steps similar to those described herein, as well as other existing or future techniques. Furthermore, although the terms “step,” “operation,” and / or “block” may be used herein to imply different elements of the method employed, they should not be construed as implying any particular order between the steps disclosed herein unless the order of the individual steps is explicitly described.

Claims

1. A computer-implemented method, the method comprising: Use one or more computing devices associated with the robot system to obtain one or more first sensor values ​​associated with one or more sensors of the robot system; Using a machine learning model and based at least on the values ​​of one or more of the sensors, determine that the robotic system requires intervention; The one or more computing devices are used to obtain one or more second sensor values ​​associated with the intervention; An intervention event is generated based on at least one or more first sensor values ​​and one or more second sensor values; as well as The intervention event database should be modified based on at least the generated intervention events.

2. The computer-implemented method as described in claim 1, further comprising generating an alarm, wherein, The alert includes a request for intervention.

3. The computer-implemented method as described in claim 1, wherein, The one or more sensor values ​​include one or more of the following: image sensor data, LiDAR sensor data, RADAR sensor data, SONAR sensor data, ultrasonic sensor data, IMU sensor data, or infrared sensor data.

4. The computer-implemented method as described in claim 1, wherein, The one or more second sensor values ​​include one or more of the position, motion, orientation, stress, strain, or torque exhibited or experienced by one or more motor control components included in the robot system.

5. The computer-implemented method of claim 1, further comprising using a machine learning model to generate a natural language description associated with the intervention.

6. The computer-implemented method of claim 1 further includes using a machine learning model to classify the intervention into one or more categories or one or more failure modes.

7. The computer-implemented method of claim 1, further comprising generating one or more entries in the training or validation dataset based at least on the intervention.

8. The computer-implemented method of claim 7, further comprising automatically retraining at least one or more functions of the autonomous software stack based on the training or validation dataset, wherein, The autonomous software stack includes one or more autonomous or semi-autonomous control routines associated with the robotic system.

9. The computer-implemented method as described in claim 1, wherein, The method is performed by at least one of the following: Control systems for autonomous or semi-autonomous machines; Sensing systems for autonomous or semi-autonomous machines; A system used to perform simulation operations; Systems used to perform digital twin operations; A system for performing optical transmission simulation; A system for creating collaborative content for 3D assets; A system used to perform deep learning operations; A system used to perform remote operations; Systems used for performing real-time streaming; A system for generating or presenting one or more of augmented reality content, virtual reality content, or mixed reality content; Systems implemented using edge devices; Systems implemented using robots; A system for performing conversational AI operations; A system that implements one or more multimodal language models; A system that implements one or more large language model LLMs; A system that implements one or more Visual Language Models (VLMs); A system for generating synthetic data; Systems for generating synthetic data using AI; A system containing one or more virtual machines (VMs); A system that is at least partially implemented in a data center; or A system that utilizes cloud computing resources at least in part.

10. One or more processors, including processing circuitry, said processing circuitry being used to: Use one or more computing devices associated with the robot to obtain one or more first sensor values ​​associated with one or more sensors of the robot; Using a machine learning model and based on one or more first sensor values, it is determined that the robot requires intervention; The one or more computing devices are used to obtain one or more second sensor values ​​associated with the intervention; An intervention event is generated based on at least one or more first sensor values ​​and one or more second sensor values; as well as The intervention event database should be modified based on at least the generated intervention events.

11. The processors of claim 10, wherein, The one or more sensor values ​​include one or more of the following: image sensor data, LiDAR sensor data, RADAR sensor data, SONAR sensor data, ultrasonic sensor data, IMU sensor data, or infrared sensor data.

12. The processors of claim 10, wherein, The one or more second sensor values ​​include one or more of the position, motion, orientation, stress, strain, or torque exhibited or experienced by one or more motor control components of the robot.

13. The one or more processors as claimed in claim 10, wherein, The processing circuitry also uses machine learning models to generate natural language descriptions associated with the intervention.

14. The processors of claim 10, wherein, The processing circuit also uses a machine learning model to classify the intervention into one or more categories or one or more failure modes.

15. One or more processors as claimed in claim 10, wherein, The processing circuitry also generates one or more entries in the training or validation dataset based on the intervention.

16. The one or more processors as claimed in claim 15, wherein, The processing circuitry also initiates automatic retraining of at least a portion of the autonomous software stack based on the training or validation dataset, wherein the autonomous software stack includes one or more autonomous or semi-autonomous control routines associated with the robot.

17. One or more processors as claimed in claim 10, wherein, The one or more processors are included in at least one of the following: Control systems for autonomous or semi-autonomous machines; Sensing systems for autonomous or semi-autonomous machines; A system used to perform simulation operations; Systems used to perform digital twin operations; A system for performing optical transmission simulation; A system for creating collaborative content for 3D assets; A system used to perform deep learning operations; A system used to perform remote operations; Systems used for performing real-time streaming; A system for generating or presenting one or more of augmented reality content, virtual reality content, or mixed reality content; Systems implemented using edge devices; Systems implemented using robots; A system for performing conversational AI operations; A system that implements one or more multimodal language models; A system that implements one or more large language model LLMs; A system that implements one or more Visual Language Models (VLMs); A system for generating synthetic data; Systems for generating synthetic data using AI; A system containing one or more virtual machines (VMs); A system that is at least partially implemented in a data center; or A system that utilizes cloud computing resources at least in part.

18. A teleoperation system, the teleoperation system comprising: One or more processors are configured to send one or more messages to a deployed robot to cause the deployed robot to perform one or more operations associated with a determined intervention event, the intervention event being determined based at least on sensor data obtained using one or more sensors of the deployed robot processed by one or more multimodal language models, wherein the one or more messages are generated based at least on one or more inputs to the teleoperation system, the one or more inputs being received in response to the presentation of event information determined using the one or more multimodal language models and corresponding to the intervention event.

19. The teleoperation system as described in claim 18, wherein, The teleoperation system is included in at least one of the following: Control systems for autonomous or semi-autonomous machines; Sensing systems for autonomous or semi-autonomous machines; A system used to perform simulation operations; Systems used to perform digital twin operations; A system for performing optical transmission simulation; A system for creating collaborative content for 3D assets; A system used to perform deep learning operations; A system used to perform remote operations; Systems used for performing real-time streaming; A system for generating or presenting one or more of augmented reality content, virtual reality content, or mixed reality content; Systems implemented using edge devices; Systems implemented using robots; Systems used to perform conversational AI operations; A system that implements one or more multimodal language models; A system that implements one or more large language model LLMs; A system that implements one or more Visual Language Models (VLMs); A system for generating synthetic data; Systems for generating synthetic data using AI; A system containing one or more virtual machines (VMs); A system that is at least partially implemented in a data center; or A system that utilizes cloud computing resources at least in part.

20. The teleoperation system as described in claim 18, wherein, The one or more operations include at least one of the following: navigating to a location indicated in the one or more messages, shutting down one or more components of the deployed robot, changing the operational state of the deployed robot, or causing an indication of a changed operational state of the deployed robot to be displayed.

Citation Information

Patent Citations

  • Method for programmable timeouts of tree traversal mechanisms in hardware

    US10885698B2