HARDWARE VIRTUALIZATION FOR FAULT MANAGEMENT
Hardware virtualization with task partitioning and redundancy management addresses the challenge of error management in diverse applications, enhancing diagnostic coverage and usability while maintaining safety integrity levels.
Patent Information
- Application Number
- DE102025133016
- Authority / Receiving Office
- DE · DE
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-08-22
- Filing Date
- 2025-08-19
- Publication Date
- 2026-02-26
AI Technical Summary
Managing errors in application deployment on hardware with diverse characteristics is challenging due to increased software design complexity and compromised functional safety when implementing redundancy, especially in autonomous or semi-autonomous machines.
Implementing hardware virtualization through task partitioning and redundancy management, using switches to divide tasks into redundant and non-redundant modes, enabling spatial and temporal redundancy based on criteria like performance and integrity levels, and utilizing hardware tools for partitioning without exposing them to applications.
Enhances diagnostic coverage, usability, and programmability while maintaining hardware utilization and safety integrity levels, allowing for efficient fault detection and management in diverse applications.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
BACKGROUND
[0001] Hardware parallelism can be used to detect errors associated with application deployment, including in autonomous or semi-autonomous machines. For example, multiple instances of applications can be run redundantly, such as on different hardware components and / or at different times, allowing errors, such as random failures, to occur and be detected. However, the diverse characteristics of applications can make it difficult to manage errors in a way that achieves application integrity levels. Furthermore, implementing redundancy may require redundancy functionality to be exposed at a low level of the target hardware, at the application and / or user level, which can increase software design complexity and compromise functional safety. SUMMARY
[0002] The invention is defined in the claims. For the purpose of illustrating the invention, aspects and embodiments are described herein that may or may not fall within the scope of protection of the claims.
[0003] Various examples disclose systems and methods related to hardware virtualization for error management. These systems and methods define a task flow for executing tasks on a first partition and a second partition. A processor can contain one or more circuits.The one or more circuits can determine that a first task of a multitude of tasks meets a criterion for execution in a redundant mode, determine a task flow for the execution of the multitude of tasks by assigning a switch before the execution of the first task, wherein the switch causes the one or more circuits to be partitioned into a first partition and a second partition, and the multitude of tasks is executed according to the task flow by executing a first instance of the first task on the first partition and a second instance of the first task on the second partition.
[0004] Implementations of the present disclosure relate to hardware virtualization for fault management. For example, systems and methods are disclosed that facilitate the virtualization of a device for random fault detection. The system can enable greater diagnostic coverage, such as by facilitating the use of spatial redundancy. The system can facilitate usability and programmability, such as by providing the functionality of one or more switches that can be implemented based on user input associated with a task and / or the processing of task features. The system can implement the switch as a high-speed switch to improve performance.The system can be scalable for multiple tasks and / or applications, as well as for multiple redundant and / or non-redundant nodes, based on the processing of an application's tasks into a scalable task flow. The system can enable greater hardware utilization based on a higher workload.
[0005] At least one aspect relates to one or more processors. In various implementations, the one or more processors may contain or consist of one or more circuits. The one or more circuits may determine that a first task of a multitude of tasks meets a criterion for execution in a redundant mode, determine a task flow for executing the multitude of tasks by assigning a switch before the execution of the first task, where the switch causes the one or more circuits to be partitioned into a first partition and a second partition, and the multitude of tasks is executed according to the task flow by executing a first instance of the first task on the first partition and a second instance of the first task on the second partition.
[0006] In various implementations, the switch can be a first switch, and the one or more circuits can determine that a second task of the multitude of tasks depends on the first task and does not meet the criterion for execution in redundant mode, and assign a second switch to the task flow between the execution of the first task and the execution of the second task, the second switch causing the one or more circuits to be non-partitioned.
[0007] In various implementations, one or more circuits can determine that the first task meets the criterion based on at least one feature of the first task received from at least one application containing the first task or from user input regarding the first task. The one or more circuits can also define the task flow as a graph containing a plurality of nodes, where the switch is between a first node for executing a non-redundant task and each of (i) a second node coupled to the first node, where the second node is to execute the first instance of the first task, and (ii) a third node coupled to the first node, where the third node is to execute the second instance of the first task.
[0008] In various implementations, the task flow can specify instructions for one or more circuits to use a hardware tool for partitioning that one or more circuits without exposing the hardware tool to an application assigned to the multitude of tasks. In various implementations, the switch is a first switch, and the one or more circuits can assign a second switch to the task flow to switch the context from the first task to executing a second, redundant task.
[0009] In various implementations, one or more circuits can configure the first partition and the second partition as concurrent multiple contexts, graphics processing unit (GPU) partitions, or multiple instances of a multi-instance GPU (MIG).The one or more processors can be included in at least one of the following: a control system for an autonomous or semi-autonomous machine, a perception system for an autonomous or semi-autonomous machine, a system that includes one or more virtual machines (VMs), a system implemented using a robot, a system implemented using an edge device, a system for generating synthetic data, a system that implements one or more large language models (LLMs), a system that implements one or more visual language models (VLMs), a system that implements one or more multimodal language models, a system for performing virtualization at the operating system (OS) level (e.g.,(using containers), which includes one or more deep learning models optimized for use, software for running the one or more deep learning models and / or telemetry software for evaluating, monitoring and / or verifying the health of the system, a system that uses one or more microservices - such as inference microservices (e.g., NVIDIA NIMs) that perform conversational AI operations, a system for performing deep learning operations, a system for performing simulation operations, a system for performing collaborative content creation for 3D assets, a system for performing digital twin operations, a system for performing light transport simulations, a system that is at least partially implemented in a data center, or a system that is at least partially implemented using cloud computing resources.
[0010] At least one aspect relates to a system. In different implementations, the system may contain one or more processing units and one or more memory units.In various implementations, one or more memory units can store instructions which, when executed by one or more processing units, cause the one or more processing units to perform operations that include: determining that a first task of a multitude of tasks satisfies a criterion for execution in a redundant mode; determining a task flow for executing the multitude of tasks by assigning a switch before the execution of the first task, the switch causing the one or more circuits to be partitioned into a first partition and a second partition; and executing the multitude of tasks according to the task flow by executing a first instance of the first task on the first partition and a second instance of the first task on the second partition.
[0011] In various implementations, the switch can be a first switch and perform operations on one or more processing units that include: determining that a second task of the multitude of tasks depends on the first task and does not meet the criterion for execution in redundant mode, and assigning the task flow between the execution of the first task and the execution of the second task of a second switch, the second switch causing the one or more circuits to be non-partitioned.
[0012] In various implementations, one or more processing units can determine that the first task meets the criterion based on at least one feature of the first task received from at least one application containing the first task or from user input regarding the first task. The one or more processing units can also define the task flow as a graph containing a plurality of nodes, where the switch is between a first node for executing a non-redundant task and each of (i) a second node coupled to the first node, where the second node is to execute the first instance of the first task, and (ii) a third node coupled to the first node, where the third node is to execute the second instance of the first task.
[0013] In various implementations, the task flow can specify instructions for one or more processing units to use a hardware tool for partitioning those units without exposing the hardware tool to an application assigned to the multitude of tasks. In some implementations, the switch is a first switch, and the one or more processing units can assign a second switch to the task flow to switch the context from the first task to executing a second, redundant task.
[0014] In various implementations, one or more processing units can configure the first partition and the second partition as concurrent multiple contexts, graphics processing unit (GPU) partitions, or multiple instances of a multi-instance GPU (MIG).The system may be included in at least one of the following: a control system for an autonomous or semi-autonomous machine, a perception system for an autonomous or semi-autonomous machine, a system that includes one or more virtual machines (VMs), a system implemented using a robot, a system implemented using an edge device, a system for generating synthetic data, a system that includes one or more large language models (LLMs), a system that implements one or more visual language models (VLMs), a system that implements one or more multimodal language models, a system for performing virtualization at the operating system (OS) level (e.g.,(using containers), which includes one or more deep learning models optimized for use, software for running the one or more deep learning models and / or telemetry software for evaluating, monitoring and / or verifying the health of the system, a system that uses one or more microservices - such as inference microservices (e.g. NVIDIA NIMs), a system for performing conversational AI operations, a system for performing deep learning operations, a system for performing simulation operations, a system for performing collaborative content creation for 3D assets, a system for performing digital twin operations, a system for performing light transport simulations, a system that is at least partially implemented in a data center, or a system that is at least partially implemented using cloud computing resources.
[0015] At least one aspect relates to a procedure. The procedure may include: determining that a first task of a plurality of tasks satisfies a criterion for execution in a redundant mode; determining a task flow for executing the plurality of tasks, in which a switch is assigned before the execution of the first task, the switch causing the one or more circuits to be partitioned into a first partition and a second partition; and executing the plurality of tasks according to the task flow by executing a first instance of the first task on the first partition and a second instance of the first task on the second partition.
[0016] In various implementations, the switch can be a first switch and the procedure can further include: determining that a second task of the multitude of tasks depends on the first task and does not meet the criterion for execution in redundant mode, and assigning to the task flow between the execution of the first task and the execution of the second task a second switch, wherein the second switch causes the one or more circuits to be unpartitioned.
[0017] In various implementations, defining the task flow as a graph can contain a multitude of nodes, with the switch between a first node for executing a non-redundant task and each of (i) a second node coupled to the first node, where the second node executes the first instance of the first task, and (ii) a third node coupled to the first node, where the third node executes the second instance of the first task. In different implementations, the first and second partitions can be configured as concurrent multiple contexts, graphics processing unit (GPU) partitions, or multiple instances of a multi-instance GPU (MIG).
[0018] The processors, systems, and / or methods described herein may be implemented by or include at least one of the following: a control system for an autonomous or semi-autonomous machine; a perception system for an autonomous or semi-autonomous machine; a system implemented using a robot; a system for performing deep learning operations; a system implemented using an edge device; a system that includes one or more virtual machines (VMs); a system for performing simulation operations; a system for performing digital twin operations; a system for performing light transport simulations; a system for performing collaborative content creation for 3D assets; a system for generating or presenting at least one type of virtual reality, augmented reality, or mixed reality content.a system for performing conversational AI operations; a system containing one or more large language models (LLMs); a system containing one or more visual language models (VLMs); a system containing one or more multimodal language models; a system for generating synthetic data;a system for performing virtualization at the operating system (OS) level (e.g., using containers) that includes one or more deep learning models optimized for use, software for running the one or more deep learning models and / or telemetry software for evaluating, monitoring, and / or verifying the health of the system, a system that uses one or more microservices—such as inference microservices (e.g., NVIDIA NIMs)—a system that is at least partially deployed in a data center, or a system that is at least partially deployed using cloud computing resources.
[0019] The revelation extends to all novel aspects or features described and / or illustrated herein.
[0020] Further features of the disclosure are characterized by the independent and dependent claims.
[0021] Any feature in one aspect of the disclosure can be applied in any suitable combination to other aspects of the disclosure. In particular, procedural aspects can be applied to apparatus or system aspects, and vice versa.
[0022] Furthermore, features implemented in hardware can be implemented in software and vice versa. Any reference to software and hardware features herein should be interpreted accordingly.
[0023] Each system or device feature described herein can also be provided as a process feature, and vice versa. System and / or device aspects that are functionally described (including means plus functional features) can alternatively be expressed in terms of their corresponding structure, such as a suitably programmed processor and allocated working memory.
[0024] It is also understood that certain combinations of the various features described and defined in each aspect of the revelation can be implemented and / or provided and / or used independently of one another.
[0025] The disclosure also provides computer programs and computer program products comprising software code designed to perform one of the methods described herein when executed on a data processing device and / or to embody one of the device and system features described herein, including one or all component steps of a method.
[0026] The disclosure also provides a computer or computer system (including networked or distributed systems) with an operating system that supports a computer program for carrying out one of the methods described herein and / or for embodying one of the device or system features described herein.
[0027] The disclosure also provides a computer-readable medium on which one or more of the aforementioned computer programs are stored.
[0028] The revelation also provides a signal that carries one or more of the aforementioned computer programs.
[0029] The disclosure extends to methods and / or devices and / or systems as described herein with reference to the accompanying drawings.
[0030] Aspects and embodiments of the disclosure will now be described purely by way of example with reference to the attached drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] The present systems and procedures for hardware use in fault management are described in detail below with reference to the attached drawings, wherein: Fig. 1 a block diagram of an example of a system for performing hardware utilization for fault management according to some implementations of the present disclosure; Fig. 2. A schematic diagram of an example of a task flow through the system of Fig. 1 can be generated and executed, according to some implementations of the present disclosure; Fig. 3 is a flowchart of an example of a fault management procedure according to some implementations of the present disclosure; Fig. 4A is an illustration of an exemplary autonomous vehicle according to some embodiments of the present disclosure; Fig. 4B Exemplary camera locations and fields of view for the exemplary autonomous vehicle of Fig. 4A according to some embodiments of the present disclosure; Fig. 4C a block diagram of an exemplary system architecture for the exemplary autonomous vehicle of Fig. 4A according to some embodiments of the present disclosure; Fig. 4D a system diagram for the communication between the cloud-based server(s) and the example autonomous vehicle of Fig. 4A according to some embodiments of the present disclosure; Fig. 5 is a block diagram of an exemplary computing device suitable for use in the implementation of at least some embodiments of the present disclosure; and Fig. 6 is a block diagram of an exemplary data center suitable for use in the implementation of some embodiments of the present disclosure. DETAILED DESCRIPTION
[0032] Systems and methods relating to hardware utilization for fault management are disclosed, such as for the execution of applications in autonomous or semi-autonomous machines, and for facilitating the virtualization of a device for the detection of random faults. Hardware parallelism can be used for fault detection, including random faults. This can be useful for detecting certain faults in safety-related applications, such as those associated with autonomous or semi-autonomous machines. For achieving safety integrity levels, such as automotive safety integrity levels (ASILs), it can be beneficial for components such as massively parallel systems (e.g., graphics processing units (GPUs)) and computational accelerators (e.g., deep learning accelerators) to utilize the built-in parallelism features of the hardware.In some cases, parallelism is achieved using redundant task execution. However, different applications have varying performance and usage goals, as well as different susceptibility to random errors. Such applications may be expected to run on the same hardware and / or operate concurrently. Therefore, managing random errors and achieving target integrity levels can be challenging. Furthermore, while the hardware may possess features and / or functions to facilitate redundancy, exposing such features to applications and / or users can lead to increased software design complexity and / or compromise integrity.
[0033] Systems and methods according to this disclosure can manage the execution of tasks on the target hardware to enable greater diagnostic coverage with respect to faults, such as by selectively implementing temporal redundancy (e.g., sequential execution of redundant workloads) and / or spatial redundancy (e.g., execution of workloads in different hardware units) based on one or more criteria (e.g., performance requirements; integrity levels) for the workloads. For example, the system can retrieve a clue about a task flow that contains one or more processing tasks. The system can update the task flow based on the one or more criteria, such as to assign one or more redundancy operations between the processing tasks of the task flow.This can include, for example, ensuring that two tasks are performed with temporal or spatial redundancy based on the criteria.
[0034] The system can assign an operating mode switch to points in the task flow to implement redundancy, such as assigning the mode switch based on predefined criteria. The operating mode switch can cause the hardware to be partitioned into sections (e.g., virtual sections) to achieve spatial redundancy (which can lead to greater diagnostic coverage, although keeping the hardware in partitioned mode may impair the performance of non-safety-related tasks). The system can execute a first task using a first section and a second task using a second section. Upon detecting the completion of both tasks, the system can cause the hardware to switch to an unpartitioned mode to execute a third task (e.g., to allow full hardware utilization).The system can detect one or more dependencies between tasks to determine the task flow (e.g., as a graph with nodes corresponding to the sequence of operations in the task flow and the hardware resources to be used for the task flow). The system can generate the task flow, modified based on one or more criteria for implementing redundancies, to logically correspond to the task flow without modification.
[0035] Operating mode switches can be exposed to an application and / or user at various levels of abstraction. For example, a library of operating mode switches can be provided for software configuration. The system can retrieve a hint regarding one or more criteria related to the tasks (e.g., based on user input and / or parsing the software code) and determine whether an operating mode switch should be assigned to the task flow according to the hint.
[0036] The system can provide output that represents runtime diagnostics of task execution, such as displaying errors detected during task execution. The system can also generate output to indicate security profiles versus performance (e.g., based on varying redundancy usage), allowing a user to select a profile for runtime operation.
[0037] Although the present disclosure relates to an exemplary autonomous or semi-autonomous vehicle or autonomous machine 400 (e.g., “Vehicle 400”, “Ego-Vehicle 400”, “Machine 400” or “Ego-Machine 400”), one example of which relates to the Fig. The fact that the systems and procedures described herein can be described (as described in sections 4A-4D) is not to be understood as a limitation. For example, the systems and procedures described herein can be used without restriction by non-autonomous vehicles or machines, semi-autonomous vehicles or machines (e.g., in one or more adaptive driver assistance systems (ADAS)), autonomous vehicles or machines, steered and unsteered robots or robotic platforms, warehouse vehicles, all-terrain vehicles, vehicles coupled to one or more trailers, flying vehicles, boats, shuttles, emergency response vehicles, motorcycles, electric or motorized bicycles, aircraft, construction vehicles, underwater vehicles, drones, and / or other types of vehicles.Furthermore, systems and methods according to the present disclosure can be used for monitoring the health status and / or fault conditions of a multi-sensor fusion system for autonomous or semi-autonomous driving and active safety systems, as well as in applications for augmented reality, virtual reality, mixed reality, robotics, security and monitoring, autonomous or semi-autonomous machines and / or in any other technology field in which a multi-sensor fusion system can be used.
[0038] With reference to Fig. 1, is Fig. 1 An exemplary system for implementing hardware utilization for fault management according to some embodiments of the present disclosure. It is noted that this and other arrangements described herein are presented only as examples. Other arrangements and elements (e.g., machines, interfaces, functions, sequences, groupings of functions, etc.) may be used in addition to or instead of those shown, and some elements may be omitted entirely. Furthermore, many of the elements described herein are functional entities that may be implemented as single or distributed components, or in conjunction with other components, in any suitable combination and at any suitable location. Various functions described herein, performed by entities, may be executed by hardware, firmware, and / or software.For example, various functions can be performed by a processor that executes instructions stored in the main memory. In some embodiments, the systems, methods, and processes described herein can be implemented using similar components, features, and / or functions to those of the exemplary autonomous vehicle 400 from [source missing]. Fig. 4A-4D, the exemplary computing device 500 of Fig. 5, and / or the exemplary data center 600 of Fig. 6 will be executed.
[0039] System 100 can contain an application layer 104. Application layer 104 can contain one or more applications. The one or more applications can, for example, be installed on the autonomous or semi-autonomous vehicle or machine 400 or the exemplary computing device 500, or they can be implemented, at least partially, by one or more devices located remotely from the autonomous vehicle 400 and / or the exemplary computing device 500. The one or more applications can be configured to perform operations, including, but not limited to, collecting data, visualizing data, monitoring, implementing controls, etc. For example, one of the applications can be configured to control a braking system of the vehicle 400.Furthermore, an application installed on the autonomous vehicle 400 can be expected to meet automotive safety integrity levels (ASIL) to ensure that the autonomous vehicle 400 maintains and adheres to ASIL. In various implementations, the user can add and / or remove applications at application layer 104.
[0040] To trigger processes (e.g., on the autonomous vehicle 400 and / or components of the autonomous vehicle 400), one or more applications at application layer 104 can contain a variety of tasks. These one or more applications can contain application data (e.g., code). The application data can define the variety of tasks. For example, an initial section of the application data can represent an initial task from the variety of tasks. The variety of tasks related to an application for detecting hazards and / or obstacles in the vehicle's environment can, for example, include collecting sensor data, generating boundary frames, and / or identifying hazards. Each of the variety of tasks can contain a variety of instances (e.g., functions).For example, if the first task is to create bounding boxes, a first instance of the first task could be to detect an object.
[0041] System 100 can contain a flow generator 108. The flow generator 108 can contain one or more processors to perform operations, such as processing application data from application layer 104 to generate a data structure that represents a task flow for executing the multitude of tasks. The task flow can specify at least one sequence or redundancy (e.g., the requirement to operate with spatial and / or temporal redundancy) for each task from the multitude of tasks.
[0042] The task flow can be a graph containing a multitude of nodes. The flow generator 108 can generate the task flow by assigning at least one node to one or more tasks from the multitude of tasks. For example, the task flow can contain nodes of various types, including mission nodes, non-redundant nodes, and redundant nodes to represent the tasks. In various implementations, the flow generator 108 can assign multiple nodes to a given task, such as one or more mission and redundant nodes. For example, each task from the multitude of tasks can contain or be assigned multiple nodes. The flow generator 108 can determine the node type of the one or more nodes to be assigned to a given task based on inputs received from the application layer 104 and / or a user.For example, flow generator 108 can parse the application data to determine the purpose of the specific task (e.g., executing steering operations) in order to determine the node type. The purposes can be determined, for example, by code libraries, features, and / or functions of the application data of the specific task, but are not limited to these.
[0043] In various implementations, the flow generator 108 can determine that a first task from a multitude of tasks meets a criterion for execution in a redundant mode, such as assigning the first task to one or more redundant nodes. For example, the flow generator 108 can classify sections of the multitude of tasks to be assigned to mission and non-redundant nodes. As mentioned earlier, the flow generator 108 can classify tasks for redundant (or non-redundant) operations based on at least one of the user inputs or by parsing the application code received by the application layer 104. In response to determining that the first task contains at least one section to be assigned to a mission node, the flow generator 108 can determine that the criterion for execution in redundant mode is met.In various implementations, the redundant mode indicates that redundant nodes and / or spatial redundancy are included in the task flow.
[0044] In various implementations, the flow generator 108 can determine whether at least one criterion is met based on a feature of the first task received from one or more applications. This feature can include user input. For example, if the user wants to achieve a higher ASIL level, they can input the task to be performed in redundant mode. In various implementations, the multitude of tasks can meet ASIL standards by being redundant and providing greater diagnostic coverage of faults within the system. In other implementations, the feature is assigned based on the function of the first task. For example, if the first task involves generating a graph, the feature can be determined based on code libraries (such as a graph code library) associated with the first task.In this case, the graph that generates the first task can be provided with a non-redundant feature and / or a graph feature.
[0045] In some implementations, the flow generator 108 can determine, in vehicle contexts, that mission nodes are assigned to perform safety-critical tasks, such as vehicle control, by sending commands to the vehicle's control systems. In this case, an initial task assigned to controlling the vehicle's brakes can be classified as a mission node. Mission nodes can be redundant (e.g., running on multiple hardware devices and / or components of a hardware device, or having a redundant node assigned to the mission node).
[0046] In some implementations, the flow generator 108 can determine that redundant nodes are assigned to the corresponding mission nodes, for example, to provide spatial and / or temporal redundancy for fault detection. In some implementations, redundant nodes can be implemented as backup systems to ensure the continuous operation of the vehicle in the event of a failure (e.g., a code error, hardware failure) of mission nodes. For example, a task flow for an application that includes steering would contain redundant nodes to maintain steering control in the event of a mission node failure. Redundant nodes also make it possible to isolate and resolve faults without affecting the system 100.Redundant nodes can be synchronized with mission nodes to ensure that in the event of a failure, the redundant nodes can take over without data or functionality loss.
[0047] In some implementations, non-redundant nodes are not critical to the vehicle and can be implemented without a backup system. An example of a non-redundant node could be one that manages media playback (e.g., music). For instance, the first task among the many tasks, prior to the execution of a mission node and a redundancy node, could be a non-redundant node.
[0048] To generate the task flow, the Flow Generator 108 can evaluate one or more dependencies of the multitude of tasks. These dependencies can include control dependencies and data dependencies. The Flow Generator 108 can determine at least one sequence of multiple tasks or a parallelization of multiple tasks based on at least the task dependencies. The Flow Generator 108 can take, for example, the application code as input and can modify the task flow of the code to output a modified task flow based on one or more criteria (e.g., ASIL) to implement redundancies (e.g., adding redundant nodes) and / or to be logically equivalent to the task flow without modification. The Flow Generator 108 can also consider user inputs and / or application performance requirements to modify the task flow.
[0049] The flow generator 108 can generate the task flow (e.g., a modified task flow compared to an original sequence of instructions or tasks specified by the application layer 104) to exhibit temporal redundancy, such as for the sequential execution of redundant workloads. For example, one or more first processing units (e.g., GPUs) can sequentially execute the multitude of tasks of the application, completing one node's task before moving on to the next.
[0050] The Flow Generator 108 can generate a task flow to exhibit spatial redundancy. For example, the Flow Generator 108 can assign tasks to be performed with redundancy to different processing units. In some cases, spatial redundancy can provide greater diagnostic coverage and a lower error rate than temporal redundancy. For example, spatial redundancy can have diagnostic coverage of approximately 99 percent, while temporal redundancy has diagnostic coverage of approximately 95 percent (although implementing spatial redundancy is not always beneficial, as it can reduce the availability of resources for tasks such as non-critical ones). The Flow Generator 108 can assign a first task from the multitude of tasks to be performed with temporal redundancy and a second task to be performed with spatial redundancy, for example.depending on one or more criteria (e.g. performance requirements, ASIL).
[0051] In various implementations, the task flow may contain a logical node. This logical node may contain an algorithm for checking the task's logic. For example, the logical node checks the task's logic (e.g., the code) and can alert the user to systematic errors within the task and / or the application data.
[0052] The flow generator 108 can assign one or more switches (e.g., a mode switch) to the task flow. The flow generator 108 can assign the switch to effect a change in the use of the hardware resources of the target hardware 120. The flow generator 108 can assign the one or more switches to a position in the task flow in response to determining that a task following that position meets the redundant mode criterion. The switches can be used to direct the hardware management of the target hardware 120, including, for example, switching between occupancy states and / or partitioning states of the target hardware 120. The switches can be nodes that allow the task flow to switch between a default state (e.g., full utilization of the target hardware 120, non-partitioned mode) and a partitioned mode (e.g.,The switch allows the target hardware 120 to change its partitioning state. For example, the switch can divide the target hardware 120 into a first partition and a second partition to run mission and / or redundant nodes. The first partition can run a mission workload (e.g., a multitude of mission nodes), and the second partition can run a corresponding redundant workload (e.g., a multitude of redundant nodes synchronized with the multitude of mission nodes). The partitions can be virtual entities (e.g., MIGs) of a physical device (e.g., GPUs). The switch can be executed, for example, by a central processing unit (CPU) or a GPU system processor (GSP).
[0053] In some implementations, the switch includes an instance-wide synchronization mechanism that enables rapid partition reconfiguration. For example, the switch can create synchronization objects for mission and redundant nodes within the task flow and, upon receiving an indication that the nodes' task execution is complete, initiate a mode change (e.g., partitioned, unpartitioned).
[0054] The Flow Generator 108 can retrieve a hint regarding one or more criteria (e.g., performance requirements, redundancy) related to the tasks. This hint can be received by parsing application code and / or through user input. For example, the user input might specify a high ASIL threshold for the multitude of tasks, resulting in more nodes in the task flow being redundant. Based on this hint, the Flow Generator 108 can assign one or more switches to the task flow. For instance, the switch can be assigned to the task flow before the execution of the first task in the multitude. In this case, the first task can be classified to run in redundant mode, and thus the switch can be assigned before the first task's execution to initiate redundant mode.
[0055] The switch can, for example, be located between a first node, which executes a non-redundant task, and a second and third node. In this case, the second node can execute a first instance of the first task, and a third node executes a second instance of the first task. The first instance can correspond to the mission node, while the second instance corresponds to the redundant node. The second node can reside in a first partition 116, while the third node resides in a second partition 116. The first partition 116 and the second partition 116 (e.g., partitions, partitioning states) can be configured as simultaneous (e.g., concurrent, parallel) multiple contexts, GPU partitions, or multiple instances of a multi-instance GPU (MIG).In this case, after the execution of the first node, the switch is triggered to change from non-partitioned mode to partitioned mode and execute the second and third nodes.
[0056] In various implementations, the nodes of the task flow are executed in one or more contexts (e.g., a computational context), and the one or more switches cause these contexts to change as well. These contexts can contain state information for the execution of the nodes on the hardware cores. Context switching can enable maximum resource utilization and allow the hardware to process multiple tasks and, in some implementations, one or more applications simultaneously.
[0057] In different implementations, one or more switches indicate only a context change. In different implementations, one or more switches only toggle between partitioned and unpartitioned modes. In different implementations, one or more switches toggle both the context and the partition mode.
[0058] In various implementations, System 100 manages one or more switches in a library of switches provided for software configuration. The switch library may include, but is not limited to, switches for changing a partition configuration of the task flow, changing the context, etc., allowing the user to modify the task flow generated by Flow Generator 108.
[0059] In various implementations, the one or more switches contain a first switch and a second switch. The second switch is included by the application after it has been determined that a second task, among the many tasks, depends on the first task and does not meet the criterion for execution in redundant mode (e.g., the second task is assigned to a non-redundant node). The second switch can then be placed between the execution of the first and second tasks, and the second task can cause hardware 120 to be unpartitioned. In other implementations, it is determined that an instance of the first task does not meet the criterion, and thus a second switch can be placed between the execution of instances of the first task.In various implementations, the one or more switches include a third switch that is placed between the nodes in partitioned mode. In this case, the third switch toggles the context of the nodes running on parts of the hardware.
[0060] With further reference to Fig. 1. System 100 can contain a hardware manager 112. The hardware manager 112 can perform operations such as processing the task flow received from the flow generator 108 to partition the target hardware 120 according to the task flow. The hardware manager 112 can execute a variety of tasks. For example, the hardware manager 112 can distribute the numerous tasks to be executed across different hardware components (e.g., MIGs or GPUs).
[0061] For example, if the task flow includes a switch, the hardware manager 112 can assign the nodes following the switch to one or more virtual entities (e.g., MIGs) and / or physical entities (e.g., streaming multiprocessors (SMCs)) of the one or more hardware 120. The hardware manager 112 assigns nodes to the various entities of the hardware 120, for example, based on whether it is a mission node or a non-redundant node. To prevent the hardware manager 112 from being overwhelmed by the application associated with the multitude of tasks, the task flow can provide the hardware manager 112 with instructions for using hardware (e.g., SMCs, MIGs) for partitioning the task flow.
[0062] Depending on whether the task flow contains one or multiple switches, the hardware manager 112 can assign the nodes to the partitions 116 to be processed and executed. The hardware manager 112 can assign one or more applications with one or more nodes to run concurrently (e.g., simultaneously, in parallel) on the partitions 116, thereby improving the utilization of the hardware 120. In various implementations, the hardware manager 112 evaluates the task flow to determine whether it is redundant or non-redundant. In non-redundant cases, the hardware manager 112 can assign the task flow to continue running on the full (e.g., unpartitioned) hardware 120, while redundant flows are partitioned to run in parallel on the partitions 116.
[0063] In various implementations, Hardware Manager 112 sequentially reuses parts (e.g., virtual entities and / or physical components) of the same hardware to perform tasks redundantly. For example, redundant nodes can be executed sequentially on the same parts (e.g., MIGs, SMCs). In other implementations, Hardware Manager 112 mutates Hardware 120 (e.g., by creating a digital twin of it) to split it into one or more virtual devices, enabling redundant and parallel task execution. Hardware 120 can then execute multiple redundant tasks simultaneously. This can improve the detectability of permanent random errors. The one or more virtual devices can then be mutated into a single or multiple virtual devices to perform non-redundant tasks.
[0064] In various implementations, the hardware manager 112 can assign different types of hardware devices (e.g., GPU, deep learning accelerator (DLA)) for inference tasks (e.g., object detection) to perform the same inference task redundantly. For example, an initial inference task can be executed multiple times on different hardware devices to improve the fault detection capabilities of each hardware device 120, thus complementing the fault detection features of the hardware 120.
[0065] In various implementations, the Hardware Manager 112 can collect activation patterns from the Hardware 120 by inserting diagnostic nodes into the task flow. Activation patterns describe how different components (e.g., texture units, MIGs) of the Hardware 120 are used during a variety of tasks. Based on the collected activation patterns, the Hardware Manager 112 can add additional runtime diagnostics for the Hardware 120. For example, if the activation patterns do not meet expectations and / or do not reach a threshold for the activation pattern, the Hardware Manager 112 adds additional diagnostic nodes to the task flow to assess whether there are any errors within the Hardware 120. Based on the runtime diagnostics received from the diagnostic nodes, a report on the various security versus performance profiles of the Hardware 120 can then be generated.The user can then select one of the various profiles on which the Hardware 120 should run. For example, if the activation patterns indicate that the Hardware 120 prioritizes performance over security, due to the lack of a switch implementation, the user can adjust the profile accordingly.
[0066] In various implementations, user input can include at least one or all descriptions of a task flow, dependencies between tasks, and / or target hardware for executing the task flow. For example, the user can decide on the task flow configuration and also which tasks depend on each other. In this case, the user can modify the task flow based on user preferences, using the task flow provided by the flow generator 108. Furthermore, the user can select on which target hardware 120 and / or which components of the target hardware 120 the multitude of tasks and / or nodes should be executed. The user can provide input to the hardware manager 112 to execute the task flow accordingly.
[0067] In various implementations, the task flow executes the multitude of nodes on physical parts of the hardware device. For example, if the hardware device contains multiple streaming multiprocessors, the task flow in redundant mode can execute the multitude of nodes in parallel and assign each of the multitude of nodes to a streaming multiprocessor of the hardware device.
[0068] Fig. Figure 2 is a schematic diagram of an example of a task flow 200, which is carried out by the system of Fig. 1 can be generated and executed, according to some implementations of the present disclosure. In Fig. 2. Task flow 200 begins with a non-redundant node in a first context C0. The non-redundant node runs on the entire hardware (e.g., in unpartitioned mode) and can also be the first instance of the first task from the multitude of tasks. Task flow 200 includes a first switch 204, which is executed after the non-redundant node completes its task. The first switch 204 can act as both a context switch and a change in the partition configuration (e.g., from unpartitioned to partitioned mode). As in Fig. As shown in Figure 2, M1 (e.g., the first mission node) and R1 (e.g., the first redundant node) are partitioned and can run on separate hardware components of the same device (e.g., MIGs of a GPU). The first mission node and the first redundant nodes reside in a second context C1 and a third context C2, respectively. The switch can also contain synchronization objects, allowing the first mission node and the first redundant node to run in parallel. This ensures that if the first mission node fails, the first redundant node takes over.
[0069] After the execution of the first mission node and the first redundant node has been completed, a second switch 212 may be included. The second switch 212 can switch contexts. The second switch 212 may be included in response to the flow generator 108 determining that a redundant task should be executed (e.g., using a second mission node and a second redundant node), following the execution of the task by the first mission node and the first redundant node.
[0070] As in Fig. As seen in section 2, the first mission node can transition to a second mission node M2, and the first redundant node can transition to a second redundant node R2. The second mission node can run on the same hardware and in the same context as the first mission node, and the second redundant node can run on the same hardware and in the same context as the second mission node. In some implementations, the switch 212 changes contexts, and the second redundant node and the second mission node run in different contexts than the first redundant node and the first mission node.
[0071] In response to the completion of the execution (task) on the second mission node and the second redundant node, a third switch 208 changes the context and partition configuration (e.g., from partitioned to non-partitioned mode). The third switch 208 can be implemented, for example, in response to a task flow transitioning from a mission and redundant node to a non-redundant node. The third switch 208 can, for instance, instruct the hardware manager 112 to switch from partitioning and execution on a first MIG and a second MIG to execution on the entire GPU.
[0072] In various implementations, the task flow (e.g., task flow 200) can contain two or more mission nodes, redundant nodes, and / or non-redundant nodes in the task flow graph. In this case, two or more switches can be implemented to accommodate multiple instances of nodes.
[0073] In various implementations, the task flow can include both temporal and spatial redundancy. For example, the task flow may begin with two or more nodes (e.g., non-redundant, redundant, or mission) that execute sequentially (e.g., temporal redundancy) and may transition to two or more nodes that execute in parallel on one or more hardware components (e.g., spatial redundancy).
[0074] With reference to Fig. 3. Each block of the Method 300 described herein comprises a computational process that can be performed using any combination of hardware, firmware, and / or software. For example, various functions can be performed by a processor executing instructions stored in memory. The methods can also be embodied as computer-usable instructions stored on computer storage media. The methods can be provided by a standalone application, a service or hosted service (alone or in combination with another hosted service), or a plug-in for another product, to name just a few. Furthermore, Method 300 is exemplified with respect to the system of Fig. 1 described. However, these procedures can additionally or alternatively be performed by any system or any combination of systems, including, but not limited to, the systems described herein.
[0075] In block 302 of procedure 300, a first task is selected from a multitude of tasks that meets a criterion for execution in a redundant mode. This criterion determines whether the first task should be redundant (e.g., run on multiple nodes). A redundant task could, for example, be the control of a vehicle's braking system. For safety reasons, a task associated with braking would meet the criterion of being executed redundantly.
[0076] Block 304 defines a task flow for executing a multitude of tasks. This is based on the number of tasks and whether each of them meets the criteria. The task flow can be determined based on redundant and non-redundant tasks. An example task flow is shown in Fig. 2. Since it has been determined that the first task meets the criterion, the switch can be assigned before the first task is executed. The switch can change the task flow from execution in a non-redundant mode (e.g., all hardware) to a redundant mode (e.g., executing tasks on separate hardware components). The switch can cause hardware on which the multitude of tasks are executed to be divided into a first partition and a second partition (e.g., partition 116) and enter redundant mode. Changing the task flow can thus enable redundant and / or mission tasks to be executed in parallel. In some implementations, the switch also changes the context.
[0077] In Block 306, the multitude of tasks is executed according to the task flow. After determining the task flow and implementing the switch, the multitude of tasks can be executed. The multitude of tasks can be executed by running a first instance of the first task on the first partition and a second instance of the first task on the second partition. For example, the first instance can run on a redundant node, and the second instance can run on a mission node in separate partitions (e.g., MIGs of the same GPU). The multitude of tasks can, for example, start with a non-redundant node running on a full GPU and then encounter a switch that causes the task flow to run on partitioned hardware. In redundant mode (e.g.,With partitioned hardware, tasks can be executed concurrently and synchronized by the switches, so that tasks in redundant mode complete simultaneously when they encounter a second switch or a second non-redundant node. Tasks in redundant mode can also be configured to complete concurrently without encountering the second switch or non-redundant node.
[0078] The systems and procedures described herein can be used without restriction by non-autonomous vehicles, semi-autonomous vehicles (e.g. in one or more adaptive driver assistance systems (ADAS)), steered and unsteered robots or robotic platforms, warehouse vehicles, all-terrain vehicles, vehicles coupled to one or more trailers, flying vehicles, boats, shuttles, emergency response vehicles, motorcycles, electric or motorized bicycles, aircraft, construction vehicles, underwater vehicles, drones and / or other types of vehicles.Furthermore, the systems and methods described herein can be used for a variety of purposes, including but not limited to machine control, machine movement, machine driving, synthetic data generation, model training, perception, augmented reality, virtual reality, mixed reality, robotics, security and surveillance, simulation and digital twinning, autonomous or semi-autonomous machine applications, deep learning, environmental simulation, object or actor simulation and / or digital twinning, data center processing, conversational AI, light transport simulation (e.g., ray tracing, path tracing, etc.), collaborative content creation for 3D assets, cloud computing, and / or other suitable applications.
[0079] The disclosed embodiments can be included in a variety of different systems, such as automotive systems (e.g., a control system for an autonomous or semi-autonomous machine, a perception system for an autonomous or semi-autonomous machine), systems implemented using a robot, aerial systems, media systems, boat systems, intelligent area surveillance systems, systems for performing deep learning operations, systems for performing simulation operations, systems for performing digital twin operations, systems implemented using an edge device, systems that include one or more virtual machines (VMs), systems for performing synthetic data generation operations, systems that are at least partially implemented in a data center, systems for performing conversational AI operations, and systems for hosting real-time streaming applications.Systems for presenting at least one type of virtual reality content, augmented reality content or mixed reality content, systems for performing light transport simulations, systems for performing collaborative content creation for 3D assets, systems that are implemented at least partially using cloud computing resources, and / or other types of systems. EXEMPLARY AUTONOMOUS VEHICLE
[0080] Fig. Figure 4A is an illustration of an exemplary autonomous vehicle 400 according to some embodiments of the present disclosure. The autonomous vehicle 400 (hereinafter referred to alternatively as "vehicle 400") may, without limitation, include: a passenger vehicle, such as a car, truck, bus, emergency service vehicle, shuttle, electric or motorized bicycle, motorcycle, fire engine, police vehicle, ambulance, boat, construction vehicle, underwater vehicle, robotic vehicle, drone, aircraft, a vehicle coupled to a trailer (e.g., a semi-trailer truck used for transporting cargo), and / or another type of vehicle (e.g., one that is unmanned and / or carries one or more passengers).Autonomous vehicles are generally described in terms of automation levels defined by the National Highway Traffic Safety Administration (NHTSA), a division of the U.S. Department of Transportation, and the Society of Automotive Engineers (SAE) standard "Taxonomy and Definitions for Terms Related to Driving Automation Systems for On-Road Motor Vehicles" (Standard No. J3016-201806, published June 15, 2018, Standard No. J3016-201609, published September 30, 2016, and earlier and future versions of this standard). The Vehicle 400 may exhibit functionality corresponding to one or more of the Levels 3 through 5 of autonomous driving levels.The Vehicle 400 may exhibit functionality corresponding to one or more of the Levels 1 through 5 of autonomous driving levels. For example, depending on its embodiment, the Vehicle 400 may be capable of driver assistance (Level 1), partial automation (Level 2), conditional automation (Level 3), high automation (Level 4), and / or full automation (Level 5). The term "autonomous," as used herein, may encompass any and / or all types of autonomy for the Vehicle 400 or any other machine, such as fully autonomous, highly autonomous, conditionally autonomous, partial autonomous, assistive autonomy, semi-autonomous, primary autonomous, or any other designation. The Vehicle 400 may incorporate ASIL for various systems described herein.
[0081] The vehicle 400 can include components such as a chassis, a vehicle body, wheels (e.g., 2, 4, 6, 8, 18, etc.), tires, axles, and other vehicle components. The vehicle 400 can include a propulsion system 450, such as an internal combustion engine, a hybrid electric power plant, a pure electric motor, and / or another type of propulsion. The propulsion system 450 can be connected to a drivetrain of the vehicle 400, which may include a transmission to enable the propulsion of the vehicle 400. The propulsion system 450 can be controlled in response to signals received from the throttle / accelerator device 452.
[0082] A steering system 454, which may include a steering wheel, can be used to steer the vehicle 400 (e.g., along a desired path or route) when the propulsion system 450 is in operation (e.g., when the vehicle is in motion). The steering system 454 can receive signals from a steering actuator 456. The steering wheel can be optional for full automation (level 5).
[0083] The brake sensor system 446 can be used to actuate the vehicle brakes in response to receiving signals from the brake actuators 448 and / or the brake sensors.
[0084] The controller(s) 436, which controls one or more systems-on-chips (SoCs) 404 ( Fig. 4C) and / or GPUs, can provide signals (e.g., representing commands) to one or more components and / or systems of the vehicle 400. For example, the controller(s) can send signals to actuate the vehicle brakes via one or more brake actuators 448, to actuate the steering system 454 via one or more steering actuators 456, and to actuate the propulsion system 450 via one or more throttle / accelerator devices 452. The controller(s) 436 can include one or more built-in (e.g., integrated) computing devices (e.g., supercomputers) that process sensor signals and issue operating commands (e.g., signals representing commands) to enable autonomous driving and / or to assist a human driver in driving the vehicle 400.The controller(s) 436 can include a first controller 436 for autonomous driving functions, a second controller 436 for functional safety functions, a third controller 436 for artificial intelligence functions (e.g., computer vision), a fourth controller 436 for infotainment functions, a fifth controller 436 for emergency redundancy, and / or other controllers. In some examples, a single controller 436 can perform two or more of the aforementioned functionalities, two or more controllers 436 can perform a single functionality, and / or any combination thereof.
[0085] The controller(s) 436 can provide the signals for controlling one or more components and / or systems of the vehicle 400 in response to sensor data received from one or more sensors (e.g. sensor inputs). The sensor data can be received, for example, without restriction, from the following: Global Navigation Satellite Systems (GNSS) sensor(s) 458 (e.g., Global Positioning System sensor(s)), radar sensor(s) 460, ultrasonic sensor(s) 462, lidar sensor(s) 464, inertial measurement unit (IMU) sensor(s) 466 (e.g., accelerometer(s), gyroscope(s), magnetic compass(s), magnetometer(s), etc.), microphone(s) 496, stereo camera(s) 468, wide-angle camera(s) 470 (e.g., fisheye cameras), infrared camera(s) 472, surround-view camera(s) 474 (e.g., 360-degree cameras), long-range and / or medium-range camera(s) 498. Speed sensor(s) 444 (e.g.B. for measuring the speed of the vehicle 400), vibration sensor(s) 442, steering sensor(s) 440, brake sensor(s) (e.g. as part of the brake sensor system 446), and / or other types of sensors.
[0086] One or more of the controllers 436 can receive inputs (e.g., represented by input data) from an instrument cluster 432 of the vehicle 400 and provide outputs (e.g., represented by output data, display data, etc.) via a human-machine interface (HMI) display 434, an audible alarm, a loudspeaker, and / or via other components of the vehicle 400. The outputs can include information such as vehicle speed, engine speed, time, map data (e.g., the high-definition (HD) map 422 of the vehicle). Fig. 4C), location data (e.g., the location of vehicle 400, as shown on a map), direction, location of other vehicles (e.g., an occupancy grid), information about objects and the status of objects as perceived by the controller(s) 436, etc. For example, the HMI display 434 can show information about the presence of one or more objects (e.g., a road sign, a warning sign, a changing traffic light, etc.) and / or information about driving maneuvers that the vehicle has performed, is currently performing, or will perform (e.g., changing lanes now, taking exit 34B in two miles, etc.).
[0087] The vehicle 400 further includes a network interface 424, which can use one or more wireless antennas 426 and / or modems for communication over one or more networks. The network interface 424 can be suitable, for example, for communication over Long-Term Evolution (LTE), Wideband Code Division Multiple Access (WCDMA), Universal Mobile Telecommunications System (UMTS), Global System for Mobile Communication (GSM), IMT-CDMA Multi-Carrier (CDMA2000), etc. The wireless antenna(s) 426 can also enable communication between objects in the environment (e.g., vehicles, mobile devices, etc.) using a local network(s) such as Bluetooth, Bluetooth Low Energy (LE), Z-Wave, ZigBee, etc.and / or a low power wide area network (LPWANs), such as LoRaWAN, SigFox, etc.
[0088] To meet ASIL requirements for various applications, the Vehicle 400 can implement System 100 for fault detection management and for efficient processing and execution of applications for the Vehicle 400 systems described above.
[0089] Fig. 4B is an example of camera locations and fields of view for the exemplary autonomous vehicle 400 from Fig. 4A according to some embodiments of the present disclosure. The cameras and respective fields of view are an exemplary embodiment and are not intended to be limiting. For example, additional and / or alternative cameras may be included and / or the cameras may be located at different locations on the vehicle 400.
[0090] The camera types may include, but are not limited to, digital cameras that are adapted for use with the components and / or systems of the Vehicle 400. The camera(s) may operate at Automotive Safety Integrity Level (ASIL) B and / or another ASIL. Depending on the configuration, the camera types may be capable of any frame rate, such as 60 frames per second (fps), 120 fps, 240 fps, etc. The cameras may use roller shutters, global shutters, another type of shutter, or a combination thereof.In some examples, the color filter array may include a red-clear-clear-clear (RCCC) color filter array, a red-clear-clear-blue (RCCB) color filter array, a red-blue-green (RBGC) color filter array, a Foveon X3 color filter array, a Bayer sensor color filter array (RGGB), a monochrome sensor color filter array, and / or another type of color filter array. In some embodiments, cameras with clear pixels, such as cameras with an RCCC, RCCB, and / or RBGC color filter array, may be used to increase light sensitivity.
[0091] In some examples, one or more of the cameras can be used to perform advanced driver assistance system (ADAS) functions (e.g., as part of a redundant or fail-safe design). For instance, a multi-function monocular camera can be installed to provide features including lane departure warning, traffic sign recognition, and intelligent headlight control. One or more of the cameras (e.g., all cameras) can simultaneously record and provide image data (e.g., video).
[0092] One or more of the cameras can be mounted in a bracket, such as a specially designed (three-dimensional ("3D") printed bracket, to eliminate stray light and reflections from inside the vehicle (e.g., reflections of the dashboard in the windshield) that could interfere with the camera's image acquisition. Referencing the mounting assemblies of exterior mirrors, the exterior mirror assemblies can be individually 3D printed so that the camera mounting plate is adapted to the shape of the exterior mirror. In some examples, the camera(s) can be integrated into the exterior mirror. For side cameras, the camera(s) can also be integrated into the four pillars at each corner of the cabin.
[0093] Cameras with a field of view that includes portions of the environment in front of the vehicle (e.g., forward-facing cameras) can be used for surround view to help identify forward paths and obstacles and to provide, with the aid of one or more controllers and / or control SoCs, information critical for creating an occupancy grid and / or determining preferred vehicle paths. Forward-facing cameras can be used to perform many of the same ADAS functions as LiDAR, including emergency braking, pedestrian detection, and collision avoidance. Forward-facing cameras can also be used for ADAS functions and systems that include lane departure warnings (LDW), autonomous cruise control (ACC), and / or other functions such as traffic sign recognition.
[0094] A variety of cameras can be used in a forward-facing configuration, including, for example, a monocular camera platform containing a complementary metal oxide semiconductor (CMOS) color imager. Another example could be a 470 wide-angle camera, which can be used to capture objects moving into the field of view from the periphery (e.g., pedestrians, crossing vehicles, or bicycles). Although in Fig. While Figure 4B illustrates only one wide-angle camera, the vehicle 400 can have any number (including zero) of wide-angle cameras 470. Furthermore, any number of long-range cameras 498 (e.g., a pair of long-range stereo cameras) can be used for depth-based object detection, particularly for objects for which a neural network has not yet been trained. The long-range camera(s) 498 can also be used for object detection and classification, as well as for basic object tracking.
[0095] Any number of stereo cameras 468 can also be included in a forward-facing configuration. In at least one embodiment, one or more of the stereo cameras 468 can include an integrated control unit comprising a scalable processing unit that can provide programmable logic (“FPGA”) and a multicore microprocessor with an integrated controller area network (“CAN”) or Ethernet interface on a single chip. Such a unit can be used to create a 3D map of the vehicle's surroundings that includes a distance estimate for all points in the image. An alternative stereo camera(s) 468 can include a compact stereo vision sensor(s) that can include two camera lenses (one each on the left and right) and an image processing chip that measures the distance between the vehicle and the target object and processes the generated information (e.g., distance, distance, and distance).B. metadata) can be used to activate the autonomous emergency braking and lane departure warning functions. Other types of stereo cameras 468 may be used in addition to or as an alternative to those described herein.
[0096] Cameras with a field of view that includes sections of the environment to the sides of the vehicle 400 (e.g., side cameras) can be used for the perimeter view and provide information used for creating and updating the occupancy grid and for generating side-impact collision warnings. For example, the perimeter camera(s) 474 (e.g., four perimeter cameras 474, as in Fig. (Figure 4B illustrates) is positioned on the vehicle 400. The surround view camera(s) 474 may include a wide-angle camera 470, a fisheye camera, a 360-degree camera, and / or the like. For example, four fisheye cameras may be positioned at the front, rear, and sides of the vehicle. In an alternative arrangement, the vehicle may use three surround view cameras 474 (e.g., left, right, and rear) and one or more other cameras (e.g., a forward-facing camera) as a fourth surround view camera.
[0097] Cameras with a field of view that includes sections of the area behind the vehicle 400 (e.g., reversing cameras) can be used for parking assistance, surround view, rear-impact warnings, and creating and updating the occupancy grid. A variety of cameras can be used, including, without limitation, cameras that are also suitable as forward-facing cameras (e.g., long-range and / or medium-range cameras 498, stereo cameras 468, infrared cameras 472, etc.), as described herein.
[0098] Fig. 4C is a block diagram of an exemplary system architecture for the exemplary autonomous vehicle 400 from Fig. 4A according to some embodiments of the present disclosure. It should be noted that these and other arrangements described herein are presented only as examples. Other arrangements and elements (e.g., machines, interfaces, functions, sequences, groupings of functions, etc.) may be used in addition to or instead of those shown, and some elements may be omitted entirely. Furthermore, many of the elements described herein are functional entities that may be implemented as single or distributed components, or in conjunction with other components, in any suitable combination and at any suitable location. Various functions described herein, which are performed by entities, may be executed by hardware, firmware, and / or software. For example, various functions may be performed by a processor executing instructions stored in memory.
[0099] Each of the components, features and systems of the 400 vehicle in Fig. 4C is illustrated as being connected via bus 402. Bus 402 may contain a Controller Area Network (CAN) data interface (referred to herein alternatively as the "CAN bus"). A CAN may be a network within the vehicle 400 that is used to support the control of various features and functions of the vehicle 400, such as the operation of brakes, acceleration, braking, steering, windshield wipers, etc. A CAN bus may be configured to have dozens or even hundreds of nodes, each with its own unique identifier (e.g., a CAN ID). The CAN bus can be read to determine the steering wheel angle, vehicle speed, engine speed (rpm), button positions, and / or other vehicle status indicators. The CAN bus may be ASIL B compliant.
[0100] Although the 402 bus is described herein as a CAN bus, this is not intended as a limitation. For example, FlexRay and / or Ethernet can be used in addition to or as alternatives to the CAN bus. Furthermore, while a single line is used to represent the 402 bus, this is not meant as a limitation. For example, there can be any number of 402 buses, which may contain one or more CAN buses, one or more FlexRay buses, one or more Ethernet buses, and / or one or more other types of buses using a different protocol. In some examples, two or more 402 buses may be used to perform different functions and / or for redundancy. For example, a first 402 bus may be used for collision avoidance functionality, and a second 402 bus may be used for actuation control.In each example, each bus 402 can communicate with one of the vehicle 400 components, and two or more buses 402 can communicate with the same components. In some examples, each SoC 404, each controller 436, and / or each computer within the vehicle can have access to the same input data (e.g., inputs from vehicle 400 sensors) and be connected to a common bus, such as the CAN bus.
[0101] The vehicle 400 can contain one or more controllers 436, as described herein with reference to Fig. 4A are described. The controller(s) 436 can be used for a variety of functions. The controller(s) 436 can be coupled with one of the various other components and systems of the vehicle 400 and can be used for controlling the vehicle 400, for the artificial intelligence of the vehicle 400, for infotainment for the vehicle 400, and / or the like.
[0102] The vehicle 400 can contain a system-on-a-chip (SoC) 404. The SoC 404 can contain CPU(s) 406, GPU(s) 408, processor(s) 410, cache(s) 412, accelerator(s) 414, data storage 416, and / or other components and features not illustrated. The SoC(s) 404 can be used to control the vehicle 400 in a variety of platforms and systems. For example, the SoC(s) 404 can be combined in a system (e.g., the system of the vehicle 400) with an HD card 422, which is accessed via a network interface 424 by one or more servers (e.g., the server(s) 478). Fig. 4D) Receive map refreshes and / or updates.
[0103] The CPU(s) 406 may contain a CPU cluster or CPU complex (hereinafter referred to as "CCPLEX"). The CPU(s) 406 may contain multiple cores and / or L2 caches. In some embodiments, the CPU(s) 406 may, for example, contain eight cores in a coherent multiprocessor configuration. In some embodiments, the CPU(s) 406 may contain four dual-core clusters, each cluster having a dedicated L2 cache (e.g., a 2 MB L2 cache). The CPU(s) 406 (e.g., the CCPLEX) may be configured to support the concurrent operation of clusters, so that any combination of the CPU(s) 406's clusters can be active at any given time.
[0104] The CPU(s) 406 can implement power management features that include one or more of the following: individual hardware blocks can be automatically clocked when idle to save dynamic power; each core clock can be controlled when the core is not actively executing instructions due to the execution of WFI / WFE instructions; each core can be independently power-controlled; each core cluster can be independently clock-controlled when all cores are clock-controlled or power-controlled; and / or each core cluster can be independently power-controlled when all cores are power-controlled. The CPU(s) 406 can also implement an improved power state management algorithm that establishes acceptable power states and expected wake-up times, and the hardware / microcode determines the best power state to input for the core, cluster, and CCPLEX.The processing kernels can support simplified sequences for inputting the performance state into the software, thereby offloading the work to the microcode.
[0105] The GPU(s) 408 may include an integrated graphics processor (referred to herein alternatively as the "iGPU"). The GPU(s) 408 may be programmable and efficient for parallel workloads. The GPU(s) 408 may, in some examples, use an extended Tensor instruction set. The GPU(s) 408 may include one or more streaming microprocessors, each of which may contain an L1 cache (for example, an L1 cache of at least 96 KB), and two or more of the streaming microprocessors may share an L2 cache (for example, an L2 cache of 512 KB). In some embodiments, the GPU(s) 408 may contain at least eight streaming microprocessors. The GPU(s) 408 can use an application programming interface (API) for calculations.Furthermore, the GPU(s) 408 can utilize one or more parallel computing platforms and / or programming models (e.g., NVIDIA's CUDA).
[0106] The GPU(s) 408 can be performance-optimized for best performance in automotive and embedded applications. For example, the GPU(s) 408 can be manufactured using a FinFET field-effect transistor. However, this is not a limitation, and the GPU(s) 408 can also be manufactured using other semiconductor manufacturing processes. Each streaming microprocessor can contain a number of mixed-precision processing cores divided into multiple blocks. For example, and without limitation, 64 PF32 cores and 32 PF64 cores can be divided into four processing blocks. In such an example, each processing block can be allocated 16 FP32 cores, 8 FP64 cores, 16 INT32 cores, two mixed-precision NVIDIA TENSOR COREs for Deep Learning matrix arithmetic, an L0 instruction cache, a warp scheduler, a dispatch unit and / or a 64 KB register file.Furthermore, the streaming microprocessors can include independent parallel integer and floating-point data paths to provide efficient execution of workloads with a mix of computations and addressing operations. The streaming microprocessors can include an independent thread scheduling function to enable fine-grained synchronization and cooperation between parallel threads. The streaming microprocessors can also include a combined L1 data cache and a shared memory unit to improve performance while simplifying programming.
[0107] The GPU(s) 408 can include high-bandwidth memory (HBM) and / or a 16 GB HBM2 memory subsystem to provide a peak memory bandwidth of approximately 900 GB / second in some examples. In some examples, in addition to or as an alternative to the HBM memory, synchronous graphics random access memory (SGRAM), such as double-data-rate type five synchronous graphics random access memory (GDDR5), can be used.
[0108] The GPU(s) 408 may incorporate a unified memory technology that includes access counters to enable more accurate migration of memory pages to the processor that accesses them most frequently, thereby improving the efficiency of memory areas shared by processors. In some examples, support for Address Translation Services (ATS) may be used so that the GPU(s) 408 can directly access the page tables of the CPU(s) 406. In such examples, if the memory management unit (MMU) of the GPU(s) 408 fails, an address translation request can be sent to the CPU(s) 406. In response, the CPU(s) 406 can search its page tables for the virtual-physical mapping for the address and send the translation back to the GPU(s) 408.Thus, the unified memory technology can enable a single unified virtual address space for the memory of both the CPU(s) 406 and the GPU(s) 408, thereby simplifying the programming of the GPU(s) 408 and the porting of applications to the GPU(s) 408.
[0109] Additionally, the GPU(s) 408 can contain an access counter that tracks the frequency of GPU 408 access to the memory of other processors. This access counter can help ensure that memory pages are moved into the physical memory of the processor that accesses them most frequently.
[0110] The SoC(s) 404 can contain any number of caches 412, including those described herein. For example, the cache(s) 412 can contain an L3 cache available to both the CPU(s) 406 and the GPU(s) 408 (e.g., connected to both the CPU(s) 406 and the GPU(s) 408). The cache(s) 412 can contain a write-back cache capable of tracking row states, such as by using a cache coherence protocol (e.g., MEI, MESI, MSI, etc.). The L3 cache can j Depending on the design, it may contain 4 MB or more, although smaller cache sizes can also be used.
[0111] The SoC(s) 404 may contain one or more arithmetic logic units (ALUs) that can be used to perform processing related to one of the many tasks or operations of the Vehicle 400, such as processing deep neural networks (DNNs). Additionally, the SoC(s) 404 may contain one or more floating-point units (FPUs) or other mathematical or numerical coprocessors for performing mathematical operations within the system. For example, the SoC(s) 104 may contain one or more FPUs integrated as execution units into a CPU(s) 406 and / or a GPU(s) 408.
[0112] The SoC(s) 404 can contain one or more Accelerators 414 (e.g., hardware accelerators, software accelerators, or a combination thereof). For example, the SoC(s) 404 can contain a hardware acceleration cluster, which may include optimized hardware accelerators and / or a large amount of on-chip memory. The large on-chip memory (e.g., 4 MB SRAM) can enable the hardware acceleration cluster to accelerate neural networks and other computations. The hardware acceleration cluster can be used in conjunction with the GPU(s) 408 and offload some of the GPU's tasks (e.g., to free up more GPU cycles for other tasks). The accelerator(s) 414 can be used, for example, for targeted workloads (e.g. perception, convolutional neural networks (CNNs), etc.).) are used that are stable enough to be suitable for acceleration. The term "CNN" as used herein can include all types of CNNs, including region-based or regional convolutional neural networks (RCNNs) and fast RCNNs (e.g., for object detection).
[0113] The Accelerator 414 (e.g., the Hardware Acceleration Cluster) can include a Deep Learning Accelerator (DLA). The DLA(s) can include one or more Tensor Processing Units (TPUs) configured to provide an additional ten trillion operations per second for deep learning applications and inference. The TPUs can be accelerators configured and optimized to perform image processing functions (e.g., CNNs, RCNNs, etc.). The DLA(s) can also be optimized for a specific set of neural network types and floating-point operations, as well as for inference. The design of the DLA(s) can deliver more performance per millimeter than a general-purpose GPU and far surpasses the performance of a CPU.The TPU(s) can perform several functions, including a convolution function for a single instance, supporting, for example, INT8, INT16 and FP16 data types for both features and weights, as well as post-processor functions.
[0114] The DLA(s) can quickly and efficiently execute neural networks, especially CNNs, on processed or unprocessed data for a variety of functions, including, but not limited to: a CNN for object identification and detection using camera sensor data; a CNN for distance estimation using camera sensor data; a CNN for emergency vehicle detection and identification using microphone data; a CNN for facial recognition and vehicle owner identification using camera sensor data; and / or a CNN for security and / or protection-related events.
[0115] The DLA(s) can execute any function of the GPU(s) 408, and by using an inference accelerator, a developer can, for example, allocate either the DLA(s) or the GPU(s) 408 to each function. For example, the developer can concentrate the processing of CNNs and floating-point operations on the DLA(s) and leave other functions to the GPU(s) 408 and / or another accelerator(s) 414.
[0116] The Accelerator 414 (e.g., the Hardware Accelerator Cluster) can include a Programmable Vision Accelerator (PVA), which may also be referred to herein as a Computer Vision Accelerator. The PVA(s) can be designed and configured to accelerate computer vision algorithms for Advanced Driver Assistance Systems (ADAS), autonomous driving, and / or augmented reality (AR) and / or virtual reality (VR) applications. The PVA(s) can provide a balance between performance and flexibility. For example, each PVA can, without limitation, include any number of Reduced Instruction Set Computer (RISC) cores, Direct Memory Access (DMA) cores, and / or any number of vector processors.
[0117] The RISC cores can interact with image sensors (e.g., the image sensors of one of the cameras described herein), image signal processor(s), and / or the like. Each RISC core can contain any amount of memory. Depending on the implementation, the RISC cores can use any number of protocols. In some examples, the RISC cores can run a real-time operating system (RTOS). The RISC cores can be implemented with one or more integrated circuits, application-specific integrated circuits (ASICs), and / or memory devices. For example, the RISC cores can include an instruction cache and / or tightly coupled RAM.
[0118] The DMA can allow components of the PVA(s) to access the system's memory independently of the CPU(s) 406. The DMA can support any number of features that serve to optimize the PVA, including, but not limited to, support for multidimensional addressing and / or circular addressing. In some examples, the DMA can support up to six or more dimensions of addressing, which may include block width, block height, block depth, horizontal block stepping, vertical block stepping, and / or depth stepping.
[0119] Vector processors can be programmable processors designed to efficiently and flexibly execute programming for computer vision algorithms and provide signal processing capabilities. In some examples, the PVA may contain a PVA core and two vector processing subsystem partitions. The PVA core may include a processor subsystem, DMA machine(s) (e.g., two DMA machines), and / or other peripheral devices. The vector processing subsystem may operate as the primary processing machine of the PVA and may include a vector processing unit (VPU), an instruction cache, and / or working memory (e.g., VMEM).A VPU core can contain a digital signal processor, such as a single instruction, multiple data (SIMD) or a very long instruction word (VLIW). The combination of SIMD and VLIW can increase throughput and speed.
[0120] Each vector processor can contain an instruction cache and can be coupled to dedicated memory. Therefore, in some examples, each vector processor can be configured to operate independently of the others. In other examples, the vector processors contained in a particular PVA can be configured to employ data parallelism. For example, in some embodiments, the multitude of vector processors contained in a single PVA can execute the same computer vision algorithm, but for different image segments. In other examples, the vector processors contained in a particular PVA can simultaneously execute different computer vision algorithms on the same image, or even different algorithms on successive images or segments of an image.Among other things, the hardware acceleration cluster can contain any number of PVAs and any number of vector processors in each PVA. Furthermore, the PVA(s) can include additional memory for error-correcting code (ECC) to enhance overall system security.
[0121] The Accelerator 414 (e.g., the Hardware Acceleration Cluster) can include an on-chip computer vision network and SRAM to provide high-bandwidth, low-latency SRAM for the Accelerator 414. In some examples, the on-chip memory can include at least 4 MB of SRAM, consisting, for example, and without limitation, of eight field-configurable memory blocks accessible to both the PVA and the DLA. Each pair of memory blocks can include an Advanced Peripheral Bus (APB) interface, configuration circuitry, a controller, and a multiplexer. Any type of memory can be used. The PVA and the DLA can access the memory via a backbone, providing high-speed memory access for both the PVA and the DLA.The backbone can include an on-chip computer vision network that connects the PVA and DLA to the main memory (e.g., using the APB).
[0122] The on-chip computer vision network can include an interface that, prior to the transmission of control signals / addresses / data, ensures that both the PVA and the DLA provide ready-to-use and valid signals. Such an interface can provide separate phases and channels for the transmission of control signals / addresses / data, as well as burst communication for continuous data transmission. This type of interface can conform to ISO 26262 or IEC 61508 standards, although other standards and protocols can also be used.
[0123] In some examples, the SoC(s) 404 may include a real-time ray tracing hardware accelerator as described in US Patent Application No. 16 / 101,232, filed on August 10, 2018. The real-time ray tracing hardware accelerator can be used to quickly and efficiently determine the positions and extents of objects (e.g., within a world model), to generate real-time visualization simulations, for radar signal interpretation, for sound propagation synthesis and / or analysis, for the simulation of SONAR systems, for general wave propagation simulation, for comparison with lidar data for localization purposes, and / or for other functions and / or purposes. In some embodiments, one or more Tree Traversal Units (TTUs) can be used to perform one or more operations related to ray tracing.
[0124] The Accelerator 414 (e.g., the hardware accelerator cluster) has a wide range of applications for autonomous driving. The PVA can be a programmable vision accelerator used for critical processing steps in ADAS and autonomous vehicles. The PVA's capabilities are well-suited to algorithmic domains requiring predictable processing with low power consumption and low latency. In other words, the PVA is well-suited for semi-dense or dense regular computations, even with small datasets, that require predictable runtimes with low latency and low power consumption. In the context of autonomous vehicle platforms, PVAs are therefore designed to execute classic computer vision algorithms, as they are efficient at object detection and operate with integer mathematics.
[0125] According to one embodiment of the technology, the PVA is used, for example, to perform computer stereo vision. In some examples, a semi-global matching-based algorithm can be used, although this is not intended as a limitation. Many applications for Level 3-5 autonomous driving require spontaneous motion estimation / stereo matching (e.g., structure from motion, pedestrian detection, lane detection, etc.). The PVA can perform computer stereo vision on input from two monocular cameras.
[0126] In some examples, the PVA can be used to perform dense optical flow processing. This involves processing raw radar data (e.g., using a 4D Fast Fourier Transform) to provide processed radar data. In other examples, the PVA is used for time-of-flight depth processing, for instance, by processing raw time-of-flight data to provide processed time-of-flight data.
[0127] The DLA can be used to power any type of network to improve control and driving safety; this includes, for example, a neural network that outputs a confidence score for each object detection. Such a confidence score can be interpreted as a probability or as providing a relative "weighting" of each detection compared to other detections. This confidence score allows the system to make further decisions about which detections should be considered true positives and not false positives. For example, the system can set a confidence threshold and consider only those detections that exceed the threshold as true positives.In an automatic emergency braking (AEB) system, false positive detections would cause the vehicle to automatically initiate emergency braking, which is obviously undesirable. Therefore, only the safest detections should be considered as triggers for AEB. The DLA can run a neural network for confidence value regression. The neural network can use as input at least a subset of parameters, such as the dimensions of the boundary frame, the ground plane estimate (obtained, for example, from another subsystem), the output of the inertial measurement unit (IMU) sensor 466, correlated with the vehicle's orientation 400, distance, and 3D position estimates of the object obtained from the neural network and / or other sensors (e.g., a LIDAR sensor 464 or a RADAR sensor 460).
[0128] The SoC(s) 404 may contain the datastore(s) 416 (e.g., main memory). The datastore(s) 416 may be on-chip main memory on the SoC(s) 404, capable of storing neural networks to be executed on the GPU and / or DLA. In some examples, the datastore(s) 416 may be large enough to store multiple instances of neural networks for redundancy and security. The datastore(s) 412 may include L2 or L3 cache(s) 412. The reference to the datastore(s) 416 may include a reference to the main memory allocated to the PVA, DLA, and / or other accelerator(s) 414, as described herein.
[0129] The SoC(s) 404 can contain one or more Processors 410 (e.g., embedded processors). The Processor(s) 410 can contain a Boot and Power Management Processor, which may be a dedicated processor and subsystem to handle boot power and management functions and associated security enforcement. The Boot and Power Management Processor can be part of the SoC(s) 404's boot sequence and can provide runtime power management services. The Boot and Power Management Processor can provide clock and voltage programming, support for system transitions to a low-power state, management of the SoC(s) 404's thermals and temperature sensors, and / or management of the SoC(s) 404's power states.Each temperature sensor can be implemented as a ring oscillator whose output frequency is proportional to the temperature, and the SoC(s) 404 can use the ring oscillators to detect the temperatures of the CPU(s) 406, the GPU(s) 408, and / or the accelerator(s) 414. If it is determined that the temperatures exceed a threshold, the boot and power management processor can enter a temperature fault routine and put the SoC(s) 404 into a reduced-power state and / or put the vehicle 400 into a chauffeur-to-safe-stop mode (e.g., bring the vehicle 400 to a safe stop).
[0130] The 410 processor(s) may also include a number of embedded processors that can serve as an audio processing engine. The audio processing engine can be an audio subsystem that provides full hardware support for multi-channel audio across multiple interfaces and a wide and flexible range of audio I / O interfaces. In some examples, the audio processing engine is a dedicated processor core with a digital signal processor and dedicated RAM.
[0131] The 410 processor(s) may also include an always-on processor machine that provides the necessary hardware functions to support low-power sensor management and wake-up of use cases. The always-on processor machine may include a processor core, tightly coupled RAM, supporting peripherals (such as timers and interrupt controllers), various I / O controller peripherals, and routing logic.
[0132] The 410 processor(s) can further include a safety cluster machine, which contains a dedicated processor subsystem for handling the safety management of automotive applications. The safety cluster machine can include two or more processor cores, tightly coupled RAM, supporting peripherals (e.g., timers, an interrupt controller, etc.), and / or routing logic. In a safety mode, the two or more cores can operate in lockstep mode and function as a single core with comparison logic that detects any differences between their operations.
[0133] The processor(s) 410 may also contain a real-time camera machine, which may include a dedicated processor subsystem for managing the real-time camera.
[0134] The 410 processor(s) may also include a high dynamic range signal processor, which may include an image signal processor, which is a hardware machine that is part of the camera processing pipeline.
[0135] The processor(s) 410 may include a video image compositor, which may be a processing block (e.g., implemented on a microprocessor) that implements video post-processing functions required by a video playback application to generate the final image for the player window. The video image compositor may perform lens distortion correction on the wide-angle camera(s) 470, the ambient light camera(s) 474, and / or on the sensors of the cabin surveillance camera. The cabin surveillance camera sensor is preferably monitored by a neural network running on a separate instance of the extended SoC and configured to detect events in the cabin and respond accordingly.A system in the cabin can lip-read to activate mobile service and make a call, dictate emails, change the destination, activate or change the infotainment system and vehicle settings, or enable voice-controlled internet browsing. Certain functions are only available to the driver when the vehicle is operating in autonomous mode and are otherwise deactivated.
[0136] The video image compositor can include enhanced temporal noise reduction for both spatial and temporal noise reduction. For example, if there is motion in a video, the noise reduction weights the spatial information accordingly and reduces the weight of information provided by neighboring frames. If a frame or a portion of a frame does not contain motion, the temporal noise reduction performed by the video image compositor can use information from the previous frame to reduce noise in the current frame.
[0137] The video image compositor can also be configured to perform stereo equalization of the input stereo lens images. Furthermore, the video image compositor can be used for user interface design when the operating system desktop is in use and the GPU(s) 408 do not need to constantly render new surfaces. Even when the GPU(s) 408 are powered on and actively performing 3D rendering, the video image compositor can be used to offload the GPU(s) 408, thus improving performance and responsiveness.
[0138] The SoC(s) 404 may further include a serial camera interface with a Mobile Industry Processor Interface (MIPI) for receiving video and camera input, a high-speed interface, and / or a video input block that can be used for camera and related pixel input functions. The SoC(s) 404 may also include an input / output controller that can be software-controlled and used for receiving I / O signals not assigned to a specific role.
[0139] The SoC(s) 404 can also include a wide range of peripheral interfaces to enable communication with peripheral devices, audio codecs, power management, and / or other devices. The SoC(s) 404 can be used to process data from cameras (e.g., via Gigabit Multimedia Serial Link and Ethernet), sensors (e.g., one LiDAR sensor 464, one RADAR sensor 460, etc., which may be connected via Ethernet), data from the 402 bus (e.g., vehicle speed 400, steering wheel position, etc.), and data from one GNSS sensor 458 (e.g., connected via Ethernet or CAN bus). The SoC(s) 404 may also include dedicated high-performance mass storage controllers, which may contain their own DMA machines and which can be used to offload routine data management tasks from the CPU(s) 406.
[0140] The SoC(s) 404 can be an end-to-end platform with a flexible architecture spanning automation levels 3-5, thereby providing a comprehensive functional safety architecture that supports and efficiently utilizes computer vision and ADAS techniques for diversity and redundancy, and provides a platform for a flexible, reliable driving software stack along with deep learning tools. The SoC(s) 404 can be faster, more reliable, and even more energy- and space-efficient than conventional systems. For example, the accelerator(s) 414, in combination with the CPU(s) 406, the GPU(s) 408, and the data storage(s) 416, can provide a fast, efficient platform for autonomous vehicles of levels 3-5.
[0141] This technology thus provides capabilities and functions that cannot be achieved by conventional systems. For example, computer vision algorithms can be run on CPUs that can be configured using a high-level programming language, such as C, to execute a wide variety of processing algorithms on a wide variety of visual data. However, CPUs are often unable to meet the performance requirements of many computer vision applications, such as execution time and power consumption. In particular, many CPUs are unable to execute complex object detection algorithms in real time, which is a requirement for in-vehicle ADAS applications and a prerequisite for practical Level 3-5 autonomous vehicles.
[0142] In contrast to conventional systems, the technology described herein, by providing a CPU complex, a GPU complex, and a hardware acceleration cluster, enables the simultaneous and / or sequential execution of multiple neural networks and the combination of their results to enable Level 3-5 autonomous driving functionality. For example, a CNN running on the DLA or the dGPU (e.g., the GPU(s) 420) can include text and word recognition, allowing the supercomputer to read and understand traffic signs, including signs for which the neural network has not been specifically trained. The DLA can further include a neural network capable of identifying and interpreting the sign, providing a semantic understanding, and passing this semantic understanding to the path planning modules running on the CPU complex.
[0143] Another example is that multiple neural networks can run simultaneously, as required for driving at levels 3, 4, or 5. For instance, a warning sign reading "Caution: Flashing lights indicate black ice" accompanied by an electric light can be interpreted independently or jointly by several neural networks. The sign itself can be identified as a traffic sign by a first neural network (e.g., a trained one), while the text "Flashing lights indicate black ice" can be interpreted by a second neural network, which then informs the vehicle's path planning software (preferably running on the CPU) that the presence of black ice indicates the presence of flashing lights.The flashing light can be identified by operating a third neural network across multiple images, which informs the vehicle's path planning software about the presence (or absence) of flashing lights. All three neural networks can run simultaneously, as within the DLA and / or on the GPU(s) 408.
[0144] In some examples, a CNN for facial recognition and vehicle owner identification can use data from camera sensors to identify the presence of an authorized driver and / or owner of the vehicle 400. The always-on sensor processing unit can be used to unlock the vehicle and turn on the lights when the owner approaches the driver's door, and to disable the vehicle in security mode when the owner leaves. In this way, the SoC(s) 404 provide security against theft and / or carjacking.
[0145] In another example, a CNN for emergency vehicle detection and identification can use data from microphones 496 to detect and identify emergency vehicle sirens. Unlike conventional systems that use general classifiers to detect sirens and manually extract features, the SoC(s) 404 uses the CNN for classifying environmental and urban noise as well as visual data. In a preferred embodiment, the CNN running on the DLA is trained to identify the relative approach speed of the emergency vehicle (e.g., by using the Doppler effect). The CNN can also be trained to identify emergency vehicles specific to the local area in which the vehicle is operating, as identified by a GNSS sensor 458.For example, the CNN will attempt to detect European sirens when operating in Europe, and when operating in the United States, the CNN will attempt to identify only North American sirens. Once an emergency vehicle is detected, a controller can be used to execute an emergency vehicle safety routine, slowing the vehicle down, pulling over to the side of the road, parking the vehicle, and / or letting the vehicle idle, using the 462 ultrasonic sensors, until the emergency vehicle(s) pass.
[0146] The vehicle may contain a CPU(s) 418 (e.g., a discrete CPU or a dCPU) which may be coupled to the SoC(s) 404 via a high-speed connection (e.g., PCIe). The CPU(s) 418 may, for example, contain an x86 processor. The CPU(s) 418 may, for example, be used to perform a variety of functions, including reconciling potentially inconsistent results between ADAS sensors and the SoC(s) 404 and / or monitoring the status and health of the Controller(s) 436 and / or the Infotainment SoC 430.
[0147] The Vehicle 400 can include a GPU(s) 420 (e.g., a discrete GPU(s) or a dGPU(s)) which can be coupled to the SoC(s) 404 via a high-speed connection (e.g., NVIDIA's NVLINK). The GPU(s) 420 can provide additional artificial intelligence functions, such as by executing redundant and / or different neural networks, and can be used to train and / or update neural networks based on inputs (e.g., sensor data) from the Vehicle 400's sensors.
[0148] The vehicle 400 may also include the network interface 424, which may contain one or more wireless antennas 426 (e.g., one or more wireless antennas for different communication protocols, such as a cellular antenna, a Bluetooth antenna, etc.). The network interface 424 can be used to enable a wireless connection over the internet to the cloud (e.g., to the server(s) 478 and / or other network devices), to other vehicles, and / or to computing devices (e.g., passenger client devices). To communicate with other vehicles, a direct connection between the two vehicles and / or an indirect connection (e.g., via networks and the internet) can be established. Direct connections can be provided using vehicle-to-vehicle communication.Vehicle-to-vehicle communication can provide the vehicle 400 with information about vehicles in its vicinity (e.g., vehicles in front of, beside, and / or behind the vehicle 400). This functionality can be part of a cooperative adaptive cruise control function of the vehicle 400.
[0149] The network interface 424 can include a system-on-a-chip (SoC) that provides modulation and demodulation functions, enabling the controller(s) 436 to communicate over wireless networks. The network interface 424 can include a high-frequency (RF) front end for upconversion from baseband to RF and downconversion from RF to baseband. The frequency conversions can be performed using well-known methods and / or superheterodyne techniques. In some examples, the RF front-end functionality can be provided by a separate chip. The network interface can include wireless functionality for communication over LTE, WCDMA, UMTS, GSM, CDMA2000, Bluetooth, Bluetooth LE, Wi-Fi, Z-Wave, ZigBee, LoRaWAN, and / or other wireless protocols.
[0150] The vehicle 400 may further include a data storage device 428, which may be located outside the chip (e.g., outside the SoC(s) 404). The data storage device(s) 428 may contain one or more memory elements, including RAM, SRAM, DRAM, VRAM, flash, hard disks, and / or other components and / or devices capable of storing at least one bit of data.
[0151] The vehicle 400 can also include one or more GNSS sensors (458). The GNSS sensor(s) (e.g., GPS, supported GPS sensors, differential GPS (DGPS) sensors, etc.) assists with mapping, perception, occupancy grid creation, and / or path planning functions. Any number of GNSS sensors (458) can be used, including, for example, a single GPS unit that uses a USB connection with an Ethernet-to-serial (RS-232) bridge.
[0152] The vehicle 400 may also include a RADAR sensor 460. The RADAR sensor 460 can be used by the vehicle 400 for long-range vehicle detection, even in darkness and / or adverse weather conditions. The functional safety level of the RADAR can be ASIL B. The RADAR sensor 460 can use the CAN bus and / or the 402 bus (e.g., for transmitting the data generated by the RADAR sensor 460) for control and access to object tracking data, with some examples using Ethernet for access to the raw data. A wide variety of RADAR sensor types can be used. The RADAR sensor 460 can be suitable for front, rear, and side RADAR applications without restriction. In some examples, a pulse-Doppler RADAR sensor is used.
[0153] The RADAR sensor(s) 460 can include various configurations, such as long-range with a narrow field of view, short-range with a wide field of view, side coverage with short-range, etc. In some examples, long-range RADAR can be used for the adaptive cruise control function. Long-range RADAR systems can provide a wide field of view, achieved through two or more independent scans, such as at a range of 250 m. The RADAR sensor(s) 460 can assist in distinguishing between static and moving objects and can be used by ADAS systems for emergency braking and forward collision warning. Long-range RADAR sensors can be a monostatic multimodal RADAR with multiple (e.g.,The system includes six or more fixed radar antennas and a high-speed CAN and FlexRay interface. In a six-antenna example, the four central antennas can create a focused beam pattern designed to detect the area around Vehicle 400 at higher speeds with minimal interference from traffic in adjacent lanes. The other two antennas can extend the field of view, enabling the rapid detection of vehicles entering or exiting Vehicle 400's lane.
[0154] Medium-range radar systems, for example, can have a range of up to 460 m (front) or 80 m (rear) and a field of view of up to 42 degrees (front) or 450 degrees (rear). Short-range radar systems can, without restriction, include radar sensors designed for installation at both ends of the rear bumper. When such a radar sensor system is installed at both ends of the rear bumper, it can generate two beams that continuously monitor the blind spot in the area behind and beside the vehicle.
[0155] Short-range radar systems can be used in an ADAS system for blind spot detection and / or as a lane change assistant.
[0156] The vehicle 400 may also include one or more ultrasonic sensors 462. The ultrasonic sensor(s) 462, which may be mounted on the front, rear, and / or sides of the vehicle 400, may be used for parking assistance and / or for creating and updating an occupancy grid. A wide variety of ultrasonic sensors 462 may be used, and different ultrasonic sensors 462 may be used for different detection ranges (e.g., 2.5 m, 4 m). The ultrasonic sensor(s) 462 may operate with functional safety levels of ASIL B.
[0157] The vehicle 400 can contain LiDAR sensor(s) 464. The LiDAR sensor(s) 464 can be used for object and pedestrian detection, emergency braking, collision avoidance, and / or other functions. The LiDAR sensor(s) 464 can meet functional safety level ASIL B. In some examples, the vehicle 400 can contain multiple LiDAR sensors 464 (e.g., two, four, six, etc.) that can use Ethernet (e.g., to provide data to a Gigabit Ethernet switch).
[0158] In some examples, the LIDAR sensor(s) 464 may be able to provide a list of objects and their distances for a 360-degree field of view. Commercially available LIDAR sensor(s) 464 may, for example, have a specified range of approximately 400 m, with an accuracy of 2 cm to 3 cm, and support for a 400 Mbit / s Ethernet connection. In some examples, one or more non-protruding LIDAR sensors 464 may be used. In such examples, the LIDAR sensor(s) 464 may be implemented as a small device that can be embedded in the front, rear, sides, and / or corners of the vehicle 400. The LIDAR sensor(s) 464 can provide a horizontal field of view of up to 120 degrees and a vertical field of view of up to 35 degrees in such examples, with a range of 200 m, even with objects of low reflectivity.The front-mounted LIDAR sensor(s) 464 can be configured for a horizontal field of view between 45 degrees and 135 degrees.
[0159] In some examples, LiDAR technologies, such as 3D flash LiDAR, can also be used. 3D flash LiDAR uses a laser pulse as a transmission source to illuminate the vehicle's surroundings up to approximately 200 m. A flash LiDAR unit contains a sensor that records the travel time of the laser pulse and the reflected light at each pixel, which in turn corresponds to the distance between the vehicle and objects. Flash LiDAR can enable the generation of highly accurate and distortion-free images of the surroundings with each laser pulse. In some examples, four flash LiDAR sensors can be used, one on each side of the vehicle. Available 3D flash LiDAR systems include a solid-state 3D focal plane array LiDAR camera that contains no moving parts other than a fan (e.g., a non-scanning LiDAR device).The flash LIDAR device can use a 5-nanosecond pulse of a Class I (eye-safe) laser per frame and capture the reflected laser light in the form of 3D distance point clouds and co-registered intensity data. By using flash LIDAR, and because flash LIDAR is a solid-state device with no moving parts, the LIDAR sensor(s) 464 may be less susceptible to motion blur, vibration, and / or shock.
[0160] The vehicle may further include an IMU sensor(s) 466. The IMU sensor(s) 466 may be located in the center of the rear axle of the vehicle 400 in some examples. The IMU sensor(s) 466 may, for example, and without limitation, include an accelerometer, a magnetometer, a gyroscope, a magnetic compass, and / or other sensor types. In some examples, such as six-axis applications, the IMU sensor(s) 466 may include accelerometers and gyroscopes, while in nine-axis applications, the IMU sensor(s) 466 may include accelerometers, gyroscopes, and magnetometers.
[0161] In some embodiments, the IMU sensor(s) 466 can be implemented as a miniaturized, high-performance GPS-aided inertial navigation system (GPS / INS) that combines inertial sensors of a microelectromechanical system (MEMS), a high-sensitivity GPS receiver, and advanced Kalman filter algorithms to provide estimates of position, velocity, and orientation. Thus, in some examples, the IMU sensor(s) 466 can enable the vehicle 400 to estimate its course without requiring input from a magnetic sensor by directly observing and correlating velocity changes from the GPS with the IMU sensor(s) 466. In some examples, the IMU sensor(s) 466 and the GNSS sensor(s) 458 can be combined in a single integrated unit.
[0162] The vehicle may contain microphone(s) 496, which is / are mounted in and / or around the vehicle 400. The microphone(s) 496 may be used, among other things, for the detection and identification of emergency vehicles.
[0163] The vehicle can further include any number of camera types, including stereo camera(s) 468, wide-angle camera(s) 470, infrared camera(s) 472, surround-view camera(s) 474, long-range and / or medium-range camera(s) 498, and / or other camera types. The cameras can be used to capture image data around the entire periphery of the vehicle 400. The types of cameras used depend on the embodiment and requirements of the vehicle 400, and any combination of camera types can be used to provide the necessary coverage around the vehicle 400. Furthermore, the number of cameras can vary depending on the embodiment. For example, the vehicle can include six cameras, seven cameras, ten cameras, twelve cameras, and / or any other number of cameras.The cameras can, for example and without limitation, support Gigabit Multimedia Serial Link (GMSL) and / or Gigabit Ethernet. Each of the camera(s) is described herein with reference to... Fig. 4A and Fig. 4B is described in more detail.
[0164] The vehicle 400 may also include a vibration sensor 442. The vibration sensor 442 may measure vibrations of vehicle components, such as the axle(s). For example, changes in vibration may indicate a change in the road surface. In another example, if two or more vibration sensors 442 are used, the differences between the vibrations may be used to determine the friction or slip on the road surface (e.g., if the difference in vibration is between a driven axle and a freely rotating axle).
[0165] The vehicle 400 may include an ADAS system 438. The ADAS system 438 may include a SoC in some examples. The ADAS system 438 may include autonomous / adaptive / automatic cruise control (ACC), cooperative adaptive cruise control (CACC), forward crash warning (FCW), automatic emergency braking (AEB), lane departure warning (LDW), lane keep assist (LKA), blind spot warning (BSW), rear cross-traffic warning (RCTW), collision warning systems (CWS), lane centering (LC), and / or other features and functions.
[0166] The ACC systems can use one or more radar sensors (460), one or more lidar sensors (464), and / or one or more cameras. The ACC systems can include longitudinal and / or lateral ACC. Longitudinal ACC monitors and controls the distance to the vehicle immediately in front of the vehicle (400) and automatically adjusts the vehicle speed to maintain a safe distance from vehicles ahead. Lateral ACC performs distance control and advises the vehicle (400) to change lanes if necessary. Lateral ACC is interrelated with other ADAS applications, such as LCA and CWS.
[0167] The CACC uses information from other vehicles, which can be received via the network interface 424 and / or the wireless antenna(s) 426 from other vehicles either wirelessly or indirectly via a network connection (e.g., the internet). Direct connections can be provided via a vehicle-to-vehicle (V2V) communication link, while indirect connections can be an infrastructure-to-vehicle (I2V) communication link. Generally, the V2V communication concept provides information about the vehicles immediately ahead (e.g., vehicles directly in front of the vehicle 400 and in the same lane), while the I2V communication concept provides information about traffic further ahead. CACC systems can incorporate both I2V and V2V information sources.Given the information about the vehicles ahead of vehicle 400, the CACC can be more reliable and has the potential to improve traffic flow and reduce congestion on the road.
[0168] FCW systems are designed to warn the driver of a hazard, allowing them to take corrective action. FCW systems utilize a forward-facing camera and / or RADAR sensor(s) 460 coupled with a dedicated processor, DSP, FPGA, and / or ASIC, which is electrically linked to the driver feedback system, such as a display, speaker, and / or vibrating component. FCW systems can provide a warning in the form of an audible signal, a visual warning, a vibration, and / or a rapid braking pulse.
[0169] AEB systems detect an impending forward collision with another vehicle or object and can automatically apply the brakes if the driver does not take corrective action within a specified time or distance parameter. AEB systems may use forward-facing camera(s) and / or radar sensor(s) coupled with a dedicated processor, DSP, FPGA, and / or ASIC. When the AEB system detects a hazard, it typically first warns the driver so they can take corrective action to avoid the collision. If the driver fails to take corrective action, the AEB system can automatically apply the brakes to prevent or at least mitigate the effects of the predicted collision. AEB systems may incorporate techniques such as dynamic brake assist and / or emergency braking for an impending collision.
[0170] Lane Departure Warning (LDW) systems provide visual, audible, and / or tactile warnings, such as steering wheel or seat vibrations, to alert the driver if the vehicle crosses lane markings. An LDW system will not activate if the driver indicates an intentional lane departure by using a turn signal. LDW systems may use forward-facing cameras coupled to a dedicated processor, DSP, FPGA, and / or ASIC, which is electrically connected to the feedback device for the driver, such as a display, speaker, and / or vibrating component.
[0171] LKA systems are a variant of LDW systems. LKA systems provide steering or braking inputs to correct vehicle 400 if the vehicle 400 begins to leave its lane.
[0172] Blind Spot Warning (BSW) systems detect and warn the driver of vehicles in the car's blind spot. BSW systems can provide a visual, audible, and / or tactile warning signal to indicate that merging into or changing lanes is unsafe. The system can provide an additional warning when the driver activates a turn signal. BSW systems can use a rear-facing camera(s) and / or radar sensor(s) coupled to a dedicated processor, DSP, FPGA, and / or ASIC, which is electrically coupled to the feedback to the driver, such as a display, speaker, and / or vibrating component.
[0173] RCTW systems can provide visual, audible, and / or tactile alerts when an object is detected outside the reversing camera's field of view while the vehicle is reversing. Some RCTW systems incorporate AEB to ensure the vehicle's brakes are applied to prevent a collision. RCTW systems can utilize one or more rear-facing radar sensors coupled with a dedicated processor, DSP, FPGA, and / or ASIC, which is electrically connected to the driver for feedback, such as a display, speaker, and / or vibrating component.
[0174] Conventional ADAS systems can produce false positives, which, while annoying and distracting for the driver, are generally not catastrophic because the ADAS systems warn the driver and give them the opportunity to decide whether a safety issue truly exists and to act accordingly. However, in an autonomous vehicle 400, the vehicle 400 itself must decide, in the event of conflicting results, whether to follow the result from a primary computer or a secondary computer (e.g., a first controller 436 or a second controller 436). In some embodiments, the ADAS system 438 can, for example, be a backup and / or secondary computer that provides information about perception to a rationality module of the backup computer.The rationality monitor of the backup computer can run redundant, diverse software on hardware components to detect errors in perception and dynamic driving tasks. The outputs of the ADAS system 438 can be provided to a monitoring MCU. If the outputs of the primary and secondary computers conflict, the monitoring MCU must determine how to resolve the conflict to ensure safe operation.
[0175] In some examples, the primary computer can be configured to provide the monitoring MCU with a confidence score indicating its confidence in the chosen outcome. If the confidence score exceeds a certain threshold, the monitoring MCU can follow the primary computer's instruction, regardless of whether the secondary computer provides a conflicting or inconsistent result. If the confidence score does not reach the threshold and the primary and secondary computers display different results (e.g., conflicting results), the monitoring MCU can mediate between the computers to determine the appropriate outcome.
[0176] The monitoring MCU can be configured to run a neural network(s) trained and configured to determine, based on the output of the primary and secondary computers, the conditions under which the secondary computer will provide false alarms. Thus, the neural network(s) in the monitoring MCU can learn when the output of the secondary computer can be trusted and when it cannot. For example, if the secondary computer is a radar-based FCW system, a neural network in the monitoring MCU can learn when the FCW system identifies metallic objects that do not actually pose a threat, such as a drain grate or manhole cover, triggering an alarm.Similarly, if the secondary computer is a camera-based lane departure warning (LDW) system, a neural network in the monitoring MCU can learn to override the LDW system when cyclists or pedestrians are present and leaving the lane is indeed the safest maneuver. In embodiments that include a neural network(s) running on the monitoring MCU, the monitoring MCU may include at least one DLA or GPU suitable for executing the neural network(s) with allocated memory. In preferred embodiments, the monitoring MCU may include and / or be included as a component of the SoC 404.
[0177] In other examples, the ADAS system 438 can include a secondary computer that executes the ADAS functionality using classical computer vision rules. Thus, the secondary computer can employ classical computer vision (if-then) rules, and the presence of a neural network(s) in the monitoring MCU can improve reliability, safety, and performance. For example, the diverse implementation and intentional non-identity make the overall system more fault-tolerant, particularly to errors caused by software functionality (or software-hardware interfaces).For example, if a software bug or error occurs in the software running on the primary computer, and the non-identical software code running on the secondary computer provides the same overall result, the monitoring MCU can have greater confidence that the overall result is correct and that the bug in the software or hardware on the primary computer is not causing a material error.
[0178] In some examples, the output of the ADAS system 438 can be fed into the perception block and / or the dynamic driving task block of the primary computer. For example, if the ADAS system 438 displays a forward collision warning due to an object immediately in front of the vehicle, the perception block can use this information in object identification. In other examples, the secondary computer may have its own trained neural network, thus reducing the risk of false positives, as described herein.
[0179] The Vehicle 400 may also include the Infotainment SoC 430 (e.g., an in-vehicle infotainment system (IVI)). Although illustrated and described as an SoC, the infotainment system may not actually be an SoC and may contain two or more discrete components. The Infotainment SoC 430 may include a combination of hardware and software that can be used to provide the Vehicle 400 with audio (e.g., music, a personal digital assistant, navigation instructions, news, radio, etc.), video (e.g., TV, movies, streaming, etc.), telephone (e.g., hands-free calling), network connectivity (e.g., LTE, Wi-Fi, etc.), and / or information services (e.g., navigation systems, rear parking sensors, a radio data system, vehicle-related information such as fuel level, total distance traveled, brake fluid level, oil level, door open / close status, air filter information, etc.).The Infotainment SoC 430 can, for example, include radios, record players, navigation systems, video players, USB and Bluetooth connectivity, car computers, in-car entertainment, Wi-Fi, steering wheel audio controls, hands-free systems, a head-up display (HUD), an HMI display 434, a telematics device, a control panel (e.g., for controlling and / or interacting with various components, functions, and / or systems), and / or other components. The Infotainment SoC 430 can also be used to provide vehicle occupant(s) with information (e.g., visual and / or audible), such as information from the ADAS system 438, autonomous driving information such as planned vehicle maneuvers, road layouts, environmental information (e.g., intersection information, vehicle information, road information, etc.), and / or other information.
[0180] The infotainment SoC 430 may include GPU functionality. The infotainment SoC 430 can communicate with other devices, systems, and / or components of the vehicle 400 via bus 402 (e.g., CAN bus, Ethernet, etc.). In some examples, the infotainment SoC 430 may be coupled with a monitoring MCU so that the infotainment system's GPU can perform some self-driving functions if the primary controller(s) 436 (e.g., the vehicle 400's primary and / or backup computers) fail. In such an example, the infotainment SoC 430 can place the vehicle 400 into a chauffeur-to-safe-stop mode, as described herein.
[0181] The vehicle 400 may also include an instrument cluster 432 (e.g., a digital instrument cluster, an electronic instrument cluster, a digital instrument panel, etc.). The instrument cluster 432 may include a controller and / or supercomputer (e.g., a discrete controller or supercomputer). The instrument cluster 432 may contain a number of instruments, such as a speedometer, fuel gauge, oil pressure gauge, tachometer, odometer, turn signals, shift position indicator, seat belt warning light(s), parking brake warning light(s), engine malfunction light(s), airbag system (SRS) information, lighting controls, safety system controls, navigation information, etc. In some examples, information from the infotainment SoC 430 and the instrument cluster 432 may be displayed and / or shared. In other words, the instrument cluster 432 may be included as part of the infotainment SoC 430, or vice versa.
[0182] Fig. 4D is a system diagram for the communication between the cloud-based server(s) and the exemplary autonomous vehicle 400. Fig. 4A according to some embodiments of the present disclosure. The system 476 may include the server(s) 478, the network(s) 490, and the vehicles, including the vehicle 400. The server(s) 478 may include multiple GPUs 484(A)-484(H) (hereinafter collectively referred to as GPUs 484), PCIe switches 482(A)-482(H) (hereinafter collectively referred to as PCIe switches 482), and / or CPUs 480(A)-480(B) (hereinafter collectively referred to as CPUs 480). The GPUs 484, the CPUs 480, and the PCIe switches may be interconnected by high-speed links, such as, but not limited to, NVIDIA's NVLink interfaces 488 and / or PCIe links 486. In some examples, the GPUs 484 are connected via NVLink and / or NVSwitch SoCs, and the GPUs 484 and PCIe switches 482 are connected via PCIe links. Although eight GPUs 484, two CPUs 480, and two PCIe switches are illustrated, this should not be interpreted as a limitation.Depending on the configuration, each Server 478 can contain any number of GPUs 484, CPUs 480, and / or PCIe switches. For example, each Server 478 can contain eight, sixteen, thirty-two, and / or more GPUs 484.
[0183] The server(s) 478 can receive image data from the network(s) 490 and the vehicles, depicting unexpected or altered road conditions, such as recently commenced roadworks. The server(s) 478 can transmit neural networks 492, updated neural networks 492, and / or map information 494 to the vehicles via the network(s) 490 and the vehicles. This map information includes traffic and road condition updates. Map information updates 494 may include updates to the HD map 422, such as information about construction sites, potholes, detours, flooding, and / or other obstructions.In some examples, the neural networks 492, the updated neural networks 492 and / or the map information 494 may result from new training and / or experience represented in the data received from any number of vehicles in the environment, and / or may be based on training performed in a data center (e.g. using the server(s) 478 and / or other servers).
[0184] Server 478 can be used to train machine learning models (e.g., neural networks) based on training data. The training data can be generated by the vehicles and / or in a simulation (e.g., using a game machine). In some examples, the training data is tagged (e.g., if the neural network benefits from supervised learning) and / or subjected to other preprocessing, while in other examples, the training data is not tagged and / or preprocessed (e.g., if the neural network does not require supervised learning).Training can be performed using one or more classes of machine learning techniques, including, but not limited to, classes such as: supervised training, semi-supervised training, unsupervised training, self-learning, reinforcement learning, federated learning, transfer learning, feature learning (including principal component and cluster analysis), multilinear subspace learning, diverse learning, representational learning (including substitute dictionary learning), rule-based machine learning, anomaly detection, and all variants or combinations thereof. Once the machine learning models are trained, they can be used by the vehicles (e.g., transmitted to the vehicles via network(s) 490) and / or used by the server(s) 478 for remote monitoring of the vehicles.
[0185] In some examples, the Server 478 can receive data from the vehicles and apply that data to current neural networks in real time for real-time intelligent inference. The Server 478 can include deep learning supercomputers and / or dedicated AI computers powered by GPUs 484, such as NVIDIA's DGX and DGX Station machines. However, in some examples, the Server 478 can include a deep learning infrastructure that uses only CPU-powered data centers.
[0186] The deep learning infrastructure of server(s) 478 is capable of performing fast, real-time inference and can use this capability to assess and verify the state of the processors, software, and / or associated hardware in vehicle 400. For example, the deep learning infrastructure can receive periodic updates from vehicle 400, such as a sequence of images and / or objects that vehicle 400 has located within that sequence (e.g., via computer vision and / or other machine learning object classification techniques).The deep learning infrastructure can run its own neural network to identify the objects and compare them with the objects identified by vehicle 400. If the results do not match and the infrastructure concludes that the AI in vehicle 400 is not working correctly, server 478 can send a signal to vehicle 400, instructing a fail-safe computer in vehicle 400 to take control, notify the passengers, and perform a safe parking maneuver.
[0187] For inference, the server(s) 478 can include the GPU(s) 484 and one or more programmable inference accelerators (e.g., NVIDIA's TensorRT). The combination of GPU-driven servers and inference accelerators can enable real-time responsiveness. In other scenarios, such as when performance is less critical, servers powered by CPUs, FPGAs, and other processors can be used for inference. EXAMPLE CALCULATION DEVICE
[0188] Fig. Figure 5 is a block diagram of an exemplary computing device(s) 500 suitable for use in implementing at least some embodiments of the present disclosure. The computing device 500 may include a connection system 502 that directly or indirectly couples the following devices: main memory 504, one or more central processing units (CPUs) 506, one or more graphics processing units (GPUs) 508, a communication interface 510, input / output (I / O) ports 512, input / output components 514, a power supply 516, one or more presentation components 518 (e.g., display(s)), and one or more logic units 520. In at least one embodiment, the computing device(s) 500 may include one or more virtual machines (VMs), and / or each of its components may include virtual components (e.g., virtual hardware components).As non-restrictive examples, one or more of the GPUs 508 can comprise one or more vGPUs, one or more of the CPUs 506 can comprise one or more vCPUs, and / or one or more of the logic units 520 can comprise one or more virtual logic units. Thus, a compute device (or devices) 500 can contain discrete components (e.g., a complete GPU allocated to the compute device 500), virtual components (e.g., a portion of a GPU allocated to the compute device 500), or a combination thereof. For example, the unpartitioned mode of the system 100 can run on the discrete components, while the partitioned mode of the system 100 runs on virtual components.
[0189] Although the various blocks of Fig. Where components 5 are shown connected via the connection system 502, this is not intended as a limitation and is for clarity only. In some embodiments, for example, a presentation component 518, such as a display device, may be considered an I / O component 514 (e.g., if the display is a touchscreen). As another example, the CPUs 506 and / or GPUs 508 may contain memory (e.g., the memory 504 may represent a storage device in addition to the memory of the GPUs 508, the CPUs 506, and / or other components). In other words, the computing device of Fig. Figure 5 is for illustrative purposes only. No distinction is made between categories such as "workstation", "server", "laptop", "desktop", "tablet", "client device", "mobile device", "handheld device", "game console", "electronic control unit (ECU)", "virtual reality system" and / or other device or system types, as all fall within the scope of the computing device of Fig. 5.
[0190] The 502 interconnect system can represent one or more connections or buses, such as an address bus, a data bus, a control bus, or a combination thereof. The 502 interconnect system can include one or more bus or interconnect types, such as an Industry Standard Architecture (ISA) bus, an Extended ISA bus, a Video Electronics Standards Association (VESA) bus, a Peripheral Component Interconnect (PCI) bus, a PCI Express (PCIe) bus, and / or another type of bus or interconnect. In some embodiments, there are direct connections between components. For example, the CPU 506 can be directly connected to the memory 504. Furthermore, the CPU 506 can be directly connected to the GPU 508.In a direct or point-to-point connection between components, the 502 connection system can include a PCIe link to establish the connection. In these examples, a PCI bus does not need to be included in the 500 computing device.
[0191] The 504 main memory can contain any of a variety of computer-readable media. Computer-readable media can be any available media that the 500 computing device can access. Computer-readable media can include both volatile and non-volatile media, and removable and non-removable media. For example, and without limitation, computer-readable media can include computer storage media and communication media.
[0192] Computer storage media can include both volatile and non-volatile media, and / or removable and non-removable media, implemented in any method or technology for storing information, such as computer-readable instructions, data structures, program modules, and / or other types of data. For example, main memory can store 504 computer-readable instructions (e.g., representing a program and / or program element, such as an operating system).Computer storage media may, but are not limited to, include RAM, ROM, EEPROM, flash memory or other memory technologies, CD-ROM, Digital Versatile Discs (DVDs) or other optical disk storage, magnetic cartridges, magnetic tapes, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to store the desired information and that the Computing Device 500 can access. As used herein, computer storage media do not per se include signals.
[0193] Computer storage media can embody computer-readable instructions, data structures, program modules, and / or other types of data in a modulated data signal, such as a carrier wave or other transport mechanism, and may include any media for transmitting information. The term "modulated data signal" can refer to a signal in which one or more of its properties are set or modified to encode information within the signal. Computer storage media may include, but are not limited to, wired media, such as a wired network or a direct-wired connection, and wireless media, such as acoustic, RF, infrared, and other wireless media. Combinations of the foregoing should also be included in the scope of protection of the computer-readable media.
[0194] The CPU(s) 506 can be configured to execute at least some of the computer-readable instructions to control one or more components of the Computing Device 500 to perform one or more of the procedures and / or processes described herein. The CPU(s) 506 can each contain one or more cores (e.g., one, two, four, eight, twenty-eight, seventy-two, etc.) capable of processing a plurality of software threads simultaneously. The CPU(s) 506 can contain any type of processor and may contain different types of processors depending on the type of Computing Device 500 implemented (e.g., processors with fewer cores for mobile devices and processors with more cores for servers).Depending on the type of computing device 500, the processor can be, for example, an Advanced RISC Machine (ARM) processor implemented using Reduced Instruction Set Computing (RISC), or an x86 processor implemented using Complex Instruction Set Computing (CISC). The computing device 500 can contain one or more CPUs 506, in addition to one or more microprocessors or additional coprocessors, such as mathematical coprocessors.
[0195] In addition to or as an alternative to the CPU(s) 506, the CPU(s) 508 may be configured to execute at least some of the computer-readable instructions to control one or more components of the computing device 500 to perform one or more of the procedures and / or processes described herein. One or more of the GPUs 508 may be an integrated GPU (e.g., with one or more of the CPU(s) 506) and / or one or more of the GPUs 508 may be a discrete GPU. In embodiments, one or more of the GPUs 508 may be a coprocessor of one or more of the CPU(s) 506. The GPU(s) 508 may be used by the computing device 500 to render graphics (e.g., 3D graphics) or to perform general-purpose calculations. The GPU(s) 508 can be used, for example, for general-purpose computing on GPUs (GPGPU).The GPU(s) 508 can contain hundreds or thousands of cores capable of processing hundreds or thousands of software threads simultaneously. The GPU(s) 508 can generate pixel data for output images in response to rendering commands (e.g., rendering commands from the CPU(s) 506 received via a host interface). The GPU(s) 508 can include graphics memory, such as display memory, for storing pixel data or other suitable data, such as GPGPU data. The display memory can be included as part of the 504 memory. The GPU(s) 508 can contain two or more GPUs operating in parallel (e.g., via a link). The link can connect the GPUs directly (e.g., using NVLINK) or connect them via a switch (e.g., using NVSwitch).When combined, each GPU can generate 508 pixel data or GPGPU data for different sections of an output or for different outputs (e.g., a first GPU for a first image and a second GPU for a second image). Each GPU can have its own dedicated memory or share memory with other GPUs.
[0196] In addition to or as an alternative to the CPU(s) 506 and / or the GPU(s) 508, the logic unit(s) 520 may be configured to execute at least some of the computer-readable instructions to control one or more components of the computing device 500 to perform one or more of the methods and / or processes described herein. In embodiments, the CPU(s) 506, the GPU(s) 508, and / or the logic unit(s) 520 may discretely or jointly perform any combination of the methods, processes, and / or sections thereof. One or more of the logic units 520 may be part of and / or integrated within one or more of the CPU(s) 506 and / or the GPU(s) 508, and / or one or more of the logic units 520 may be discrete components or otherwise separate from the CPU(s) 506 and / or the GPU(s) 508.In embodiments, one or more of the logic units 520 can be a co-processor of one or more of the CPUs 506 and / or one or more of the GPUs 508. For example, the hardware manager 112 can use the GPU(s) 508 to run applications and the multitude of tasks in the task flow.
[0197] Examples of Logic Unit(s) 520 include one or more processing cores and / or components thereof, such as Data Processing Units (DPUs), Tensor Cores (TCs), Tensor Processing Units (TPUs), Pixel Visual Cores (PVCs), Vision Processing Units (VPUs), Graphics Processing Clusters (GPCs), Texture Processing Clusters (TPCs), Streaming Multiprocessors (SMs), Tree Traversal Units (TTUs), Artificial Intelligence Accelerators (AIAs), Deep Learning Accelerators (DLAs), Arithmetic Logic Units (ALUs), and application-specific integrated circuits. (Application-Specific Integrated Circuits, ASICs), Floating Point Units (FPUs), Input / Output (I / O) elements,Peripheral Component Interconnect (PCI) or PCI Express (PCIe) elements, and / or the like.
[0198] The Communications Interface 510 can include one or more receivers, transmitters, and / or transceivers that enable the Computing Device 500 to communicate with other computers over an electronic network, including wired and / or wireless communication. The Communications Interface 510 can include components and functions that enable communication over a variety of different networks, such as wireless networks (e.g., Wi-Fi, Z-Wave, Bluetooth, Bluetooth LE, ZigBee, etc.), wired networks (e.g., Ethernet or InfiniBand communication), low-power wide-area networks (e.g., LoRaWAN, SigFox, etc.), and / or the Internet.In one or more embodiments, the logic unit(s) 520 and / or the communication interface 510 may contain one or more data processing units (DPUs) to directly transfer data received via a network and / or the connection system 502 to one or more GPUs 508 (e.g., a memory thereof).
[0199] The I / O ports 512 enable the computing device 500 to be logically coupled with other devices, including the I / O components 514, the presentation component(s) 518, and / or other components, some of which may be built into (e.g., integrated with) the computing device 500. Illustrative I / O components 514 include a microphone, mouse, keyboard, joystick, gamepad, game controller, satellite dish, scanner, printer, wireless device, etc. The I / O components 514 can provide a natural user interface (NUI) that processes air gestures, speech, or other physiological inputs generated by a user. In some cases, the inputs can be transmitted to a suitable network element for further processing.A NUI can implement any combination of speech capture, stylus capture, face capture, biometric capture, gesture capture (both on-screen and off-screen), air gestures, head and eye tracking, and touch capture (as described in more detail below) associated with a display of the Computing Device 500. The Computing Device 500 can include depth cameras, such as stereoscopic camera systems, infrared camera systems, RGB camera systems, touchscreen technology, and combinations thereof, for gesture capture and recognition. Additionally, the Computing Device 500 can include accelerometers or gyroscopes (e.g., as part of an inertial measurement unit (IMU)) that enable motion detection. In some examples, the output from the accelerometers or gyroscopes can be used by the Computing Device 500 to render immersive augmented reality or virtual reality.
[0200] The power supply 516 can include a hardwired power supply, a battery power supply, or a combination thereof. The power supply 516 can power the computing device 500 to enable the operation of the computing device 500's components.
[0201] The presentation component(s) 518 can include a display (e.g., a monitor, a touchscreen, a television screen, a head-up display (HUD), other display types, or a combination thereof), speakers, and / or other presentation components. The presentation component(s) 518 can receive data from other components (e.g., the GPU(s) 508, the CPU(s) 506, DPUs, etc.) and output the data (e.g., as an image, video, sound, etc.). EXEMPLARY DATA CENTER
[0202] Fig. Figure 6 illustrates an exemplary data center 600 that can be used in at least one embodiment of the present disclosure. The data center 600 can include an infrastructure layer 610, a framework layer 620, a software layer 630, and / or an application layer 640. The application layer 640 can be application layer 104.
[0203] As in Fig. As shown in Figure 6, the infrastructure layer 610 of the data center can contain a resource orchestrator 612, clustered computer resources 614 and node computer resources (“node CRs”) 616(1)-616(N), where “N” is any positive integer. In at least one embodiment, the node CRs 616(1)-616(N) may include, but are not limited to, any number of central processing units (CPUs) or other processors (including DPUs, accelerators, field programmable gate arrays (FPGAs), graphics processing units or graphics processing units (GPUs), etc.), memory devices (e.g., dynamic solid memory), storage devices (e.g., solid-state or disk drives), network input / output (NW I / O) devices, network switches, virtual machines (VMs), power supply modules and / or cooling modules, etc.In some embodiments, one or more node CRs among node CRs 616(1)-616(N) may correspond to a server that has one or more of the aforementioned computing resources. Furthermore, in some embodiments, node CRs 616(1)-616(N) may contain one or more virtual components, such as vGPUs, vCPUs, and / or the like, and / or one or more of node CRs 616(1)-616(N) may correspond to a virtual machine (VM).
[0204] In at least one embodiment, the grouped compute resources 614 can contain separate groupings of node CRs 616, which are housed in one or more racks (not shown) or in many racks in data centers at different geographic locations (also not shown). Separate groupings of node CRs 616 within grouped compute resources 614 can contain grouped compute, network, memory, or storage resources that can be configured or allocated to support one or more workloads. In at least one embodiment, multiple node CRs 616, including the CPUs, GPUs, DPUs, and / or other processors, can be grouped in one or more racks to provide compute resources for supporting one or more workloads.The one or more racks can also contain any number of power supply modules, cooling modules and / or network switches in any combination.
[0205] The resource orchestrator 612 can configure or otherwise control one or more node CRs 616(1)-616(N) and / or grouped computing resources 614. In at least one embodiment, the resource orchestrator 612 can include a software design infrastructure (SDI) management entity for the data center 600. The resource orchestrator 612 can include hardware, software, or a combination thereof.
[0206] In at least one embodiment, as in Fig. As shown in Figure 6, the framework layer 620 can contain a job scheduler 633, a configuration manager 634, a resource manager 636, and / or a distributed file system 638. The framework layer 620 can contain a framework that supports the software 632 of software layer 630 and / or application(s) 642 of application layer 640. The software 632 or the application(s) 642 can each contain web-based service software or applications, such as those provided by Amazon Web Services, Google Cloud, and Microsoft Azure. The framework layer 620 can be a type of free and open-source software web application framework, such as Apache Spark™ (hereinafter "Spark"), which can utilize a distributed file system 638 for processing large amounts of data (e.g., "Big Data"), but is not limited to it.In at least one embodiment, the job scheduler 633 can include a Spark driver to facilitate the scheduling of workloads supported by different layers of the data center 600. The configuration manager 634 can be capable of configuring different layers, such as the software layer 630 and the framework layer 620, which includes Spark and the distributed file system 638, to support the processing of large amounts of data. The resource manager 636 can be capable of managing clustered or grouped computing resources allocated or assigned to support the distributed file system 638 and the job scheduler 633. In at least one embodiment, the clustered or grouped computing resources can include the grouped computing resource 614 on the infrastructure layer 610 of the data center.The Resource Manager 636 can coordinate with the Resource Orchestrator 612 to manage these allocated or assigned computing resources.
[0207] In at least one embodiment, the software contained in software layer 630 may include software 632 that is used by at least sections of the node CRs 616(1)-616(N), the grouped computing resources 614, and / or the distributed file system 638 of framework layer 620. One or more types of software may include, but are not limited to, web page search software, email virus scanning software, database software, and streaming video content software.
[0208] In at least one embodiment, the application(s) 642 contained in the application layer 640 may contain one or more types of applications used by at least sections of the nodes CRs 616(1)-616(N), the grouped compute resources 614, and / or the distributed file system 638 of the framework layer 620. One or more types of applications may include, but are not limited to, any number of genome applications, cognitive computations, and machine learning applications, including training or inference software, machine learning framework software (e.g., PyTorch, TensorFlow, Caffe, etc.), and / or other machine learning applications used in conjunction with one or more embodiments. The application(s) 642 of the application layer may execute and control systems of the vehicle 400 and may contain various ASILs.
[0209] In at least one embodiment, a configuration manager 634, resource manager 636, and resource orchestrator 612 can implement any number and type of self-modifying actions based on any set and type of data acquired in any technically feasible manner. Self-modifying actions can relieve a data center operator of data center 600 of potentially making poor configuration decisions and potentially avoiding underutilized and / or poorly functioning sections of a data center.
[0210] The Data Center 600 may contain tools, services, software, or other resources to train one or more machine learning models or to predict or infer information using one or more machine learning models according to one or more embodiments described herein. For example, a machine learning model(s) may be trained by calculating weighting parameters according to a neural network architecture, using software and / or computing resources described above with reference to the Data Center 600.In at least one embodiment, trained or deployed machine learning models corresponding to one or more neural networks can be used to infer or predict information using the resources described above with reference to the Computing Center 600 by using weighting parameters calculated by one or more training techniques such as, but not limited to, those described herein.
[0211] In at least one embodiment, the data center can use 600 CPUs, application-specific integrated circuits (ASICs), GPUs, FPGAs, and / or other hardware (or equivalent virtual computing resources) to perform training and / or inference using the resources described above. Furthermore, one or more of the software and / or hardware resources described above can be configured as a service to allow users to train or infer information, such as image capture, speech capture, or other artificial intelligence services. EXEMPLARY NETWORK ENVIRONMENTS
[0212] Network environments suitable for implementing embodiments of the disclosure may include one or more client devices, servers, network-attached storage (NAS), other backend devices, and / or other device types. The client devices, servers, and / or other device types (e.g., each device) may run on one or more instances of the computing device(s). Fig. 5. Implemented - e.g., each device may contain similar components, features, and / or functionality to the computing device(s) 500. If backend devices (e.g., servers, NAS, etc.) are implemented, the backend devices may also be included as part of a data center 600, an example of which is given herein with reference to Fig. 6 is described in more detail.
[0213] The components of a network environment can communicate with each other over a network, which can be wired, wireless, or both. The network can contain multiple networks or a network of networks. For example, the network can include one or more wide area networks (WANs), one or more local area networks (LANs), one or more public networks such as the internet and / or a public switched telephone network (PSTN), and / or one or more private networks. If the network includes a wireless telecommunications network, components such as a base station, a communications tower, or even access points (as well as other components) can provide wireless connectivity.
[0214] Compatible network environments can include one or more peer-to-peer network environments—in which case a server cannot be included in a network environment—and one or more client-server network environments—in which case one or more servers can be included in a network environment. In peer-to-peer network environments, the functionality described herein can be implemented on any number of client devices with reference to a server.
[0215] In at least one embodiment, a network environment can include one or more cloud-based network environments, a distributed computing environment, a combination thereof, etc. A cloud-based network environment can include a framework layer, a job scheduler, a resource manager, and a distributed file system implemented on one or more servers, which may include one or more core network servers and / or edge servers. A framework layer can include a framework for supporting software of a software layer and / or one or more applications of an application layer. The software or application(s) can each include web-based service software or applications. In embodiments, one or more of the client devices can use the web-based service software or applications (e.g.,by accessing the service software and / or applications via one or more application programming interfaces (APIs). The framework layer can be a type of free and open-source software web application framework, such as one that uses a distributed file system for processing large amounts of data (e.g., "Big Data"), but is not limited to that.
[0216] A cloud-based network environment can provide cloud computing and / or cloud storage, performing any combination (or parts thereof) of the computing and / or data storage functions described herein. Each of these different functions can be distributed across multiple locations of central or core servers (e.g., one or more data centers, which may be distributed across a state, region, country, the globe, etc.). If a connection to a user (e.g., a client device) is relatively close to an edge server, the core server(s) may offload at least some functionality to the edge server(s). A cloud-based network environment can be private (e.g., restricted to a single organization), public (e.g., available to many organizations), and / or a combination thereof (e.g., a hybrid cloud environment).
[0217] The client device(s) may include at least some of the components, features, and functions described herein with respect to Fig.The exemplary computing device(s) described in Section 5 contain 500. By way of example, and not as a limitation, a client device may be a personal computer (PC), a laptop, a mobile device, a smartphone, a tablet computer, a smartwatch, a portable computer, a personal digital assistant (PDA), an MP3 player, a virtual reality headset, a global positioning system (GPS) or global positioning device, a video player, a video camera, a surveillance device or surveillance system, a vehicle, a boat, a hydrofoil, a virtual machine, a drone, a robot, a handheld communication device, a hospital device, a gaming device or gaming system, an entertainment system, a vehicle computer system, an embedded system controller, a remote control, a device, a consumer electronics device, a workstation, an edge device,any combination of these described devices or any other suitable device may be embodied.
[0218] The disclosure of this application also contains the following numbered clauses: Clause 1. One or more processors comprising the following: one or more circuits for the following: Determine that a first task from a multitude of tasks meets a criterion for execution in a redundant mode; Determining a task flow for the execution of the multitude of tasks, in which a switch is assigned prior to the execution of the first task, wherein the switch causes the one or more circuits to be divided into a first partition and a second partition; and Executing the multitude of tasks according to the task flow by executing a first instance of the first task on the first partition and a second instance of the first task on the second partition. Clause 2. The one or more processors according to Clause 1, wherein the switch is a first switch and the one or more circuits serve the following purpose: Determine that a second task in the set of tasks is dependent on the first task and does not meet the criterion for execution in redundant mode; and Assigning the task flow between the execution of the first task and the execution of the second task of a second switch, wherein the second switch causes one or more circuits to be unpartitioned. Clause 3. The one or more processors according to any of the preceding clauses, wherein the one or more circuits are to determine that the first task satisfies the criterion, based on at least one feature of the first task received from at least one application containing the first task or user input relating to the first task. Clause 4. The one or more processors of any of the preceding clauses, wherein the one or more circuits shall define the task flow as a graph containing a plurality of nodes, wherein the switch between a first node for performing a non-redundant task and each of (i) a second node coupled to the first node, wherein the second node shall perform the first instance of the first task, and (ii) a third node coupled to the first node, wherein the third node shall perform the second instance of the first task. Clause 5. The one or more processors according to any of the preceding clauses, wherein the task flow specifies instructions for the one or more circuits to use a hardware tool for partitioning the one or more circuits without exposing the hardware tool to an application allocated to the multitude of tasks. Clause 6. The one or more processors according to any of the preceding clauses, wherein the switch is a first switch and the one or more circuits are to assign a second switch to the task flow to switch the context from the first task to the execution of a second, redundantly executed task. Clause 7. The one or more processors according to any of the preceding clauses, wherein the one or more circuits shall configure the first partition and the second partition as concurrent multiple contexts, graphics processing unit (GPU) partitions or multiple instances of a multi-instance GPU (MIG). Clause 8. The one or more processors according to any of the preceding clauses, wherein the one or more processors are included in at least one of the following: a control system for an autonomous or semi-autonomous machine; a perception system for an autonomous or semi-autonomous machine; a system that embodies one or more virtual machines (VMs); a system that is implemented using a robot; a system implemented using an edge device; a system for generating synthetic data; a system that includes one or more large language models (LLMs); a system that includes one or more large visual language models (VLMs); a system that includes one or more multimodal language models; a system that performs virtualization at the operating system (OS) level, which includes at least one of the following: one or more deep learning models, software for running the one or more deep learning models, or telemetry software for at least one of the evaluation, monitoring, or health checks of the system; a system that uses one or more microservices; a system for the use of one or more inference microservices; a system for performing conversational AI operations; a system for performing deep learning processes; a system for carrying out simulation processes; a system for conducting collaborative content creation for 3D assets; a system for performing digital twin operations; a system for performing light transport simulations; a system that is at least partially implemented in a data center; or a system that is implemented at least partially using cloud computing resources. Clause 9. System, which includes the following: one or more processors for performing operations that include the following: Determine that a first task from a multitude of tasks meets a criterion for execution in a redundant mode; Determining a task flow for the execution of the multitude of tasks, in which a switch is assigned prior to the execution of the first task, wherein the switch causes the one or more circuits to be divided into a first partition and a second partition; and Executing the multitude of tasks according to the task flow by executing a first instance of the first task on the first partition and a second instance of the first task on the second partition. Clause 10. System according to Clause 9, wherein the switch is a first switch, wherein the one or more processing processors are to perform operations which further include: Determine that a second task in the set of tasks is dependent on the first task and does not meet the criterion for execution in redundant mode; and Assigning the task flow between the execution of the first task and the execution of the second task of a second switch, wherein the second switch causes one or more circuits to be unpartitioned. Clause 11. System according to Clause 9 or Clause 10, wherein the one or more processors shall determine that the first task satisfies the criterion, based on at least one feature of the first task received from at least one application containing the first task or user input relating to the first task. Clause 12. System according to one of Clauses 9-11, wherein the one or more processors shall determine the task flow as a graph containing a plurality of nodes, wherein the switch between a first node for the execution of a non-redundant task and each of (i) a second node coupled to the first node, wherein the second node shall execute the first instance of the first task, and (ii) a third node coupled to the first node, wherein the third node shall execute the second instance of the first task. Clause 13. System according to any of Clauses 9-12, wherein the task flow specifies instructions for the one or more processing units to use a hardware tool for partitioning the one or more processing units without exposing the hardware tool to an application allocated to the multitude of tasks. Clause 14. System according to one of clauses 9-13, wherein the switch is a first switch and the one or more processors are to assign a second switch to the task flow to switch the context from the first task to the execution of a second, redundantly executed task. Clause 15. System according to any of Clauses 9-14, wherein the one or more processors shall configure the first partition and the second partition as concurrent multiple contexts, graphics processing unit (GPU) partitions or multiple instances of a multi-instance GPU (MIG). Clause 16. System according to any of Clauses 9-15, wherein the system is contained in at least one of the following: a control system for an autonomous or semi-autonomous machine; a perception system for an autonomous or semi-autonomous machine; a system that embodies one or more virtual machines (VMs); a system that is implemented using a robot; a system implemented using an edge device; a system for generating synthetic data; a system that includes one or more large language models (LLMs); a system that includes one or more large visual language models (VLMs); a system that includes one or more multimodal language models; a system for performing virtualization at the operating system (OS) level, which includes at least one of the following: one or more deep learning models, software for running the one or more deep learning models, or telemetry software for at least one of the evaluation, monitoring, or health checks of the system; a system that uses one or more microservices; a system for the use of one or more inference microservices; a system for performing conversational AI operations; a system for performing deep learning processes; a system for carrying out simulation processes; a system for conducting collaborative content creation for 3D assets; a system for performing digital twin operations; a system for performing light transport simulations; a system that is at least partially implemented in a data center; or a system that is implemented at least partially using cloud computing resources. Clause 17. Procedure, which includes the following: Determine that a first task from a multitude of tasks meets a criterion for execution in a redundant mode; Determining a task flow for the execution of the multitude of tasks, in which a switch is assigned prior to the execution of the first task, wherein the switch causes the one or more circuits to be divided into a first partition and a second partition; and Executing the multitude of tasks according to the task flow by executing a first instance of the first task on the first partition and a second instance of the first task on the second partition. Clause 18. The procedure according to Clause 17, wherein the switch is a first switch, further comprising the following: Determine that a second task in the set of tasks is dependent on the first task and does not meet the criterion for execution in redundant mode; and Assigning the task flow between the execution of the first task and the execution of the second task of a second switch, wherein the second switch causes one or more circuits to be unpartitioned. Clause 19. Procedure according to Clause 17 or Clause 18, wherein determining the task flow as a graph contains a plurality of nodes, wherein the switch is between a first node for the execution of a non-redundant task and each of (i) a second node coupled to the first node, wherein the second node is to execute the first instance of the first task, and (ii) a third node coupled to the first node, wherein the third node is to execute the second instance of the first task. Clause 20. Method according to any of Clauses 17-19, wherein the first partition and the second partition are configured as concurrent multiple contexts, graphics processing unit (GPU) partitions or multiple instances of a multi-instance GPU (MIG).
[0219] The revelation can be described in the general context of computer code or machine-usable instructions, including computer-executable instructions such as program modules that are executed by a computer or other machine, such as a personal data assistant or other handheld device. Generally, program modules contain routines, programs, objects, components, data structures, etc., and refer to code that performs specific tasks or implements certain abstract types of data. The revelation can be practiced in a variety of system configurations, including handheld devices, consumer electronics, general-purpose computers, more specialized computing devices, etc. The revelation can also be practiced in distributed computing environments, where tasks are performed by remote processing devices connected to each other via a network for communication.
[0220] As used herein, any mention of "and / or" in relation to two or more elements should be interpreted as referring to only one element or a combination of elements. For example, "element A, element B and / or element C" may include only element A, only element B, only element C, element A and element B, element A and element C, element B and element C, or elements A, B, and C. Furthermore, "at least one of element A or element B" may include at least one of element A, at least one of element B, or at least one of element A and at least one of element B. Additionally, "at least one of element A and element B" may include at least one of element A, at least one of element B, or at least one of element A and at least one of element B.
[0221] The subject matter of this disclosure is specifically described herein to satisfy legal requirements. However, the description itself is not intended to limit the scope of protection afforded by this disclosure. Rather, the inventors have considered that the claimed subject matter may also be embodied in other ways to include various steps or combinations of steps similar to those described in this document, in conjunction with other present or future technologies. Although the terms "step" and / or "block" may be used herein to denote various elements of the methods employed, these terms should not be interpreted as implying any particular sequence among or between the various steps disclosed herein, except where the sequence of each step is expressly described.
[0222] It is understood that the aspects and embodiments described above are only exemplary and that changes to details may be made within the scope of protection of the claims.
[0223] Each device, method and feature disclosed in the description and (where applicable) in the claims and drawings may be provided independently or in any suitable combination.
[0224] The reference numerals appearing in the claims are for illustrative purposes only and are not intended to have any limiting effect on the scope of protection of the claims. QUOTES INCLUDED IN THE DESCRIPTION
[0000] This list of documents cited by the applicant was automatically generated and is included solely for the reader's convenience. The list is not part of the German patent or utility model application. The DPMA accepts no liability for any errors or omissions. Cited patent literature
[0000] US 16 / 101,232
[0123] Cited non-patent literature
[0000] Society of Automotive Engineers, SAE) (Standard No. J3016-201806, published on June 15, 2018, Standard No. J3016-201609, published on September 30, 2016
[0080]
Claims
[1] One or more processors comprising the following: one or more circuits for the following: Determine that a first task from a multitude of tasks meets a criterion for execution in a redundant mode; Determining a task flow for the execution of the multitude of tasks, in which a switch is assigned prior to the execution of the first task, wherein the switch causes the one or more circuits to be divided into a first partition and a second partition; and Executing the multitude of tasks according to the task flow by executing a first instance of the first task on the first partition and a second instance of the first task on the second partition. [2] The one or more processors according to claim 1, wherein the switch is a first switch and the one or more circuits serve to: Determine that a second task in the set of tasks is dependent on the first task and does not meet the criterion for execution in redundant mode; and Assigning the task flow between the execution of the first task and the execution of the second task of a second switch, wherein the second switch causes one or more circuits to be unpartitioned. [3] The one or more processors according to claim 1 or 2, wherein the one or more circuits are to determine that the first task satisfies the criterion, on the basis of at least one feature of the first task that was received from at least one application containing the first task or user input relating to the first task. [4] The one or more processors according to any of the preceding claims, wherein the one or more circuits are to define the task flow as a graph containing a plurality of nodes, wherein the switch between a first node for the execution of a non-redundant task and each of (i) a second node coupled to the first node, wherein the second node is to execute the first instance of the first task, and (ii) a third node coupled to the first node, wherein the third node is to execute the second instance of the first task. [5] The one or more processors according to any of the preceding claims, wherein the task flow specifies instructions for the one or more circuits to use a hardware tool for partitioning the one or more circuits without exposing the hardware tool to an application that is assigned to the plurality of tasks. [6] The one or more processors according to any of the preceding claims, wherein the switch is a first switch and the one or more circuits are to assign a second switch to the task flow in order to switch the context from the first task to the execution of a second, redundantly executed task. [7] The one or more processors according to any of the preceding claims, wherein the one or more circuits are to configure the first partition and the second partition as concurrent multiple contexts, graphics processing unit (GPU) partitions or multiple instances of a multi-instance GPU (MIG). [8] The one or more processors according to any of the preceding claims, wherein the one or more processors are included in at least one of the following: a control system for an autonomous or semi-autonomous machine; a perception system for an autonomous or semi-autonomous machine; a system that embodies one or more virtual machines (VMs); a system that is implemented using a robot; a system implemented using an edge device; a system for generating synthetic data; a system that includes one or more large language models (LLMs); a system that includes one or more large Visual Language Models (VLMs); a system that includes one or more multimodal language models; a system that performs virtualization at the operating system (OS) level, which includes at least one of the following: one or more deep learning models, software for running the one or more deep learning models, or telemetry software for at least one of the evaluation, monitoring, or health checks of the system; a system that uses one or more microservices; a system for the use of one or more inference microservices; a system for performing conversational AI operations; a system for performing deep learning processes; a system for carrying out simulation processes; a system for conducting collaborative content creation for 3D assets; a system for performing digital twin operations; a system for performing light transport simulations; a system that is at least partially implemented in a data center; or a system that is implemented at least partially using cloud computing resources. [9] System comprising the following: one or more processors for performing operations that include the following: Determine that a first task from a multitude of tasks meets a criterion for execution in a redundant mode; Determining a task flow for the execution of the multitude of tasks, in which a switch is assigned prior to the execution of the first task, wherein the switch causes the one or more circuits to be divided into a first partition and a second partition; and Executing the multitude of tasks according to the task flow by executing a first instance of the first task on the first partition and a second instance of the first task on the second partition. [10] System according to claim 9, wherein the switch is a first switch, wherein the one or more processing processors are to perform operations which further comprise: Determine that a second task in the set of tasks is dependent on the first task and does not meet the criterion for execution in redundant mode; and Assigning the task flow between the execution of the first task and the execution of the second task of a second switch, wherein the second switch causes one or more circuits to be unpartitioned. [11] System according to claim 9 or claim 10, wherein the one or more processors are to determine that the first task satisfies the criterion, on the basis of at least one feature of the first task that was received from at least one application containing the first task or user input relating to the first task. [12] System according to one of claims 9-11, wherein the one or more processors are to determine the task flow as a graph containing a plurality of nodes, wherein the switch between a first node for the execution of a non-redundant task and each of (i) a second node coupled to the first node, wherein the second node is to execute the first instance of the first task, and (ii) a third node coupled to the first node, wherein the third node is to execute the second instance of the first task. [13] System according to one of claims 9-12, wherein the task flow specifies instructions for the one or more processing units to use a hardware tool for partitioning the one or more processing units without exposing the hardware tool to an application that is assigned to the plurality of tasks. [14] System according to one of claims 9-13, wherein the switch is a first switch and the one or more processors are to assign a second switch to the task flow in order to switch the context from the first task to the execution of a second, redundantly executed task. [15] System according to one of claims 9-14, wherein the one or more processors are to configure the first partition and the second partition as concurrent multiple contexts, graphics processing unit (GPU) partitions or multiple instances of a multi-instance GPU (MIG). [16] System according to any one of claims 9-15, wherein the system is contained in at least one of the following: a control system for an autonomous or semi-autonomous machine; a perception system for an autonomous or semi-autonomous machine; a system that embodies one or more virtual machines (VMs); a system that is implemented using a robot; a system implemented using an edge device; a system for generating synthetic data; a system that includes one or more large language models (LLMs); a system that includes one or more large Visual Language Models (VLMs); a system that includes one or more multimodal language models; a system for performing virtualization at the operating system (OS) level, which includes at least one of the following: one or more deep learning models, software for running the one or more deep learning models, or telemetry software for at least one of the evaluation, monitoring, or health checks of the system; a system that uses one or more microservices; a system for the use of one or more inference microservices; a system for performing conversational AI operations; a system for performing deep learning processes; a system for carrying out simulation processes; a system for conducting collaborative content creation for 3D assets; a system for performing digital twin operations; a system for performing light transport simulations; a system that is at least partially implemented in a data center; or a system that is implemented at least partially using cloud computing resources. [17] Method comprising the following: Determine that a first task from a multitude of tasks meets a criterion for execution in a redundant mode; Determining a task flow for the execution of the multitude of tasks, in which a switch is assigned prior to the execution of the first task, wherein the switch causes the one or more circuits to be divided into a first partition and a second partition; and Executing the multitude of tasks according to the task flow by executing a first instance of the first task on the first partition and a second instance of the first task on the second partition. [18] The method of claim 17, wherein the switch is a first switch, further comprising: Determine that a second task in the set of tasks is dependent on the first task and does not meet the criterion for execution in redundant mode; and Assigning the task flow between the execution of the first task and the execution of the second task of a second switch, wherein the second switch causes one or more circuits to be unpartitioned. [19] Method according to claim 17 or claim 18, wherein determining the task flow as a graph includes a plurality of nodes, wherein the switch is between a first node for performing a non-redundant task and each of (i) a second node coupled to the first node, wherein the second node is to perform the first instance of the first task, and (ii) a third node coupled to the first node, wherein the third node is to perform the second instance of the first task. [20] Method according to one of claims 17-19, wherein the first partition and the second partition are configured as concurrent multiple contexts, graphics processing unit (GPU) partitions or multiple instances of a multi-instance GPU (MIG).
Citation Information
Patent Citations
US-PATENTANMELDUNGNR.16/101,232