Adapting performance levels of runnable programs based on security ratings
By using UCIe specifications and functional safety procedures, and based on safety rating management of chip interconnects and system degradation in autonomous vehicles, the problem of insufficient safety and reliability in autonomous vehicles is solved, and safe degradation and stable operation are achieved in complex environments.
Patent Information
- Application Number
- CN202480050464.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-08-02
- Filing Date
- 2024-07-25
- Publication Date
- 2026-03-06
AI Technical Summary
Existing technologies have failed to effectively manage the interconnection between components from different silicon wafer manufacturers in autonomous vehicles, and lack effective scheduling of security ratings and system degradation for computing systems, resulting in insufficient safety and reliability in complex driving environments.
By monitoring the execution of workload processing cores through core layout and functional safety procedures in accordance with UCIe specifications, and degrading the system based on safety ratings (such as ASIL), including reducing execution frequency, changing connections or trunculating non-critical programs, and managing workload scheduling using reservation tables, the system can be safely degraded under conditions such as overload or high temperature.
It realizes a dynamic scheduling program based on safety rating in autonomous vehicles, which improves the safety and reliability of the system in complex environments and ensures that vehicles can be safely stopped or parked in case of failure.
Smart Images

Figure CN121620751A_ABST
Abstract
Description
Background Technology
[0001] Universal Chip Interconnect (UCIe) provides an open specification for interconnects and serial buses between chips, enabling the production of large system-on-a-chip (SoC) packages that mix components from different silicon manufacturers. It is envisioned that autonomous vehicle computing systems could operate using chip arrangements conforming to the UCIe specification. One goal in manufacturing such computing systems is to achieve robust safety integrity levels for other critical electrical and electronic (E / E) automotive components of the vehicle. Summary of the Invention
[0002] This paper describes systems and methods for scheduling a set of runnables included in a software architecture for execution by a set of workload processing chips (chiplets). Depending on the scheduler, each runnable in the software architecture and / or each connection between runnables can be associated with a safety rating (e.g., an Automotive Safety Integrity Level (ASIL) rating) to facilitate degradation of the software architecture. In various examples, a set of workload processing chips can be included on a system-on-a-chip (SoC) comprising a central chip and a reservation table containing a scheduler and workload information (e.g., dependency information on when a workload can be executed as a runnable on a workload processing chiplet).
[0003] In some implementations, the central chip may include a Functional Safety (FuSa) procedure that: (i) monitors communications corresponding to the execution of operable programs by a set of workload processing chips, and (ii) triggers a degradation of the software architecture upon detecting data corresponding to system overload, processing latency, overheating, or other issues affecting the SoC that require system degradation. In response to the FuSa procedure triggering degradation, the scheduler performs the software architecture degradation based on the security rating of each operable program in the software architecture and / or each connection between operable programs. In some examples, the scheduler's degradation of the software architecture may include reducing the execution frequency of a selected set of operable programs in the software architecture. In other examples, the scheduler may use a reservation table to reduce the execution frequency of a selected set of operable programs, the reservation table identifying when said selected set of operable programs is ready to execute (e.g., based on dependency information being met). Additionally or alternatively, degradation may include deforming (transforming) and / or truncating the software architecture to change the connections between operable programs or exclude the execution of certain operable programs (e.g., operable programs with low security ratings). In another example, degradation can include layered switching between computation graphs or software structures (e.g., using a first-level degradation computation graph, a second-level degradation computation graph, a third-level degradation computation graph, etc.).
[0004] Depending on the implementation, adaptive performance of the runnable program in the SoC can be achieved for autonomous vehicle operation. For example, the SoC may include sensor data input chips, a set of workload processing chips, and a central chip, and may be included on a SoC arrangement in UCIe. The SoC arrangement may also include one or more machine learning (ML) accelerator chips, high-bandwidth memory (HBM) chips, and / or autonomous driver chips. Attached Figure Description
[0005] The disclosure herein is illustrated by way of example rather than limitation in the accompanying drawings, and similar reference numerals in the drawings refer to similar elements, wherein:
[0006] Figure 1 It is a block diagram depicting an exemplary computing system according to the examples described herein, in which the embodiments described herein can be implemented;
[0007] Figure 2 It is a block diagram depicting a system-on-chip (SoC) that can implement the examples described herein, according to the examples described herein;
[0008] Figure 3 This is a block diagram illustrating an exemplary central chip for executing an operable program in accordance with the examples described herein;
[0009] Figure 4 The software architecture, consisting of a set of operable programs to be executed by the workload processing core, is described according to the example described herein;
[0010] Figure 5 It is a block diagram depicting an exemplary multi-chip system-on-a-chip (mSoC) based on the examples described herein;
[0011] Figure 6 This is a block diagram depicting a performance network and FuSa network, based on the examples described herein, for performing health monitoring, error correction, and system degradation; and
[0012] Figure 7 and Figure 8 This is a flowchart describing exemplary methods for adapting runnable programs based on various examples and security ratings. Detailed Implementation
[0013] In experimental and controlled testing environments, system redundancy and Automotive Safety Integrity Level (ASIL) ratings are generally not priority considerations for autonomous systems. As autonomous driving features continue to evolve (e.g., beyond Level 3 autonomy) and autonomous vehicles begin to operate more commonly on public road networks, the qualification and certification of E / E components related to the autonomous operation of vehicles will contribute to ensuring the operational safety of these vehicles. Furthermore, novel methods for the qualification and certification of hardware, software, and / or hardware / software combinations will also contribute to increasing public confidence and ensuring that the safety of autonomous driving systems exceeds current standards. For example, some safety standards for autonomous driving systems include safety thresholds corresponding to average human ability and level of caution. However, these statistics include the incidence of vehicles with driver inattention or distraction but do not account for specific time windows where vehicle operation is inherently riskier (e.g., adverse weather conditions, late-night driving, winding mountain roads, etc.).
[0014] The Automotive Safety Integrity Level (ASIL) is a risk classification scheme defined by ISO 26262 (Standard for Functional Safety of Road Vehicles). It is typically established for the vehicle's electrical and electronic components (E / E) through a risk analysis of potential hazards. This risk analysis involves determining the severity of the vehicle's operating scenarios (i.e., the severity of damage that a hazard is expected to cause; classified between S0 (no damage) and S3 (life-threatening damage)), the level of exposure (i.e., the relative expected frequency of operating conditions where damage may occur; classified between E0 (extremely unlikely) and E4 (high probability of damage under most operating conditions)), and the level of controllability (i.e., the relative likelihood that the driver can take action to prevent damage; classified between C0 (generally controllable) and C3 (difficult to control or uncontrollable)). Therefore, safety objectives for any potential hazardous event include a set of ASIL requirements.
[0015] Hazards identified as Quality Management (QM) do not determine any safety requirements. As an example, these QM hazards can be any combination of a low probability of exposure to the hazard, a low severity level of potential damage caused by the hazard, and a high level of controllability for the driver in avoiding the hazard and / or preventing damage. Other hazardous events are classified as ASIL-A, ASIL-B, ASIL-C, or ASIL-D based on the severity, exposure level, and controllability corresponding to the various levels of potential hazard. ASIL-D events correspond to the highest integrity requirements (ASIL requirements) for safety systems or E / E components of safety systems, while ASIL-A includes the lowest integrity requirements. For example, vehicle airbags, anti-lock braking systems (ABS), and power steering systems would typically have an ASIL-D level, where the risks associated with failure of these components (e.g., the potential severity of damage and the lack of vehicle controllability to prevent such damage) are relatively high.
[0016] As provided herein, ASIL can refer to both risk requirements and risk-dependent requirements, where various combinations of severity, exposure, and controllability are quantified to form an expression of risk (e.g., an airbag system for a vehicle may have a relatively low exposure classification but high values for severity and controllability). As provided above, the quantities of severity, exposure, and controllability for a given hazard are traditionally determined using values for severity (e.g., S0 to S3), exposure (e.g., E0 to E4), and controllability (e.g., C0 to C3) from the ISO 26262 series, where these values are then used to classify the ASIL requirements for components of a particular safety system. As provided herein, certain safety systems can perform variable mitigation measures, which may include alarms (e.g., visual, audible, or tactile alarms), minor interventions (e.g., brake assist or steering assist), major interventions and / or evasive maneuvers (e.g., taking over control of one or more control mechanisms, such as steering, acceleration, or braking systems), and full autonomous control of the vehicle.
[0017] Based on the examples described herein, a software architecture for performing autonomous or semi-autonomous vehicle functions (e.g., perception, object detection and classification, scene understanding ML inference, etc.) may include a set of runnable programs that may be connected to other runnable programs based on associations or input / output dependencies. As an example, a first runnable program may be responsible for identifying and classifying pedestrians in image data, and a second runnable program may be responsible for predicting the motion of each pedestrian. In this example, the second runnable program will receive the output of the first runnable program as input. Therefore, in the software architecture, the first and second runnable programs are connected.
[0018] In the various examples described herein, operable programs and / or connections between operable programs within a software architecture can be associated with safety ratings, such as QM ratings, or ASIL-A, ASIL-B, ASIL-C, or ASIL-D ratings. These safety ratings can be defined within the software architecture and can indicate the criticality of an operable program or a connection between two operable programs. As an example, an ASIL-D rated operable program or a connection between two operable programs can correspond to the identification of traffic signals and the classification of traffic signal states (e.g., red, yellow, or green lights), which is crucial for preventing car collisions. As another example, a connection between two operable programs rated ASIL-B can correspond to the detection, classification, and speed difference calculation of vehicles behind autonomous vehicles in the same lane (e.g., because these vehicles do not have priority over the autonomous vehicle's right-of-way).
[0019] It is conceivable that defining or associating executable programs or connections between executable programs within a software architecture can facilitate system degradation (e.g., when a computing system experiences a critical temperature due to overload computing and / or high ambient temperatures). Furthermore, the degradation of the execution of the software architecture can be managed by the functional safety (FuSa) component of the computing system, which can be responsible for: (i) monitoring communication between executable programs (e.g., the cores that organize data and execute executable programs based on that data), (ii) communicating with the thermal management component of the computing system, and (iii) communicating with a workload scheduler that manages a reservation table including workload entries indicating whether a workload is available for execution in one or more executable programs.
[0020] As presented herein, "degradation" of software architecture, or the degradation of the execution of software architecture, generally refers to the selective reduction of certain computational tasks based on a safety rating associated with the operable program and / or the connections between operable programs. As an example, ISO 26262 refers to the concept of degradation in the context of automotive safety, where the functionality of a given system (e.g., an E / E system) is degraded to achieve a safe state. Therefore, autonomous systems of vehicles should be able to handle degraded functionality appropriately. Alternatively, as a last resort, the autonomous driving system can safely pull the vehicle over or park it.
[0021] In various specific implementations, the degradation of software architecture or the degradation of the execution of software architecture can involve a scheduler that reduces the execution frequency of a selected set of runnable programs within the software architecture. In another example, the scheduler may use a reservation table to reduce the execution frequency of a selected set of runnable programs, the reservation table indicating when the selected set of runnable programs is ready to execute (e.g., based on dependency information being satisfied). In yet another example, degradation may include: reducing the processing frequency of a selected device (specific chip) on which runnable programs run, runnable programs on the selected device or runnable program connections having a lower security rating (e.g., ignoring every other image or sensor data iteration), reducing or increasing the frequency at which runnable programs are executed, ignoring data from certain non-critical sensors, temporarily blocking the execution of certain runnable programs, and so on.
[0022] As described herein, degradation can involve altering or editing connections between operable programs and / or truncating selected portions of the computation graph corresponding to the software architecture to exclude the execution of one or more operable programs. In additional examples, degradation can involve complete changes between the software architecture or computation graphs. For example, a vehicle may store multiple computation graphs corresponding to different SAE autonomy levels and / or different degradation levels.
[0023] As further described herein, degradation can be implemented by the scheduler using a reservation table that indicates when certain workloads are available for execution as runnable programs. For example, the scheduler can flush certain workloads in the unordered buffer of the reservation table without executing them (e.g., data corresponding to a workload can be moved to an HBM kernel instead of being processed by the workload processing kernel), or the reservation table can selectively remove certain workloads associated with runnable programs that have low security ratings.
[0024] As provided herein, a “degradation event” can include any event capable of triggering a degradation of the software architecture or the execution of the software architecture, and can include excessive heat in the computing system or parts thereof, computer hardware failure or malfunction, sensor failure or malfunction, specified driving scenarios (e.g., extreme conditions in traffic for vulnerable road users), vehicle failure or malfunction, etc. For example, when other thermal mitigation measures (such as activating fans and / or water cooling systems and switching the SoC between primary and backup roles) are insufficient to successfully cool the computing system, the thermal management and functional safety components of the computing system can trigger a degradation of the software architecture.
[0025] In some specific implementations, the example computing system may implement one or more of the functions described herein using learning-based methods, such as by executing artificial neural networks (e.g., recurrent neural networks, convolutional neural networks, etc.) or one or more machine learning models. Such learning-based methods may also correspond to computing system storage or include one or more machine learning models. In implementations, the machine learning model may include an unsupervised learning model. In implementations, the machine learning-derived model may include neural networks (e.g., deep neural networks) or other types of machine learning-derived models, including nonlinear and / or linear models. Neural networks may include feedforward neural networks, recurrent neural networks (e.g., long short-term memory recurrent neural networks), convolutional neural networks, or other forms of neural networks. Some example machine learning-derived models may utilize attention mechanisms such as self-attention. For example, some example machine learning-derived models may include multi-head self-attention models (e.g., transformer models).
[0026] As provided herein, "network" or "one or more networks" can include any type of network or combination of networks that allows communication between devices. In embodiments, a network can include one or more of a local area network (LAN), a wide area network (WAN), the Internet, a secure network, a cellular network, a mesh network, a peer-to-peer communication link, or some combination thereof, and can include any number of wired or wireless links. For example, communication over a network can be achieved via a network interface using any type of protocol, protection scheme, encoding, format, packetization, etc.
[0027] In one or more examples described herein, methods, techniques, and actions performed by a computing device are executed in a programmatic manner, or as methods implemented by a computer. As used herein, "programmatically" means using code or computer-executable instructions. These instructions may be stored in one or more memory resources of the computing device. The steps executed in a programmatic manner may or may not be automatic.
[0028] One or more examples described herein may be implemented using a programming module, engine, or component. A programming module, engine, or component may include a program, subroutine, part of a program, or a software or hardware component capable of performing one or more of the described tasks or functions. As used herein, a module or component may exist independently of other modules or components on a hardware component. Alternatively, a module or component may be a shared element or process of other modules, programs, or machines.
[0029] Some of the examples described herein may typically require the use of computing devices, including processing and memory resources. For example, one or more examples described herein may be implemented wholly or partially on computing devices such as servers and / or on personal computers using networking equipment (e.g., routers). Memory resources, processing resources, and network resources may all be used in conjunction with the creation, use, or execution of any of the examples described herein, including with the execution of any method or with the implementation of any system.
[0030] Furthermore, one or more examples described herein can be implemented using instructions executable by one or more processors. These instructions may be carried on a non-transitory computer-readable medium. Machines illustrated or described with reference to the accompanying drawings provide examples of processing resources and computer-readable media on which instructions for implementing the examples disclosed herein may be carried and / or executed. In particular, many machines illustrated with reference to the examples of the invention include processors and various forms of memory for storing data and instructions. Examples of non-transitory computer-readable media include permanent memory storage devices, such as hard disk drives on personal computers or servers. Other examples of computer storage media include portable storage units, such as flash memory or magnetic storage. Computers, terminal devices, and network-enabled devices are examples of machines and devices utilizing processors, memory, and instructions stored on computer-readable media. Additionally, the examples may be implemented in the form of a computer program or a computer-usable carrier medium capable of carrying such a program.
[0031] Exemplary computing system
[0032] Figure 1 This is a block diagram depicting an example computing system 100, which implements the embodiments described herein, according to the examples described herein. In embodiments, the computing system 100 may include one or more control circuitry 110, which may include one or more processors (e.g., microprocessors), one or more processing cores, programmable logic circuitry (PLC) or programmable logic array (PLA) / programmable gate array (PGA), field-programmable gate array (FPGA), application-specific integrated circuit (ASIC), system-on-a-chip (SoC), or any other control circuitry. In some specific embodiments, the control circuitry 110 and / or the computing system 100 may be part of, or may form part of, a vehicle control unit (also referred to as a vehicle controller), which is embedded in or otherwise disposed in a vehicle (e.g., a Mercedes-Benz). ®In a car, truck, or van. For example, the vehicle controller may be or may include an infotainment system controller (e.g., infotainment head unit), telematics control unit (TCU), electronic control unit (ECU), central powertrain controller (CPC), central external and internal controller (CEIC), zone controller, autonomous vehicle control system, or any other controller (the term "or" may be used interchangeably with "and / or" herein).
[0033] In an implementation, control circuitry 110 may be programmed by one or more computer-readable or computer-executable instructions stored on a non-transitory computer-readable medium 120. The non-transitory computer-readable medium 120 may be a memory device (also referred to as a data storage device), which may include electronic storage devices, magnetic storage devices, optical storage devices, electromagnetic storage devices, semiconductor storage devices, or any suitable combination thereof. The non-transitory computer-readable medium 120 may be formed, for example, a computer floppy disk, hard disk drive (HDD), solid-state drive (SDD) or solid-state integrated memory, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), dynamic random access memory (DRAM), portable optical disc read-only memory (CD-ROM), digital versatile optical disc (DVD), and / or memory stick. In some cases, the non-transitory computer-readable medium 120 may store computer-executable or computer-readable instructions, such as those for performing the following combined Figure 7 and Figure 8 The instructions for the described method.
[0034] In various implementations, the terms "computer-readable instructions" and "computer-executable instructions" are used to describe software instructions or computer code configured to perform various tasks and operations. In various implementations, if the computer-readable instructions or computer-executable instructions form a module, the term "module" broadly refers to a collection of software instructions or code configured to cause control circuitry 110 to perform one or more functional tasks. When control circuitry 110 or other hardware components execute modules or computer-readable instructions, these modules and computer-readable / executable instructions can be described as performing various operations or tasks.
[0035] In another embodiment, computing system 100 may include a communication interface 140 that enables communication over one or more networks 150 to send and receive data. In various examples, computing system 100 may use communication interface 140 to communicate with fleet vehicles over one or more networks 150 to receive sensor data and implement the methods described throughout this disclosure. In some embodiments, communication interface 140 may be used to communicate with one or more other systems. Communication interface 140 may include any circuitry, components, software, etc., for communication via one or more networks 150 (e.g., local area networks, wide area networks, the Internet, secure networks, cellular networks, mesh networks, and / or peer-to-peer links). In some specific embodiments, communication interface 140 may include one or more of, for example, a communication controller, receiver, transceiver, transmitter, port, conductor, software and / or hardware for conveying data / information.
[0036] As an example implementation, the control circuitry 110 of the computing system 100 may include a System-on-a-Chip (SoC) arrangement that facilitates the various methods and techniques described throughout this disclosure. In various examples, the SoC may include a set of chips, including a central chip, which includes shared memory where various autonomous driving workloads are executed as runnable programs in independent pipelines using reserved tables. According to the implementation described herein, the shared memory of the central chip may include a FuSa program that can be executed by the control circuitry 110 to perform functional safety tasks of the SoC arrangement, as described in detail below.
[0037] Exemplary System-on-Chip
[0038] Figure 2 This is a block diagram illustrating an exemplary system-on-chip (SoC) 200 according to the examples described herein. Figure 2 The example SoC 200 shown may include additional components, and the components of the SoC 200 can be arranged in various alternative configurations different from the example shown. Therefore, Figure 2 The SoC 200 is described herein as an example arrangement for illustrative purposes and is not intended to limit the scope of this disclosure in any way.
[0039] refer to Figure 2The sensor data input chip 210 of the SoC 200 can receive sensor data from various vehicle sensors 205 of the vehicle. These vehicle sensors 205 can include any combination of image sensors (e.g., single-camera, binocular, fisheye lens, etc.), LiDAR sensors, radar sensors, ultrasonic sensors, proximity sensors, etc. The sensor data input chip 210 can automatically transfer the received sensor data as is to the cache memory 231 of the central chip 220. The sensor data input chip 210 may also include an image signal processor (ISP), which is responsible for capturing, processing, and enhancing images acquired from the various vehicle sensors 205. The ISP acquires raw image data and performs a series of complex image processing operations, such as color, contrast and brightness correction, noise reduction, and image enhancement, to create higher quality images ready for further processing or analysis by other chips of the SoC 200. The ISP may also include features such as autofocus, image stabilization, and advanced scene recognition to further enhance the quality of the captured images. The ISP can then store the higher quality images in the cache memory 231.
[0040] In some aspects, sensor data input core 210 publishes identification information for each piece of sensor data (e.g., image, point cloud, etc.) to a shared memory 230 of central core 220, which acts as a central mailbox for synchronizing the workloads of various cores. The identification information may include details such as the address of the stored data in cache memory 231, the type of sensor data, the sensor that captured the data, and the timestamp when the data was captured.
[0041] To communicate with the central chip 220, the sensor data input chip 210 transmits data via interconnect 211a. Interconnects 211a-f each represent a die-to-die (D2D) interface between chips of the SoC 200. In some aspects, interconnects 211a-f may include a high-bandwidth data path to cache memory 231 for general data purposes, and a high-reliability data path for sending functional safety (FuSa) and scheduler information to shared memory 230. Depending on bandwidth requirements, interconnects 211a-f may include more than one die-to-die interface. For example, interconnect 211a may include two interfaces to support higher bandwidth communication between the sensor data input chip 210 and the central chip 220.
[0042] In one aspect, interconnect 211a-f implements the Universal Chip Interconnect High Speed (UCIe) standard and communicates via an indirect mode to allow each chip host processor in the chip host processor to access remote memory as if it were local memory. This is achieved using a dedicated Network on-Chip (NoC) Network Interface Unit (NIU) (e.g., which allows devices connected to the network to communicate without interference), which provides hardware-level support for Remote Direct Memory Access (RDMA) operations. In UCIe indirect mode, the host processor sends a request to the NIU, which then accesses the remote memory and returns the data to the host processor. This approach allows for efficient and low-latency access to remote memory, which can be particularly useful in distributed computing and data-intensive applications. Additionally, UCIe indirect mode offers high flexibility because it can be used with a wide range of different network topologies and protocols.
[0043] In various examples, SoC 200 may include additional chips that can store, modify, or otherwise process sensor data cached by sensor data input chip 210. SoC 200 may include an autonomous driving chip 240 that can perform perception, sensor fusion, trajectory prediction, and / or other autonomous driving algorithms for autonomous vehicles. Autonomous driving chip 240 may connect to a dedicated HBM-RAM chip 235, in which it can publish all conditional information, variables, statistics, and / or processed sensor data as processed by autonomous driving chip 240.
[0044] In various examples, the system-on-chip 200 may also include a machine learning (ML) accelerator chip 240 specifically designed to accelerate machine learning workloads or AI workloads, such as image inference or other sensor inference using machine learning, to achieve high performance and low power consumption for these workloads. The ML accelerator chip 240 may include an engine and a highly parallel processor designed to efficiently process graph-based data structures (typically used in AI workloads), allowing for efficient processing of large amounts of data. The ML accelerator chip 240 may also include dedicated hardware accelerators for common AI operations such as matrix multiplication and convolution, and a memory hierarchy designed to optimize memory access for AI workloads that typically have complex memory access patterns.
[0045] The general-purpose computing chip 245 can provide general-purpose computing for the system-on-chip 200. For example, the general-purpose computing chip 245 may include a high-performance central processing unit and / or graphics processing unit capable of supporting computing tasks of the central chip 220, the autonomous driving chip 240, and / or the ML accelerator chip 250.
[0046] In various specific implementations, shared memory 230 may store programs and instructions for performing autonomous driving tasks. The shared memory 230 of the central core 220 may also include a reservation table that provides various cores with the information needed to perform their respective tasks (e.g., sensor data items and their locations in memory). In various aspects, the central core 220 also includes a large cache memory 231 that supports invalidation and refresh operations on the stored data. See below for further details. Figure 3 Further description of the shared memory 230 is provided in the context of the central core 220.
[0047] Cache misses and evicts from cache memory 231 are transferred by high-bandwidth memory (HBM) RAM core 255 connected to central core 220. HBM-RAM core 255 may include status information, variables, statistics, and / or sensor data from all other cores. In some examples, information stored in HBM-RAM core 255 may be stored for a predetermined period of time (e.g., ten seconds) before being deleted or otherwise refreshed. For example, when a malfunction occurs on an autonomous vehicle, the information stored in HBM-RAM core 255 may include all the information necessary to diagnose and resolve the malfunction. Compared to accessing data from HBM-RAM core 255, cache memory 231 keeps up-to-date data available with lower latency and less power consumption.
[0048] As provided herein, shared memory 230 can accommodate a mailbox architecture, including a set of reflective instructions for executing workloads via central chip 220, general-purpose computing chip 245, and / or autonomous driving chip 240. In some examples, central chip 220 may further execute a FuSa procedure for comparing and verifying the output of the corresponding pipeline to ensure the consistency of ML inference operations. In yet another example, central chip 220 may execute a thermal management procedure to ensure that the various components of SoC 200 operate within normal temperature ranges. See below for reference. Figure 3 Further description of the shared memory 230 is provided in the context of workload execution and system degradation.
[0049] Exemplary central core
[0050] Figure 3This is a block diagram illustrating an example central chip 300 of an MSoC arrangement according to the examples described herein. Figure 3 The central core 300 shown can correspond to, for example: Figure 2 The central chip 220 of the SoC 200 is shown. Furthermore, Figure 3 The sensor data input core 310 can correspond to Figure 2 The sensor data input core 210 is shown, and Figure 3 The workload processing core 320 shown can correspond to Figure 2 Any or more of the general-purpose computing chip 245, the ML accelerator chip 250, and / or the autonomous driving chip 240 shown. In yet another example, the processor 340 of the central chip 300 can also execute the workload as a runnable program with the support of the workload processing chip 320.
[0051] refer to Figure 3 The central core 300 may include a shared memory 360, which includes a reflection program 330, an application program 335, a thermal management program 337, and a functional safety (FuSa) program 338. As provided herein, the reflection program 330 may include a set of instructions for executing reflection workloads in a separate workload pipeline. Reflection workloads may include sensor data acquisition, sensor fusion, and inference tasks that facilitate scene understanding of the vehicle's surroundings. These tasks may include two-dimensional image processing, sensor fusion data processing (e.g., three-dimensional LiDAR, radar, and image fusion data), neural radiation field (NeRF) scene reconstruction, occupancy grid determination, object detection and classification, motion prediction, and other scene understanding tasks for autonomous vehicle operation.
[0052] As further provided herein, application 335 may include a set of instructions for operating vehicle control of the autonomous vehicle based on the output of the reflective workload pipeline. For example, application 335 may be powered by one or more processors 340 of central core 300 and / or one or more workload processing cores 320 (e.g., Figure 2 The autonomous driving core 240 is used to execute the motion plan of the vehicle dynamically based on the execution of the reflective workload, and to operate the vehicle's controls (e.g., acceleration, braking, steering and signaling systems) to execute the motion plan accordingly.
[0053] Thermal management program 337 can be executed by one or more processors 340 of the central chip to manage the heat generated by the SoC and trigger role switching between the primary SoC and the backup SoC in a dual SoC arrangement, for example, refer to Figure 5As described. In various examples, thermal management program 337 can operate a set of cooling components (e.g., heatsinks, fans, cold plates, heat pipes, synthesis jets, etc.) included with the SoC to manage the local temperature of the computing components within the normal operating range. In another example, thermal management program 337 can be executed by a dedicated CPU of central chip 300 and can communicate with workload processing chip 320 and sensor data input chip 310 for adjusting timing or current limiting of computing components, thereby further managing temperature. In some specific implementations, thermal management program 337 can execute on each SoC in a dual-SoC or multi-SoC arrangement and can further communicate with FuSa program 338 to, for example, power down the primary SoC when the temperature exceeds the nominal range and enable the standby SoC to take over autonomous inference and driving tasks.
[0054] According to the examples described herein, FuSa program 338 can be executed by one or more processors 340 of central chip 300 (e.g., dedicated FuSa CPU) to perform functional safety tasks of the SoC. As described throughout this disclosure, these tasks may include acquiring and comparing outputs from multiple independent pipelines corresponding to inference and / or autonomous vehicle control tasks. For example, one independent pipeline may include a workload corresponding to other vehicles operating around the vehicle identified in image data, and a second independent pipeline may include a workload corresponding to other vehicles operating around the vehicle identified in radar and LiDAR data. FuSa program 338 may execute the FuSa workload in another independent pipeline that acquires the outputs of the first and second independent pipelines to dynamically verify that they have identified the same vehicle.
[0055] In another example, the FuSa program 338 can operate to perform SoC monitoring in a dual-SoC arrangement, where the primary SoC performs inference and vehicle control tasks, and the standby SoC performs health monitoring on the primary SoC, with its dies in a low-power standby mode, ready to take over these tasks if any errors, faults, or failures are detected in the primary SoC. Further description of these FuSa functions is provided below. Figure 6 In yet another example, the FuSa program 338 can operate to monitor communication between cores and provide redundancy (e.g., via error correction code techniques) to ensure the reliability of communication between cores. Further description of these techniques is provided below. Figure 6 supply.
[0056] In various specific implementations, central core 300 may include a group of one or more processors 340 (e.g., transient-resistant CPUs and general-purpose computing CPUs) that can execute scheduler 342 to execute workloads as runnable programs in a set of independent pipelines. In some examples, one or more processors of processor 340 may execute reflective workloads according to reflective program 330 and / or application workloads according to application program 335. Thus, the processors 340 of central core 300 may reference, monitor, and update dependency information in workload entries of reservation table 350 as workloads become available and are executed accordingly. For example, when a workload is executed by a particular core, that core updates the dependency information of other workloads in reservation table 350 to indicate that the workload has been completed. This may include changing the bitwise operators or binary values (e.g., 0 to 1) representing the workload to indicate in reservation table 350 that the workload has been completed. Thus, the dependency information of all workloads that depend on the completed workload is updated accordingly.
[0057] According to the examples described herein, reservation table 350 may include workload entries, each indicating a workload identifier describing the workload to be executed, the address in cache memory 315 and / or HBM-RAM of the location of the raw or processed sensor data required to execute the workload as a runnable program, and any dependency information corresponding to dependencies that need to be resolved before executing the workload. In some respects, dependencies may correspond to other runnable programs that need to be executed before processing that particular workload. Once the dependencies of a particular workload are resolved, the workload entry may be updated (e.g., by the kernel executing the dependent workload, or by the processor 240 of central kernel 300 by executing scheduler 342). When a particular workload, as referenced in reservation table 350, has no dependencies, the workload may be executed in the corresponding pipeline by the corresponding workload processing kernel 320 as a runnable program or as part of a runnable program that includes multiple workloads.
[0058] In various specific implementations, sensor data input core 310 acquires sensor data from the vehicle's sensor system and stores the sensor data (e.g., image data, LiDAR data, radar data, ultrasonic data, etc.) in cache 315 of central core 300. Sensor data input core 310 can generate workload entries for reservation table 350, which include identifiers for the sensor data (e.g., identifiers for each image acquired from the various cameras of the vehicle's sensor system) and provide the addresses of the sensor data in cache memory 315. An initial workload can be performed on the raw sensor data by the processor 340 of central core 300 and / or workload processing core 320, which can update reservation table 350 to indicate that an initial workload has been completed.
[0059] As described herein, workload processing core 320 monitors reservation table 350 to determine whether a specific workload in its corresponding pipeline is ready to be executed as a runnable program. As an example, workload processing core 320 may use a workload window (e.g., a command window for multimedia data) to continuously monitor the reservation table, where a pointer can sequentially read each workload entry to determine if the workload has any unresolved dependencies. If one or more dependencies remain in a workload entry, the pointer proceeds to the next entry without executing the workload. However, if the workload indicates that all dependencies have been resolved (e.g., all workloads that a particular workload depends on have been executed), the relevant workload processing core 320 and / or processor 340 of central core 300 may execute the workload accordingly.
[0060] Therefore, workloads can be executed in an out-of-order manner, with some workloads buffered until their dependencies are resolved. Thus, to facilitate out-of-order execution of workloads, reservation table 350 includes an out-of-order buffer that allows workload processing core 320 to execute workloads in a specific order, governed by the deterministic resolution of their dependencies. It is conceivable that out-of-order execution of workloads can improve speed, increase power efficiency, and reduce the overall complexity of workload processing.
[0061] In some implementations, the workload processing kernel 320 can execute workloads as executable programs in a deterministic manner in each independent pipeline, such that successive workloads in the pipeline depend on the output of previous workloads in the pipeline. In various examples, the processor 340 and the workload processing kernel 320 can execute multiple independent workload pipelines in parallel, where each workload pipeline includes multiple workloads to be executed as runnable programs in a deterministic manner. Each workload pipeline can provide sequential outputs (e.g., for use in other workload pipelines, or for processing by the application 335 for autonomous operation of a vehicle). By simultaneously executing reflective workloads in independent pipelines, the application 335 can autonomously operate control of the vehicle along a travel route.
[0062] Exemplary software architecture
[0063] Figure 4 A software structure 400, comprising a set of executable programs 405 to be executed by a workload processing core 320, is depicted according to the example described herein. In various specific embodiments, the executable programs 405 may be executed by designated hardware components according to scheduler 342 and / or FuSa program 338, as referenced herein. Figure 3 As shown and described. For example, FuSa program 338 can monitor communications corresponding to the execution of runnable program 405 by a set of workload processing cores 320, and trigger software degradation when one or more degradation events (such as system overload, processing delay, overheating, excessive latency, etc.) are detected.
[0064] Figure 4 The software architecture 400 shown can correspond to autonomous driving software for autonomously operating a vehicle and is provided for illustrative purposes. Each of the runnable programs 405 in the software architecture 400 can correspond to one or more workload entries in the reservation table 350. Thus, when the dependency of a particular workload entry is resolved, the workload can be executed as a runnable program in an independent workload pipeline (e.g., by a designated workload processing core 320). Therefore, the collective execution of the runnable programs 405 in the software architecture 400 can correspond to each autonomous driving task described herein, such as perception, object detection and classification, scene understanding, ML inference, motion prediction, motion planning, and / or autonomous vehicle control tasks.
[0065] As described herein, the FuSa program 338 can monitor the output of each workload pipeline to verify them against each other (e.g., verifying consistency between inference runnables). The FuSa program 338 can further monitor communication within the performance network to determine if any errors have occurred, as referenced below. Figure 6 Described. Additionally, the FuSa program 338 can operate in conjunction with the thermal management program 337 to manage the heat generated by the SoC 200, switching between the primary and backup SoCs in the mSoC arrangement (see [link]). Figure 5 Furthermore, the execution of the software structure 400 is downgraded based on the security rating of the executable program 405 in the software structure 400 and / or the connections between executable programs 405.
[0066] In response to the FuSa procedure 338 that triggers system degradation, scheduler 342 may perform system degradation of software architecture 400 based on the security rating associated with each runnable program 405 and / or each connection between runnable programs 405. In some scenarios, scheduler 342 may perform degradation by reducing the execution frequency of a selected set of runnable programs in software architecture 400 (e.g., runnable programs with lower security ratings for communication connections). As an example, when an autonomous vehicle enters a pedestrian-intensive area, SoC 200 may perform scene understanding and inference operations. The central chip 300 and workload processing chip 320 may begin to overheat due to increased computational demands in pedestrian-intensive areas, causing FuSa procedure 338 to initiate degradation.
[0067] As presented herein, some executable programs may depend on the output of other executable programs. For example, a first executable program may involve detecting external dynamic entities near a vehicle (e.g., pedestrians, other vehicles, cyclists, etc.). A second executable program may involve predicting the motion of each external dynamic entity. Thus, the second executable program receives the output of the first executable program as input, and therefore these executable programs include connections within the software architecture.
[0068] Based on the examples described herein, software architecture 400 may include security ratings (e.g., ASIL ratings) for certain operable programs and connections between operable programs 405. Connections between operable programs may correspond to specified communication and / or dependencies between operable programs in software architecture 400. These associated security ratings may specify to FuSa program 338 and scheduler 342 which operable programs, communications, and / or connections between operable programs take precedence over other operable programs, communications, and / or connections when system degradation is required (e.g., by flow limiting to address overheating). Security ratings may further specify the importance of communication between operable programs 405 in terms of security priority ranking. For example, when degradation of the autonomous driving system is required, a connection 410 between two operable programs with an ASIL-D rating may take precedence over, for example, a connection 415 between two operable programs with an ASIL-B rating. In such examples, when system degradation occurs, communication or connections between operable programs with an ASIL-B rating may be degraded or flow limited, while communication between operable programs with an ASIL-D rating may remain robust.
[0069] In another example, each executable program or subset of executable programs can be associated with a security rating (e.g., ASIL rating). Figure 4 As shown, operable procedure 407 may be associated with an ASIL-B rating and may include relatively non-essential computational tasks, while operable procedure 406 may be associated with an ASIL-D rating and may include prioritized computational tasks. For example, operable procedure 407 may involve classifying the colors of fire hydrants and / or curbs in sensor data for parking purposes. Operable procedure 406 may involve predicting the movement of pedestrians near vehicles and is therefore crucial to the safety of those pedestrians.
[0070] For illustration, the software architecture 400 can be designed as an arrangement of nodes in a computational graph corresponding to the executable program 405, and connections between nodes representing dependencies or communication between the executable programs 405 (e.g., connections 410 and 415). In this document, it is conceivable that establishing safety ratings for the executable program 405 and / or for each connection between the executable programs 405 could facilitate an adaptive degradation scheme designed to maintain a high level of safety for the entire autonomous driving system (e.g., an overall ASIL-D rating).
[0071] In some implementations, the FuSa program 338 running on the central chip 300 of the computing system (e.g., SoC 200) executing the software architecture 400 can detect when the computing system experiences problems, such as system overload, processing latency, overheating, or any abnormal event (e.g., heavy rain or snow, lightning strike, tire failure, braking or steering failure, etc.). Based on the security rating established for node connections between executable programs 405, the FuSa program 338 can cause the central chip 300 to hierarchically degrade processes or executable programs 405 corresponding to node connections with lower security ratings (e.g., ASIL-B rating) and maintain the robustness of processes or executable programs corresponding to node connections with higher security ratings (e.g., ASIL-D rating).
[0072] In another example, degrading the execution of software architecture 400 may involve FuSa program 338 adaptively truncating or deforming the computation graph (e.g., a portion of software architecture 400 executed by workload processing core 320). For example, the connections between runnable programs 405 in software architecture 400 may be adaptively altered such that one or more truncated portions of the computation graph are temporarily ignored. In various embodiments, certain connections between runnable programs may be rearranged or otherwise edited such that the output of some runnable programs becomes the input of different runnable programs. As an example, when FuSa program 338 adaptively deforms the computation graph, the output of runnable program 406 may be switched from being the input of runnable program 412 to being the input of runnable program 411. According to the examples described herein, the connections between any number of runnable programs in software architecture 400 may be changed during the degradation process.
[0073] As provided in this article, degrading the execution of software architecture 400 can involve truncating a portion of the computation graph or software architecture 400. Figure 4 As shown, when FuSa program 338 performs a degradation, one or more runnable programs or combinations of runnable programs can be excluded from execution, as depicted by truncation 420 of software architecture 400. In such an example, workload processing core 320 will ignore the truncated runnables in truncation 420 until the degradation is reversed.
[0074] Additionally or alternatively, SoC 200 may store multiple software structures and / or computational graphs that can be executed based on the vehicle's level of autonomy. For example, when the vehicle switches from SAE Level 3 autonomy to SAE Level 4 autonomy, SoC 200 may switch from executing SAE Level 3 autonomous software structures to executing SAE Level 4 software structures. According to the embodiments provided herein, SoC 200 may also store multiple software structures and / or computational graphs for degrading purposes. Therefore, when FuSa program 338 triggers degrading, different software structures and / or computational graphs can be executed in the degraded state.
[0075] In some scenarios, the safety rating of the connection between the executable program and / or the executable program 405 can be dynamically adapted (e.g., based on the driving scenario). The process of acquiring sensor data from vehicle sensors (e.g., LiDAR sensors, image sensors, radar sensors, etc.), performing sensor data preprocessing (e.g., adjusting the contrast on the acquired images), combining sensor data (e.g., image stitching, sensor fusion, etc.), performing inference tasks (e.g., detecting and classifying objects of interest, such as other vehicles, pedestrians, traffic signs and signals, etc.), to performing motion prediction, motion planning, and vehicle control tasks can involve a set of safety priority rankings at any given time.
[0076] In the provided example, the vehicle may approach an area with very dense pedestrian traffic, where pedestrians are in the direction the vehicle is moving forward—while the area behind the vehicle may be relatively devoid of external entities. In this scenario, the hardware computing components may experience heavy workloads, potentially leading to processing latency and / or overheating. Furthermore, the safety rating of connections between runnable programs 405 can be dynamically adjusted based on the driving scenario. Scheduler 342 can detect the driving scenario and prioritize runnable programs and connections between runnable programs that involve pedestrian detection, which can be associated with an ASIL-D safety rating. In another example, scheduler 342 and / or FuSa program 338 can dynamically adjust the safety rating of runnable programs and / or runnable connections used for non-essential computational tasks. As provided herein, degradation may include reducing the inference or execution frequency of certain runnable programs, discarding or skipping images, ignoring radar data, etc.
[0077] Therefore, in some examples, the scheduler 342 of the central core 300 can dynamically change the schedule of runnable programs based on the system's performance when executing the software architecture. The node connections between nodes (runnable programs) in the software architecture 400 and runnable programs 405 can include fixed safety ratings or dynamically adjustable safety ratings (e.g., ASIL ratings) based on driving scenarios. It is conceivable that using the scheduler 342 to downgrade various tasks associated with lower safety ratings in the fuel tank components of the central core 300 (e.g., in ASIL-D rated memory) can maintain the effective operation of the autonomous vehicle in various driving scenarios and the overall high safety level of the autonomous driving system.
[0078] Multi-chip system
[0079] Figure 5 This is a block diagram depicting an example computing system 500 implementing a multi-system-on-a-chip (mSoC) according to the examples described herein. In various examples, computing system 500 may include a first SoC 510 having a first memory 515 and a second SoC 520 having a second memory 525, the first and second SoCs coupled via interconnects 540 (e.g., ASIL-D rated interconnects), enabling each of the first SoC 510 and the second SoC 520 to read each other's memories 515, 525. During any given session, the roles of the first SoC 510 and the second SoC 520 may alternate between a primary SoC and a backup SoC. As provided herein, the primary SoC may perform various autonomous driving tasks such as perception, object detection and classification, grid occupancy determination, sensor data fusion and processing, motion prediction (e.g., of dynamic external entities), motion planning, and vehicle control tasks. The backup SoC may maintain a set of computing components (e.g., CPU, ML accelerator, and / or memory chips) in a low-power state and continuously or periodically read the memory of the primary SoC.
[0080] For example, if the first SoC 510 is the primary SoC and the second SoC 520 is the backup SoC, the first SoC 510 executes a set of autonomous driving tasks and publishes status information corresponding to these tasks in the first memory 515. The second SoC 520 reads the published status information from the first memory 515 to continuously check whether the first SoC 510 is operating within nominal thresholds (e.g., temperature thresholds, bandwidth and / or memory thresholds, etc.) and whether the first SoC 510 is properly executing the set of autonomous driving tasks. Therefore, the second SoC 520 performs health monitoring and error management tasks for the first SoC 510 and takes over control of the set of autonomous driving tasks when a trigger condition is met. As provided herein, the trigger condition may correspond to a fault, failure, or other error experienced by the first SoC 510 that may affect the first SoC 510's execution of the set of tasks.
[0081] In various specific implementations, the second SoC 520 may publish status information corresponding to the computing units that are being maintained in a standby state (e.g., a low-power state, where the second SoC 520 maintains a ready state to take over the set of tasks from the first SoC 510). In such examples, the first SoC 510 may monitor the status information of the second SoC 520 by continuously or periodically reading the memory 525 of the second SoC 520 to similarly perform health check monitoring and error management on the second SoC 520. For example, if the first SoC 510 detects a fault, failure, or other error in the second SoC 520, the first SoC 510 may trigger the second SoC 520 to perform a system reset or restart.
[0082] In some examples, the first SoC 510 and the second SoC 520 may each include a functional safety (FuSa) component that performs health monitoring and error management tasks (e.g., a FuSa program 338 executed by one or more processors 340 of the central chip 300, as referenced). Figure 3 (As shown and described herein). The FuSa component can remain powered for each SoC, whether the SoC is operating in primary or standby mode. Thus, a standby SoC can keep its other components in a low-power state, where its FuSa component is powered on and performs the health monitoring and error management tasks described herein.
[0083] In various aspects, when the first SoC 510 operates as the main SoC, the status information published in the first memory 515 may correspond to the set of tasks being performed by the first SoC 510. For example, the first SoC 510 may publish any information corresponding to the surrounding environment of the vehicle (e.g., any external entities identified by the first SoC 510, their locations and predicted trajectories, detected objects such as traffic signals, signs, lane markings, and pedestrian crossings). The status information may also include the operating temperature of the computing components of the first SoC 510, the bandwidth usage and available memory of the first SoC 510's chips, and / or any faults or errors or information indicating faults or errors in these components.
[0084] In another aspect, when the second SoC 520 operates as a backup SoC, the status information published in the second memory 525 may correspond to the status of each computing unit of the second SoC 520. Specifically, these units may operate in a low-power state, in which they are ready to take over the set of tasks being performed by the first SoC 510. The status information may include whether the unit is operating within nominal temperature and other nominal ranges (e.g., available bandwidth, power, memory, etc.).
[0085] As described throughout this disclosure, the first SoC 510 and the second SoC 520 can switch between operating as a primary SoC and as a backup SoC (e.g., whenever the system 500 restarts). For example, in a computing session following a session in which the first SoC 510 operates as the primary SoC and the second SoC 520 operates as the backup SoC, the second SoC 520 can assume the role of the primary SoC, and the first SoC 510 can assume the role of the backup SoC. It is conceivable that this process of switching roles between the two SoCs can provide substantially uniform wear on the hardware components of each SoC, which can extend the overall lifespan of the computing system 500.
[0086] According to the implementation scheme, the first SoC 510 may be powered by a first power source (power supply), and the second SoC 520 may be powered by a second power source, which is independent of or isolated from the first power source. For example, in an electric vehicle, the first power source may include a battery pack for driving an electric motor of the vehicle, and the second power source may include an auxiliary power source for the vehicle (e.g., a 12-volt battery). In other specific implementations, the first and second power sources may include other types of power sources, such as dedicated batteries for each SoC 510, 520, or other power sources that are electrically isolated from each other or otherwise do not depend on each other.
[0087] It is conceivable that an mSoC arrangement of computing system 500 could be provided to improve the safety integrity level (e.g., ASIL rating) of computing system 500 and the entire autonomous driving system of the vehicle. As described herein, the autonomous driving system may include any number of dual SoC arrangements, each of which can perform a set of autonomous driving tasks. In doing so, a backup SoC dynamically monitors the health status of the primary SoC according to a set of functional safety operations, such that when a fault, failure, or other error is detected, the backup SoC can easily power on its components and take over the set of tasks from the primary SoC.
[0088] Functional safety accounting
[0089] Figure 6 This is a block diagram depicting a performance network and FuSa network for performing health monitoring, error correction, and system degradation, according to the examples described herein. In various examples, the FuSa CPU 600 may be included on the central die 300 of the SoC, or on each central die 300 of the mSoC 500 as described herein. Each FuSa CPU 600 may execute a FuSa program 602, which may correspond to a reference... Figure 3 The FuSa program 338 shown and described herein. As described herein, execution of the FuSa program 602 can cause the FuSa CPU 600 to execute references. Figure 5 The descriptions of the primary and backup SoC monitoring tasks, as well as the execution of the FuSa workload in the FuSa pipeline 420, are provided for reference. Figure 4 The comparison and verification of the described independent pipeline outputs.
[0090] In addition, Figure 6 In the example shown, multiple dies of a SoC can communicate with each other via a high-bandwidth performance network, which includes multiple sets of corresponding interconnects (e.g., interconnects 610 and 660) and network hubs (e.g., network hubs 615, 635, and 665). As described herein, multiple dies may include Figure 3 The sensor data input core 310, the central core 300, and the workload processing core 320 are composed of... Figure 6 Core A 605, core B 655, and any number of additional cores (not shown) are indicated. Furthermore, Figure 6 The cache memories 625 and 675 shown may represent cache memories associated with multiple dies and / or as referenced Figure 3 The cache memory 315 of the central chip 300 shown and described.
[0091] In various examples, raw sensor data, processed sensor data, and various communications between kernel A 605, kernel B 655, and FuSa CPU 600 can be transmitted via a high-bandwidth performance network including interconnects 610, 660, network hubs 615, 635, 665, and caches 625, 675. For example, if kernel A 605 includes a sensor data input kernel, kernel A 605 can obtain sensor data from various sensors of the vehicle and send the sensor data to cache 625 via interconnect 610 and network hub 615. In this example, if kernel B 655 includes a workload processing kernel, kernel B 655 can obtain sensor data from cache 625 via network hubs 615, 635, 665, and interconnect 660 to perform corresponding inference workloads based on the sensor data.
[0092] In some implementations, through the execution of FuSa program 602, FuSa CPU 600 can communicate with a high-bandwidth performance network via a performance on-chip network (NoC) 607 coupled to network hub 635. This communication may include, for example, acquiring output data from a separate pipeline to perform the comparison and verification steps described herein. Communication via the high-bandwidth performance network may also include communication for accessing the shared memory 515, 525 of each of the multiple SoCs 500, including a primary SoC and a backup SoC. In such examples, the FuSa CPU 600 of each SoC 510, 520 accesses each other's shared memory 515, 525 to determine if any fault, failure, or other error has occurred. As described herein, when a backup SoC detects a fault, failure, or error, the backup SoC takes over the tasks of the primary SoC (e.g., inference, scene understanding, vehicle control tasks, etc.).
[0093] In some aspects, interconnects 610 and 660 are used as high-bandwidth data paths for general data purposes to cache memories 625 and 675, and health control modules 620 and 670 and FuSa compute hubs 630, 640, and 680 are used as high-reliability data paths to send functional safety and scheduler information to the SoC's shared memory. The NoC and Network Interface Unit (NIU) on die A 605 and die B 655 can be configured to generate error correction code (ECC) data on both the high-bandwidth and high-reliability data paths. Each corresponding NIU on each paired die has the same ECC configuration, which generates and checks ECC data to ensure end-to-end error correction coverage.
[0094] According to various implementation schemes, the FuSa CPU 600 communicates via a FuSa network, which includes FuSa computing hubs 630, 640, and 680 and health control modules 620 and 670 via FuSa NoC 609. As provided herein, the FuSa network facilitates communication monitoring and error correction coding techniques. Figure 6 As shown, FuSa computing hubs 630, 640, and 680 can monitor communications sent through each network hub 615, 635, and 665 of the high-bandwidth network. Each of core A 605 and core B 655 can communicate with or include the health control module 620 and 670, through which ECC data, workload start and end communications, and scheduling information can be sent.
[0095] For the FuSa network data path, the NIU can send functional safety and scheduler information via health control modules 620 and 670 in two redundant transactions, where the second transaction sorts the bits in reverse order of the first transaction (e.g., from bit 31 to 0 on a 32-bit bus). Furthermore, if an error is detected during data transmission between chip A 605 and chip B 655 via the high-reliability FuSa network, the NIU can reduce the transmission rate to improve reliability.
[0096] In some examples, certain processors of the Chip A 605, Chip B 655, and / or FuSa CPU 600 may include those for running Figure 3 The transient-resistant CPU core is configured with a scheduler 342 that schedules workloads belonging to the reflection program 330, application program 335, thermal management program 337, and / or FuSa program 602. The transient-resistant CPU core is designed to resist and recover from transient faults caused by environmental factors such as cosmic radiation, power surges, and electromagnetic interference. These faults can cause the CPU to malfunction or produce incorrect results, potentially leading to system failure or security vulnerabilities. To address these issues, the transient-resistant CPU core may include a range of hardware-based fault detection and recovery mechanisms, such as redundant execution units, error correction code (ECC) memory, and register copying. These mechanisms can detect and correct errors in real time, ensuring that the CPU continues to operate correctly even in the presence of transient faults. Additionally, the transient-resistant CPU core may include various software-based fault-tolerance techniques, such as checkpointing and rollback, to further enhance system reliability and resilience.
[0097] In some respects, health control modules 620, 670 and FuSa computing hubs 630, 640, 680 can detect and correct errors in real time, ensuring that the CPU continues to operate correctly even in the presence of transient faults. For example, workload processing core A and the central core can perform error correction checks to verify that processed data is transmitted completely and without damage and stored in cache memories 625, 675. For example, for each processed data communication, the workload processing core can use the processed data to generate an error correction code (ECC) and send the ECC to the central core. As the data itself is transmitted along the high-bandwidth performance network between cores, the ECC is transmitted along the high-reliability FuSa network via FuSa computing hubs 630, 640, 680. Upon receiving the processed data, the central core can use the processed data to generate its own ECC, and the FuSa CPU 600 can perform a functional safety call in the central core mailbox to compare the two ECCs to ensure they match, verifying that the data was transmitted correctly.
[0098] Based on the examples described herein, the FuSa CPU 600 and FuSa program 602 can further monitor communications in the performance network and reliability network to obtain evidence that the system is experiencing system overload, network latency, low bandwidth, and / or overheating. Upon detecting these issues, the FuSa program 602 can perform any number of mitigation measures, such as switching the primary and standby roles of the SoC in the mSoC 500 arrangement, or initiating system degradation in a manner consistent with the description herein.
[0099] method
[0100] Figure 7 and Figure 8 This is a flowchart describing example methods for adapting runnable programs based on various examples and security ratings. Below are examples... Figure 7 and Figure 8 In the discussion of the methods, reference can be made to indicate the reference. Figures 1 to 6 The figures depict certain features with reference numerals. Furthermore, reference numerals may be used as... Figures 1 to 6 The computing system 100, SoC 200, and / or mSoC 500 shown and described perform the workload processing core 320 and central core 300 to execute the reference. Figure 7 and Figure 8 The steps are described in the flowchart. Further, refer to... Figure 7 and Figure 8 Some steps described in the flowchart can be performed before, in combination with, or after any other step, without necessarily in the corresponding order shown. In other specific implementations, in combination with Figure 7 and Figure 8 The term "computing system" as described can refer to any computing system, and is not limited to onboard computing systems in autonomous or semi-autonomous vehicles. For example, computing systems can be included in robotic systems, aircraft, marine vehicles, autonomous agricultural equipment, server farms or data centers, personal computing devices, etc.
[0101] refer to Figure 7 At block 700, the computing system can acquire sensor data from the sensor system. As provided herein, the sensor system may include one or more sensor types, such as LiDAR sensors, image sensors or cameras, radar sensors, ultrasonic sensors, microphones, proximity sensors, etc. In other embodiments, the sensor system may be included on an autonomous or semi-autonomous vehicle and may provide a continuous sensor view of the vehicle's surrounding environment. At block 705, based on the sensor data, the computing system may schedule the execution of runnable program 405 according to software architecture 400 (e.g., via scheduler 342). As provided herein, software architecture 400 may include all the software required to perform the aforementioned reflection functions and / or application functions, which may be combined to perform all sensor data acquisition, perception, object detection and classification, scene understanding, ML inference, motion prediction, motion planning, and / or vehicle control tasks required for autonomous operation of the vehicle along a driving path.
[0102] In various implementations, software architecture 400 can define the connections between runnable programs 405, where associations or dependencies exist. For example, the output of a first runnable program (e.g., an object detection algorithm that identifies any objects of interest in sensor fusion data) can be included as input to a second runnable program (e.g., an object classification algorithm that classifies each detected object), the output of which can be included as input to a third runnable program (e.g., a motion prediction algorithm responsible for predicting the motion of each dynamic object classified by the second runnable program), and so on. As provided herein, these runnable programs can be executed according to a dynamic scheduling and reservation table 350, which identifies when the workload corresponding to each runnable program is available (e.g., when workload dependency information has been satisfied or otherwise resolved).
[0103] As provided herein, each connection or subset of connections between operable programs 405 in software architecture 400 can be associated with a safety rating (e.g., ASIL rating). More critical connections between operable programs can be associated with a higher safety rating compared to less critical connections. For example, connections between operable programs that process data in the forward operating direction of a vehicle can be associated with a higher safety rating compared to connections between operable programs that process data behind a vehicle. As another example, connections between operable programs that detect pedestrians or other vulnerable road users (VRUs) can include a higher safety rating compared to connections between operable programs that detect fire hydrants, curb colors, parking meters, etc.
[0104] At box 710, during the execution of the executable program 405 in software architecture 400, the computing system can detect degradation events in the computing system. As provided herein, degradation events can correspond to system overload, such as overload of computational tasks or the executable program required to safely navigate a given travel path. Such system overload scenarios may occur when a vehicle enters areas of high processing load (such as dense urban environments (e.g., extreme cases involving pedestrians, bike lanes, external vehicles, traffic signs, etc.)) and / or areas with complex traffic and right-of-way rules. In another example, degradation events can correspond to one or more sensor failures, such as camera or LiDAR sensor malfunction, forcing the computing system to rely on other forms of sensor data.
[0105] As another example, a degradation event could correspond to a computing system temperature rise or the onset of overheating (e.g., due to increased processing, hardware wear, hardware failure, thermal runaway, etc.). In such an example, thermal management program 337 could perform a set of thermal mitigation tasks, such as enabling heatsinks, cooling fans, radiators, and / or switching between the primary and backup SoCs in an mSoC 500 computing system implementation. In the event of thermal runaway, the temperature rise could cause a decrease in resistance in wires and other computer components, which could lead to a further increase in temperature during critical feedback cycles. In this case, FuSa program 338 could detect thermal runaway (e.g., by monitoring the performance network) and initiate degradation in the manner described herein. For example, at block 715, if the heat level generated after thermal mitigation measures still exceeds the critical temperature, FuSa program 338 could selectively degrade the execution of executable program 405 in software architecture 400 based on a security rating of the executable program and / or the connections between executable programs in software architecture 400.
[0106] As described herein, such degradation may include: reducing the processing frequency of a selected device (e.g., a specific workload processing kernel), executing a runnable on that selected device and the runnable or runnable connection on that selected device having a relatively low security rating (e.g., ignoring every other image or sensor data iteration), reducing the frequency at which a specific computational task or runnable is executed, ignoring data from certain non-critical sensors, temporarily blocking the execution of certain runnables, and so on. As further described herein, such degradation may be implemented by a scheduler 342 using a reservation table 350, which may indicate when certain workloads are available for execution as runnables. For example, the scheduler 342 may flush certain workloads in the unordered buffer of the reservation table 350 without execution (e.g., data corresponding to a workload may be transferred to an HBM kernel instead of being processed by the workload processing kernel 320), or the reservation table 350 may selectively remove certain workloads associated with runnables having low security rating connections. It is conceivable that such methods could be incorporated into autonomous vehicle computing systems to prevent system failures, malfunctions, or increased wear due to overheating. This could provide additional redundancy to the thermal management program 337 and mSoC 500 arrangement, which could increase the overall ASIL rating of the entire computing system.
[0107] Figure 8 This is another flowchart describing a method for adapting an operable program based on security ratings, using various examples. (Reference) Figure 8 At block 800, the FuSa procedure 338 of the computing system can detect a degradation event when the computing system executes the runnable program 405 in the software architecture 400 according to a schedule. At block 805, the computing system can then determine the driving scenario to infer whether the driving scenario is associated with a degradation event. As provided herein, the connection between the runnable program in the software architecture 400 and / or the runnable program 405 can be associated with a safety rating (e.g., an ASIL rating), which can be dynamically adjusted or configured based on the driving scenario.
[0108] At box 810, the computing system can configure or adjust the security rating of the operands and / or operand connections in software architecture 400. For example, a driving scenario may involve a shift in the driving environment (e.g., from a highway where operands processing forward sensor data are prioritized to an urban environment where operands processing closer proximity data are more important). In another example, the degradation trigger for initiating degradation may include the changing driving scenario itself (e.g., the opposite of computing system overheating). In such examples, a reflection procedure (performing inference operations) may trigger the computing system's scheduler 342 to preemptively initiate the degradation process described herein.
[0109] In various specific implementations, at block 815, the computing system can selectively degrade the execution of software architecture 400 based on the reconfigured or adjusted safety ratings of the connections between operable programs 405. In some examples, the safety ratings of operable programs and / or operable program connections can be hierarchically adaptive or adjustable. For example, when the system performs a "Level 1" degradation, some operable program connections that may be irrelevant in a given driving scenario may still include high safety ratings.
[0110] In such an example, at decision box 820, the FuSa procedure 338 of the computing system can determine whether the degradation event has been resolved. If not, the FuSa procedure 338 can reconfigure the safety rating and cause the scheduler 342 to hierarchically degrade additional runnable programs or runnable program connections (e.g., "secondary" degradation, etc.) until the degradation event has been resolved (e.g., the system cools down below a critical temperature threshold). When the degradation event has been resolved (e.g., by a decrease in temperature and / or a change in the driving scenario), at box 825, the scheduler can selectively and / or hierarchically reverse the degradation accordingly.
[0111] It is conceivable that the examples described herein can be extended to the individual elements and concepts described herein (independent of other concepts, ideas, or systems), as well as to combinations of elements described anywhere in this application. Although the examples are described in detail herein with reference to the accompanying drawings, it should be understood that these concepts are not limited to those precise examples. Therefore, many modifications and variations will be apparent to those skilled in the art. Accordingly, the scope of these concepts is intended to be defined by the following claims and their equivalents. Furthermore, it is conceivable that a particular feature described separately or as part of an example can be combined with other separately described features or as part of other examples, even if the other features and examples do not mention that particular feature.
Claims
1. A computing system comprising: a sensor data input corelet to obtain sensor data from a sensor system; a set of workload processing corelets; and a central corelet comprising a shared memory comprising a scheduler to schedule a set of runnable programs to be executed based at least in part on the sensor data, the set of runnable programs being included in a software structure for execution by the set of workload processing corelets, wherein each runnable program in the software structure and each connection between runnable programs is associated with a safety rating to facilitate a degradation of execution of the software structure.
2. The computing system of claim 1, wherein, The safety rating of each runnable program and each connection between runnable programs comprises an automotive safety integrity level (ASIL) rating.
3. The computing system of claim 1, wherein, The central corelet comprises a functional safety (FuSa) program to: (i) monitor communications corresponding to execution of the runnable programs by the set of workload processing corelets, and (ii) trigger the degradation upon detecting data corresponding to a system overload, processing delay, or overheating.
4. The computing system of claim 3, wherein, In response to the FuSa program triggering the degradation, the scheduler executes the degradation based on the safety rating of each runnable program and each connection between the runnable programs.
5. The computing system of claim 4, wherein, The degradation by the scheduler comprises reducing a frequency of execution of a selected set of runnable programs in the software structure.
6. The computing system of claim 5, wherein, The scheduler uses a reservation table to reduce the frequency of execution of the selected set of runnable programs, the reservation table identifying when the selected set of runnable programs are ready for execution.
7. The computing system of claim 3, wherein, The FuSa program triggers the degradation by performing at least one of: (i) morphing the software structure to change one or more connections between the runnable programs, or (ii) truncating the software structure to exclude execution of one or more runnable programs.
8. The computing system of claim 1, wherein, The sensor data input corelet, the set of workload processing corelets, and the central corelet are included on a universal corelet interconnect express (UCIe) system on a chip (SoC) arrangement.
9. The computing system of claim 1, wherein, The computing system comprises an in-vehicle vehicle computer for an autonomous vehicle.
10. A non-transitory computer-readable medium storing instructions that, when executed by one or more processors of a computing system, cause the computing system to: obtain sensor data from a sensor system using a sensor data input corelet; scheduling, via execution of a scheduler on a central core including shared memory, a set of runnable programs to be executed based at least in part on the sensor data, the set of runnable programs included in a software fabric by a set of workload processing cores, wherein, each runnable program in the software structure and each connection between runnable programs is associated with a safety rating to facilitate a degradation of execution of the software structure.
11. The non-transitory computer readable medium of claim 10, wherein, The safety rating of each runnable program and each connection between runnable programs comprises an automotive safety integrity level (ASIL) rating.
12. The non-transitory computer readable medium of claim 10, wherein, The central core particle includes a functional safety (FuSa) program that (i) monitors communications corresponding to execution of the set of workload processing core particles on the runnable programs and (ii) triggers the degradation upon detecting data corresponding to system overload, processing delay, or overheating.
13. The non-transitory computer-readable medium of claim 12, wherein, The scheduler performs the degradation based on the security rating of each runnable program and each connection between the runnable programs in response to the FuSa program triggering the degradation.
14. The non-transitory computer-readable medium of claim 13, wherein, The degradation by the scheduler includes reducing a frequency of execution of a selected set of runnable programs in the software structure.
15. The non-transitory computer-readable medium of claim 14, wherein, The scheduler uses a reservation table to reduce the frequency of execution of the selected set of runnable programs, the reservation table identifying when the selected set of runnable programs are ready for execution.
16. The non-transitory computer readable medium of claim 12, wherein, The FuSa program triggers the degradation by performing at least one of (i) morphing the software structure to change one or more connections between the runnable programs or (ii) truncating the software structure to exclude execution of one or more runnable programs.
17. The non-transitory computer-readable medium of claim 10, wherein, The sensor data input core particle, the set of workload processing core particles, and the central core particle are included on a universal core interconnect express (UCIe) system on a chip (SoC) arrangement.
18. The non-transitory computer-readable medium of claim 10, wherein, The computing system includes an on-board vehicle computer for an autonomous vehicle.
19. A computer-implemented method of adapting runnable programs, the method performed by one or more processors and comprising: obtaining sensor data from a sensor system using a sensor data input core particle; scheduling, via a scheduler executing on a central core particle that includes a shared memory, a set of runnable programs to be executed based at least in part on the sensor data, the set of runnable programs included in a software structure by a set of workload processing core particles, wherein each runnable program in the set of runnable programs and each connection between runnable programs is associated with a security rating to facilitate degradation of execution of the software structure.
20. The method of claim 19, wherein, The security rating of each runnable program and each connection between the runnable programs includes an automotive safety integrity level (ASIL) rating. The security rating of each runnable program and each connection between the runnable programs includes an automotive safety integrity level (ASIL) rating.