Adaptation of performance levels of feasible elements based on safety assessments
Patent Information
- Application Number
- JP2026506076
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-08-02
- Filing Date
- 2024-07-25
- Publication Date
- 2026-09-01
AI Technical Summary
【0003】 ある一定の実装形態では、中央チップレットは、(i)ワークロード処理チップレットのセットによる実行可能要素の実行に対応する通信を監視し、(ii)システム過負荷、処理遅延、過熱、又はシステムデグラデーションを必要とするSoCに影響を及ぼす他の問題に対応するデータを検出すると、ソフトウェア構造のデグラデーションをトリガする機能安全(FuSa)プログラムを含むことができる。FuSaプログラムがデグラデーションをトリガすることに応答して、スケジューリングプログラムは、ソフトウェア構造内の各実行可能要素及び/又は実行可能要素間の各接続の安全性評価に基づいて、ソフトウェア構造のデグラデーションを実行する。いくつかの例では、スケジューリングプログラムによるソフトウェア構造のデグラデーションは、ソフトウェア構造内の選択された実行可能要素のセットの実行頻度を低減することを含むことができる。更なる例では、スケジューリングプログラムは、選択された実行可能要素のセットが実行の準備ができているときを(例えば、満たされている従属関係情報に基づいて)識別する予約テーブルを使用して、選択された実行可能要素のセットの実行頻度を低減することができる。追加的又は代替的に、デグラデーションは、実行可能要素間の接続を変更するために、又はある一定の実行可能要素(例えば、低い安全性評価を有する実行可能要素)を実行から除外するために、ソフトウェア構造をモーフィング及び/又は切り捨てることを含むことができる。更なる例では、デグラデーションは、計算グラフ又はソフトウェア構造を階層的に切り替えること(例えば、レベル1のデグラデーション計算グラフ、レベル2のデグラデーション計算グラフ、レベル3のデグラデーション計算グラフなどを利用すること)を含むことができる。
Smart Images

Figure 2026529568000001_ABST
Abstract
Description
Background Art
[0001] Universal Chiplet Interconnect Express (UCIe) is an open standard for interconnection between multiple chiplets and serial buses that enables the manufacture of large-scale system-on-chip (SoC) packages with mixed components from various semiconductor manufacturers. It is contemplated that autonomous vehicle computing systems may operate using chiplet arrangements that comply with the UCIe standard. One goal of creating such computing systems is to achieve a robust Safety Integrity Level (SIL) for other critical electrical and electronic (E / E) automotive components of the vehicle. Summary of the Invention Means for Solving the Problems
[0002] Disclosed herein are systems and methods for scheduling a set of executable elements included in a software structure for execution by a set of workload processing chiplets. In accordance with a scheduling program, each executable element in the software structure and / or each connection between executable elements may be associated with a safety assessment (e.g., an Automotive Safety Integrity Level (ASIL) assessment) to facilitate degradation of the software structure. In various examples, the set of workload processing chiplets may be included in a system-on-chip (SoC) including a central chiplet that comprises the scheduling program and a reservation table containing workload information (e.g., dependency information when a workload is executable as an executable element by the workload processing chiplet).
[0003] In certain implementations, a central chiplet may include a Functional Safety (FuSa) program that triggers a degradation of the software structure when it (i) monitors communications corresponding to the execution of executable elements by a set of workload processing chiplets, and (ii) detects data corresponding to system overload, processing delay, overheating, or other issues affecting the SoC that require system degradation. In response to the FuSa program triggering the degradation, a scheduling program performs a degradation of the software structure based on the safety evaluation of each executable element and / or connection between executable elements in the software structure. In some examples, the degradation of the software structure by the scheduling program may include reducing the execution frequency of a selected set of executable elements in the software structure. In further examples, the scheduling program may reduce the execution frequency of a selected set of executable elements by using a reservation table that identifies when a selected set of executable elements is ready for execution (e.g., based on fulfilled dependency information). Additionally or alternatively, the degradation may include morphing and / or truncating the software structure to change connections between executable elements or to exclude certain executable elements (e.g., executable elements with a low safety evaluation) from execution. In further examples, degradation can include hierarchically switching between computation graphs or software structures (for example, using a Level 1 degradation computation graph, a Level 2 degradation computation graph, a Level 3 degradation computation graph, and so on).
[0004] According to various embodiments, the adaptable performance of executable elements in an SoC can be implemented for autonomous vehicle operation. For example, an SoC may include a sensor data input chiplet, a set of workload processing chiplets, and a central chiplet, and may be included in a UCIe SoC configuration. This SoC configuration may further include one or more machine learning (ML) accelerator chiplets, high-bandwidth memory (HBM) chiplets, and / or autonomous driving chiplets. [Brief explanation of the drawing]
[0005] The disclosures herein are illustrative and not limiting, and in the figures of the accompanying drawings, similar reference numbers refer to similar elements. [Figure 1] This block diagram shows an example of a computing system in which an embodiment described herein may be implemented, according to the examples described herein. [Figure 2] This is a block diagram showing an example of a system-on-a-chip (SoC) in which the examples described herein may be implemented. [Figure 3] This block diagram shows an example of a central chiplet in an SoC layout for executing executable elements, as described in this specification. [Figure 4] This specification illustrates a software structure consisting of a set of executable elements executed by a workload processing chiplet, as illustrated by the examples provided herein. [Figure 5] This block diagram shows an example of a multiple systems-on-a-chip (mSoC) according to the examples described herein. [Figure 6] This block diagram shows a performance network and a FuSa network for performing health monitoring, error correction, and system degradation, as illustrated in the examples described herein. [Figure 7] This flowchart illustrates an exemplary method for adapting feasible elements based on safety assessments, using various examples. [Figure 8]This flowchart illustrates an exemplary method for adapting feasible elements based on safety assessments, using various examples. [Modes for carrying out the invention]
[0006] In experimental and controlled test environments, system redundancy and Automotive Safety Level (ASIL) assessments of autonomous systems are typically not priority considerations. As autonomous driving capabilities continue to advance (e.g., beyond Level 3 autonomy) and autonomous vehicles become commonplace on public road networks, qualification and certification of E / E components related to the autonomous operation of these vehicles will be advantageous in ensuring their operational safety. Furthermore, novel methods for certifying hardware, software, and / or hardware / software combinations as qualified will also be advantageous in increasing public confidence and assurance that autonomous driving systems are safer than current standards. For example, certain safety standards for autonomous driving systems include safety thresholds corresponding to average human capabilities and attention. However, these statistics include vehicle incidents associated with drivers with reduced driving ability or distracted drivers, and do not take into account certain time windows where the risk of vehicle operation is inherently increased (e.g., bad weather conditions, nighttime driving, winding mountain roads, etc.).
[0007] The Automotive Safety Level (ASIL) is a risk classification framework defined by ISO 26262 (Functional Safety Standard for Automotive), typically established for the E / E components of a vehicle by conducting a risk analysis of potential hazards. This analysis includes determining the severity (i.e., the degree of injury that the hazard is expected to cause, classified from S0 (no injury) to S3 (life-threatening injury)), the frequency (i.e., the relative expected frequency of operating conditions under which injury may occur, classified from E0 (very unlikely) to E4 (high probability of injury under most operating conditions)), and the avoidability (i.e., the relative likelihood that the driver can act to avoid injury, classified from C0 (generally avoidable) to C3 (difficult or impossible to avoid)). Therefore, any safety objective for any potential hazard event includes a set of ASIL requirements.
[0008] Hazards identified as quality management (QM) do not indicate safety requirements. These QM hazards may be any combination of low frequency of occurrence, low severity of potential injury resulting from the hazard, and high likelihood of avoidance by the driver in preventing the hazard and / or injury. Other hazard events are classified as ASIL-A, ASIL-B, ASIL-C, or ASIL-D depending on the varying levels of severity, frequency, and avoidability corresponding to the potential hazard. ASIL-D events correspond to the highest safety requirements (ASIL requirements) for safety systems or E / E components of safety systems, while ASIL-A includes the lowest integrity requirements. For example, vehicle airbags, anti-lock brakes, and power steering systems typically have an ASIL-D grade, and the risks associated with failure of these components (e.g., the likely severity of injury and lack of vehicle controllability to prevent those injuries) are relatively high.
[0009] As described herein, ASIL relates to both risk and risk-dependent requirements, and various combinations of severity, frequency, and avoidability are quantified to form a representation of the risk (for example, a vehicle airbag system may have a relatively low frequency classification but high values for severity and avoidability). As stated above, the amounts of severity, frequency, and avoidability of a given hazard have traditionally been determined using values for severity (e.g., S0-S3), frequency (e.g., E0-E4), and avoidability (e.g., C0-C3) in the ISO 26262 series, and these values are then used to classify ASIL requirements for components of a particular safety system. As described herein, certain safety systems can implement a variety of mitigation measures, which may range from warnings (e.g., visual, auditory, or tactile warnings) to mild interventions (e.g., brake assistance or steering assistance), severe interventions and / or evasive maneuvers (e.g., taking over control of one or more control mechanisms such as the steering system, acceleration system, or braking system), and fully autonomous control of the vehicle.
[0010] As illustrated in the examples described herein, a software structure for performing functions of an autonomous or semi-autonomous vehicle (e.g., perception, object detection and classification, scene understanding, ML reasoning, etc.) may include a set of executable elements that can be connected to other executable elements based on relevance or input / output dependency relationships. For example, a first executable element may be tasked with identifying and classifying pedestrians in image data, and a second executable element may be tasked with predicting the movement of each pedestrian. In this example, the second executable element receives the output of the first executable element as input. Thus, the first and second executable elements are connected in the software structure.
[0011] In the various examples described herein, executable elements within a software structure and / or connections between executable elements may be associated with a QM rating, or a safety rating such as ASIL-A, ASIL-B, ASIL-C, or ASIL-D. These safety ratings can be defined within the software structure and may indicate the importance of an executable element or a connection between two executable elements. For example, an executable element rated ASIL-D or a connection between two executable elements may correspond to the identification of traffic signals and the classification of traffic signal states (e.g., red, yellow, or green), which may be important for preventing vehicle collisions. In another example, an ASIL-B rated connection between two executable elements may correspond to the detection, classification, and calculation of speed differences of vehicles behind an autonomous vehicle in the same lane (e.g., those vehicles do not have the right of way over the autonomous vehicle).
[0012] It is intended that system degradation (e.g., when a computing system experiences critical temperatures due to overloaded computing and / or high ambient temperatures) can be facilitated by defining or relating executable elements or connections between executable elements within a software structure. Furthermore, degradation of the execution of a software structure can be managed by a functional safety (FuSa) component of the computing system, which may be tasked with (i) monitoring communication between executable elements (e.g., chiplets that organize data and execute executable elements based on that data), (ii) communicating with the thermal management component of the computing system, and (iii) communicating with a workload scheduling program that manages a reservation table containing workload entries indicating whether a workload is executable by one or more executable elements.
[0013] As provided herein, “degradation” of a software structure or degradation of the execution of a software structure generally refers to the selective reduction of certain computational tasks based on safety assessments associated with executable elements and / or connections between executable elements. For example, ISO 26262 refers to the degradation concept in the context of automotive safety, where the functionality of a given system (e.g., an E / E system) degrades to reach a safe state. Therefore, the vehicle’s autonomous system should be able to handle the degraded functionality in an appropriate manner. Otherwise, the autonomous driving system could, as a last resort, safely stop or park the vehicle.
[0014] In various implementations, degradation of a software structure or degradation of the execution of a software structure may include a scheduling program that reduces the execution frequency of a selected set of executable elements within the software structure. In a further example, the scheduling program may reduce the execution frequency of a selected set of executable elements by using a reservation table that identifies when the selected set of executable elements is ready for execution (for example, based on fulfilled dependency information). In yet another example, degradation may include reducing the processing frequency of a selected device (a specific chiplet) on which an executable element or executable element connection with a lower safety rating is executed (for example, ignoring every other image or sensor data iteration), decreasing or increasing the frequency with which an executable element is executed, ignoring data from certain non-critical sensors, or temporarily preventing the execution of certain executable elements.
[0015] As provided herein, degradation may include changing or editing connections between executable elements and / or truncating selected portions of a computation graph corresponding to a software structure in order to exclude one or more executable elements from being executed. In additional examples, degradation may include entire changes between software structures or computation graphs. For example, a vehicle may store multiple computation graphs corresponding to different SAE autonomy levels and / or different degradation levels.
[0016] As further described herein, this degradation can be implemented by a scheduling program using a reservation table, which can indicate when certain workloads are executable as executable elements. For example, a scheduling program can flush certain workloads in an unordered buffer of the reservation table without executing them (e.g., data corresponding to the workload may be transferred to an HBM chiplet without being processed by the workload processing chiplet), or the reservation table can selectively remove certain workloads associated with executable elements that have low safety rating connections.
[0017] As provided herein, a “degradation event” can include any event that can trigger degradation of a software structure or execution of a software structure, and may include excessive heat generation in the computing system or a part of the computing system, computer hardware failure or malfunction, sensor failure or malfunction, a specified driving scenario (e.g., an extreme case of a vulnerable road user), vehicle failure or malfunction, etc. For example, the thermal management components and functional safety components of the computing system may trigger degradation of a software structure if other thermal mitigation measures, such as starting a fan and / or water cooling system and switching the SoC between primary and backup roles, are insufficient to properly cool the computing system.
[0018] In certain implementations, an exemplary computing system can perform one or more of the functions described herein using a learning-based approach, such as by running an artificial neural network (e.g., a recurrent neural network, a convolutional neural network, etc.) or one or more machine learning models. Such a learning-based approach can further be adapted to a computing system that stores or includes one or more machine learning models. In one embodiment, the machine learning models may include unsupervised learning models. In one embodiment, the machine learning models may include neural networks (e.g., deep neural networks) or other types of machine learning models, including nonlinear and / or linear models. The neural network may include a feedforward neural network, a recurrent neural network (e.g., a long-term memory recurrent neural network), a convolutional neural network, or other forms of neural networks. Some exemplary machine learning models may leverage attentional mechanisms such as self-attention. For example, some exemplary machine learning models may include a multi-head self-attention model (e.g., a transformer model).
[0019] As described herein, “Network” or “one or more networks” may include any type of network or combination of networks that enable communication between devices. In one embodiment, the network may include one or more of the following: a local area network, a wide area network, the Internet, a secure network, a cellular network, a mesh network, a peer-to-peer communication link, or a combination thereof, and may include any number of wired or wireless links. Communication over the network(s) may be achieved via a network interface, for example, using any type of protocol, protection scheme, encoding, format, packaging, etc.
[0020] One or more examples described herein provide that methods, techniques, and operations performed by a computing device are performed by a program or as a computer implementation. “By a program,” as used herein, means through the use of code or computer executable instructions. These instructions can be stored in one or more memory resources of the computing device. The steps performed by the program may be automatic or not.
[0021] One or more examples described herein can be implemented using a programmatic module, engine, or component. A programmatic module, engine, or component may include a program, a subroutine, a portion of a program, or a software or hardware component capable of performing one or more defined tasks or functions. As used herein, a module or component may reside on a hardware component independent of other modules or components. Alternatively, a module or component may be a shared element or shared process of other modules, programs, or machines.
[0022] Some examples described in the present specification can generally require the use of a computing device including processing resources and memory resources. For example, one or more examples described in the present specification may be implemented wholly or partially using a network device (e.g., a router) in a computing device such as a server and / or a personal computer. Memory resources, processing resources, and network resources may be used in connection with establishing, using, or performing any of the examples described in the present specification, including performing any method or implementing any system.
[0023] Further, one or more examples described in the present specification may be implemented via the use of instructions executable by one or more processors. These instructions may be carried on a non-transitory computer-readable medium. Machines illustrated in or described with reference to the various figures below provide examples of processing resources and computer-readable media that can carry and / or execute instructions for implementing the examples disclosed herein. In particular, numerous machines illustrated with the examples of the present invention include a processor and various forms of memory for holding data and instructions. Examples of non-transitory computer-readable media include persistent memory storage devices such as hard drives in personal computers or servers. Other examples of computer storage media include portable storage units such as flash memory or magnetic memory. Computers, terminals, and network-enabled devices are all examples of machines and devices that utilize instructions stored in processors, memory, and computer-readable media. In addition, examples may be implemented in the form of a computer program, or in the form of a computer-usable carrier medium that can carry such a program.
[0024] Example Computing System Figure 1 is a block diagram showing an example computing system 100 in which embodiments described herein may be implemented, according to examples described herein. In one embodiment, the computing system 100 may include one or more control circuits 110, which may include one or more processors (e.g., microprocessors), one or more processing cores, programmable logic circuits (PLCs), or programmable logic / gate arrays (PLAs / PGAs), field programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), system-on-a-chip (SoCs), or any other control circuits. In some implementations, the control circuits 110 and / or the computing system 100 may be part of or form part of a vehicle control unit (also called a vehicle controller) that is embedded in or otherwise located in a vehicle (e.g., a Mercedes-Benz® car, truck, or van). For example, the vehicle controller may be an infotainment system controller (e.g., an infotainment head unit), a telematics control unit (TCU), an electronic control unit (ECU), a central powertrain controller (CPC), a central exterior and interior controller (CEIC), a zone controller, an autonomous vehicle control system, or any other controller (the term "or" is used herein interchangeably with "and / or"), or may include such controllers.
[0025] In one embodiment, the control circuit 110 may be programmed by one or more computer-readable instructions or computer-executable instructions stored in a non-temporary computer-readable medium 120. The non-temporary computer-readable medium 120 may be a memory device, also called a data storage device, and the memory device may include electronic storage devices, magnetic storage devices, optical storage devices, electromagnetic storage devices, semiconductor storage devices, or any preferred combination thereof. The non-temporary computer-readable medium 120 may, for example, form a floppy disk, a hard disk drive (HDD), a solid state drive (SDD) or solid state integrated memory, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), dynamic random access memory (DRAM), portable compact disc read-only memory (CD-ROM), digital versatile disk (DVD), and / or memory stick. In some cases, the non-temporary computer-readable medium 120 may store computer-executable instructions or computer-readable instructions, such as instructions that perform the methods described below in relation to Figures 7 and 8.
[0026] In various embodiments, the terms “computer-readable instructions” and “computer-executable instructions” are used to describe software instructions or computer code configured to perform various tasks and operations. In various embodiments, where computer-readable or computer-executable instructions form a module, the term “module” broadly refers to a set of software instructions or code configured to cause the control circuit 110 to perform one or more functional tasks. Modules and computer-readable / executable instructions may be described as performing various operations or tasks on which the control circuit 110 or other hardware components execute the module or computer-readable instructions.
[0027] In further embodiments, the computing system 100 may include a communication interface 140 that enables communication over one or more networks 150 to send and receive data. In various examples, the computing system 100 can use the communication interface 140 to communicate with fleet vehicles over one or more networks 150 to receive sensor data and implement the methods described throughout this disclosure. In certain embodiments, the communication interface 140 may be used to communicate with one or more other systems. The communication interface 140 may include any circuitry, components, software, etc., for communicating over one or more networks 150 (e.g., local area networks, wide area networks, the Internet, secure networks, cellular networks, mesh networks, and / or peer-to-peer communication links). In some implementations, the communication interface 140 may include, for example, one or more of the following for communicating data / information: communication controllers, receivers, transceivers, transmitters, ports, conductors, software, and / or hardware.
[0028] As an example of one embodiment, the control circuit 110 of the computing system 100 may include an SoC arrangement that facilitates various methods and techniques described throughout this disclosure. In various examples, the SoC may include a set of chiplets, which may include a central chiplet with shared memory where a reservation table is used to execute various autonomous driving workloads as executable elements in independent pipelines. According to embodiments described herein, the shared memory of the central chiplet may include a FuSa program that can be executed by the control circuit 110 to perform functional safety tasks for the SoC arrangement, as will be described in detail below.
[0029] System-on-a-chip example Figure 2 is a block diagram showing an example of a system-on-a-chip (SoC) 200 according to the examples described herein. The SoC 200 illustrated in Figure 2 may include additional components, and the components of the SoC 200 may be arranged in various alternative configurations other than those shown. Therefore, the SoC 200 in Figure 2 is described herein as an example arrangement for illustrative purposes and is not intended to limit the scope of this disclosure in any way.
[0030] Referring to Figure 2, the SoC200's sensor data input chiplet 210 can receive sensor data from various vehicle sensors 205 of the vehicle. These vehicle sensors 205 may include any combination of image sensors (e.g., single camera, binocular camera, fisheye lens camera, etc.), LIDAR sensors, radar sensors, ultrasonic sensors, proximity sensors, etc. Simultaneously with receiving the sensor data, the sensor data input chiplet 210 may automatically dump the sensor data into the cache memory 231 of the central chiplet 220. The sensor data input chiplet 210 may also include an image signal processor (ISP) responsible for capturing, processing, and improving images acquired from the various vehicle sensors 205. The ISP acquires raw image data and performs a series of complex image processing operations, such as color, contrast, and brightness correction, noise reduction, and image enhancement, to produce higher quality images ready for further processing or analysis by other chiplets of the SoC200. The ISP may also include features such as autofocus, image stabilization, and advanced scene recognition to further improve the quality of the captured images. The ISP can then store the high-quality images in the cache memory 231.
[0031] In some embodiments, the sensor data input chiplet 210 exposes identification information for each item of sensor data (e.g., image, point cloud map, etc.) to the shared memory 230 of the central chiplet 220, which functions as a central mailbox for synchronizing the workloads of various chiplets. The identification information may include details such as the address where the data is stored in the cache memory 231, the type of sensor data, which sensor captured the data, and a timestamp when the data was captured.
[0032] To communicate with the central chiplet 220, the sensor data input chiplet 210 transmits data through interconnection 211a. Interconnections 211a-f each represent a die-to-die (D2D) interface between chiplets of the SoC200. In some embodiments, interconnections 211a-f may include a high-bandwidth data path used for general data to the cache memory 231 and a high-reliability data path for transmitting functional safety (FuSa) and scheduler information to the shared memory 230. Depending on bandwidth requirements, interconnections 211a-f may include two or more die-to-die interfaces. For example, interconnection 211a may include two interfaces to support relatively high-bandwidth communication between the sensor data input chiplet 210 and the central chiplet 220.
[0033] In one embodiment, interconnects 211a-f implement the Universal Chiplet Interconnect Express (UCIe) standard and communicate via indirect mode, allowing each of the chiplet's host processors to access remote memory as if it were local memory. This is achieved by using a dedicated network-on-chip (NoC) network interface unit (NIU) that provides hardware-level support for remote direct memory access (RDMA) operation (e.g., enabling interference freedom between network-connected devices). In UCIe indirect mode, the host processor sends a request to the NIU, which then accesses the remote memory and returns the data to the host processor. This approach enables efficient, low-latency access to remote memory, which can be particularly useful in distributed computing and data-intensive applications. Furthermore, UCIe indirect mode can be used with a wide range of different network topologies and protocols, providing a high degree of flexibility.
[0034] In various examples, the SoC200 may include additional chiplets capable of storing, modifying, or otherwise processing sensor data cached by the sensor data input chiplet 210. The SoC200 may also include an autonomous driving chiplet 240 capable of executing autonomous vehicle perception, sensor fusion, trajectory prediction, and / or other autonomous driving algorithms. The autonomous driving chiplet 240 can be connected to a dedicated HBM-RAM chiplet 235, to which the autonomous driving chiplet 240 may expose all status information, variables, statistics, and / or processed sensor data processed by the autonomous driving chiplet 240.
[0035] In various applications, the system-on-chip 200 may further include a machine learning (ML) accelerator chiplet 240, specifically designed to accelerate AI workloads such as image inference or other sensor inference using machine learning to achieve high performance and low power consumption. The ML accelerator chiplet 240 may include an engine designed to efficiently process graph-based data structures commonly used in AI workloads, and a highly parallel processor that enables efficient processing of large amounts of data. The ML accelerator chiplet 240 may also include a dedicated hardware accelerator for common AI operations such as matrix multiplication and convolution, as well as a memory hierarchy designed to optimize memory access for AI workloads that often have complex memory access patterns.
[0036] The general-purpose computing chiplet 245 may provide general-purpose computing for the system-on-chip 200. For example, the general-purpose computing chiplet 245 may include a high-power central processing unit and / or a graphical processing unit, which may support the computing tasks of the central chiplet 220, the autonomous driving chiplet 240, and / or the ML accelerator chiplet 250.
[0037] In various implementations, the shared memory 230 may store programs and instructions for performing autonomous driving tasks. The shared memory 230 of the central chiplet 220 may further include a reservation table that provides various chiplets with the information necessary to perform individual tasks (e.g., sensor data items and their locations in memory). In various embodiments, the central chiplet 220 also includes a large cache memory 231 that supports invalidation and flashing of stored data. A further explanation of the shared memory 230 in the context of the central chiplet 220 is provided below with respect to Figure 3.
[0038] Cache misses and evicting from the cache memory 231 are transmitted by a high-bandwidth memory (HBM) RAM chiplet 255 connected to the central chiplet 220. The HBM-RAM chiplet 255 may contain status information, variables, statistics, and / or sensor data for all other chiplets. In certain examples, the information stored in the HBM-RAM chiplet 255 may be stored for a predetermined period (e.g., 10 seconds), after which the data is deleted or otherwise flushed. For example, if an autonomous vehicle experiences a malfunction, the information stored in the HBM-RAM chiplet 255 may contain all the information necessary to diagnose and resolve the malfunction. The cache memory 231 holds new data that is available with lower latency and lower power consumption compared to accessing data from the HBM-RAM chiplet 255.
[0039] As described herein, the shared memory 230 can accommodate a mailbox architecture in which an immediate program containing a set of instructions is used to execute workloads by the central chiplet 220, the general-purpose compute chiplet 245, and / or the autonomous operation chiplet 240. In certain examples, the central chiplet 220 may further execute a FuSa program that operates to compare and verify the output of each pipeline to ensure consistency in ML inference operations. In further examples, the central chiplet 220 may execute a thermal management program to ensure that various components of the SoC200 operate within a normal temperature range. A further explanation of the shared memory 230 in the context of workload execution and system degradation is provided below with respect to Figure 3.
[0040] Exemplary central chiplet Figure 3 is a block diagram showing an exemplary central chiplet 300 of a certain SoC configuration according to the examples described herein. The central chiplet 300 shown in Figure 3 may correspond to the central chiplet 220 of the SoC 200 shown in Figure 2. Furthermore, the sensor data input chiplet 310 in Figure 3 may correspond to the sensor data input chiplet 210 shown in Figure 2, and the workload processing chiplet 320 shown in Figure 3 may correspond to any one or more of the general-purpose computing chiplet 245, ML accelerator chiplet 250, and / or autonomous driving chiplet 240 shown in Figure 2. In a further example, the processor 340 of the central chiplet 300 may also execute a workload as an executable element supporting the workload processing chiplet 320.
[0041] Referring to Figure 3, the central chiplet 300 may include a shared memory 360 containing a reflex program 330, an application program 335, a thermal management program 337, and a functional safety (FuSa) program 338. As described herein, the reflex program 330 may include a set of instructions for executing reflex workloads in multiple independent workload pipelines. Reflex workloads may include sensor data acquisition, sensor fusion, and inference tasks, which facilitate scene understanding of the vehicle's surrounding environment. These tasks may include two-dimensional image processing, sensor fusion data processing (e.g., three-dimensional LiDAR, radar, and image fusion data), neural radiance field (NeRF) scene reconstruction, occupancy grid determination, object detection and classification, motion prediction, and other scene understanding tasks for autonomous vehicle operation.
[0042] As further described herein, the application program 335 may include a set of instructions for operating the vehicle controls of an autonomous vehicle based on the output of the immediate workload pipeline. For example, the application program 335 may be executed by one or more processors 340 and / or workload processing chiplets 320 of the central chiplet 300 (e.g., the autonomous driving chiplet 240 in Figure 2) to dynamically generate a vehicle motion plan based on the execution of the immediate workload and to operate the vehicle controls (e.g., acceleration, braking, steering, and signaling systems) to execute the motion plan accordingly.
[0043] The thermal management program 337 is executed by one or more processors 340 in the central chiplet and can manage the heat generated by the SoC and trigger role switching between the primary SoC and the backup SoC in a dual SoC configuration, as illustrated with reference to Figure 5. In various examples, the thermal management program 337 can operate a set of cooling components included in the SoC (e.g., heat sinks, fans, cold plates, heat pipes, synthetic jets, etc.) to manage the local temperature of the computing components within their normal operating range. In further examples, the thermal management program 337 may be executed by a dedicated CPU in the central chiplet 300 and can communicate with the workload processing chiplet 320 and the sensor data input chiplet 310 to further manage the temperature by adjusting the clocking or throttling of the computing components. In certain implementation forms, the thermal management program 337 may run on each SoC in a dual or multi-SoC configuration and can further communicate with the FuSa program 338 to, for example, power off the primary SoC when the temperature exceeds a nominal range, allowing the backup SoC to take over autonomous inference and operation tasks.
[0044] As illustrated herein, the FuSa program 338 may be executed by one or more processors 340 (e.g., dedicated FuSa CPUs) of the central chiplet 300 to perform functional safety tasks for the SoC. As described throughout this disclosure, these tasks may include acquiring and comparing outputs from multiple independent pipelines corresponding to inference and / or autonomous vehicle control tasks. For example, one independent pipeline may include a workload corresponding to identifying other vehicles operating around a vehicle in image data, and a second independent pipeline may include a workload corresponding to identifying other vehicles operating around a vehicle in radar and LIDAR data. The FuSa program 338 can execute the FuSa workload in another independent pipeline that acquires the outputs of the first and second independent pipelines to dynamically verify that they identified the same vehicle.
[0045] In a further example, FuSa program 338 can operate to perform SoC monitoring in a dual SoC configuration, where the primary SoC performs inference and vehicle control tasks, and the backup SoC performs health monitoring on the primary SoC while its chiplets are in low-power standby mode, and is ready to take over these tasks if an error, fault, or failure is detected on the primary SoC. A further description of these FuSa functions is provided below with reference to Figure 6. In yet another example, FuSa program 338 can operate to monitor communication between chiplets and provide redundancy (e.g., via error correction coding techniques) to ensure the reliability of communication between chiplets. A further description of these techniques is provided below with reference to Figure 6.
[0046] In various implementations, the central chiplet 300 may include a set of one or more processors 340 (e.g., transient-tolerant CPUs and general-purpose computing CPUs) capable of executing a scheduling program 342 for executing workloads as executable elements within a set of independent pipelines. In certain examples, one or more of the processors 340 may execute an immediate workload according to an immediate program 330 and / or an application workload according to an application program 335. Thus, the processors 340 of the central chiplet 300 can refer to, monitor, and update dependency information in the workload entries of the reservation table 350 as workloads become available and are executed accordingly. For example, when a workload is executed by a particular chiplet, the chiplet updates the dependency information of other workloads in the reservation table 350 to indicate that its workload has been completed. This may involve changing the bitwise operator or binary value representing the workload (e.g., from 0 to 1) to indicate in the reservation table 350 that the workload has been completed. Consequently, the dependency information of all workloads that are dependent on the completed workload is updated accordingly.
[0047] As described herein, the reservation table 350 may include workload entries, each of which includes a workload identifier describing the workload to be executed, the address in the cache memory 315 and / or HBM-RAM of the location of the unprocessed or processed sensor data required to execute the workload as an executable element, and any dependency information corresponding to dependencies that need to be resolved before the workload is executed. In certain embodiments, dependencies may correspond to other executable elements that need to be executed before processing that particular workload. When a dependency of a particular workload is resolved, the workload entry may be updated (for example, by the chiplet executing the dependent workload or by the processor 240 of the central chiplet 300 through the execution of the scheduling program 342). If no dependencies exist for a particular workload referenced in the reservation table 350, the workload may be executed by the corresponding workload processing chiplet 320 as an executable element in its respective pipeline or as part of an executable element containing multiple workloads.
[0048] In various implementations, the sensor data input chiplet 310 acquires sensor data from the vehicle's sensor system and stores that sensor data (e.g., image data, LiDAR data, radar data, ultrasonic data, etc.) in the cache 315 of the central chiplet 300. The sensor data input chiplet 310 can generate a workload entry in the reservation table 350, which includes an identifier for the sensor data (e.g., an identifier for each image acquired from various cameras in the vehicle's sensor system), and provide the address of that sensor data in the cache memory 315. An initial set of the workload may be performed by the processor 340 of the central chiplet 300 and / or the workload processing chiplet 320 for any unprocessed sensor data, thereby updating the reservation table 350 to indicate that the initial set of the workload is complete.
[0049] As described herein, the workload processing chiplet 320 monitors the reservation table 350 to determine whether a particular workload in each pipeline is ready to be executed as an executable element. For example, the workload processing chiplet 320 may continuously monitor the reservation table using a workload window (e.g., an instruction window for multimedia data), in which a pointer can sequentially read each workload entry to determine if the workload has any outstanding dependencies. If a workload entry still has one or more dependencies, the workload is not executed, and the pointer moves on to the next entry. However, if the workload indicates that all dependencies have been resolved (e.g., all workloads to which a particular workload depends have been executed), the associated workload processing chiplet 320 and / or processor 340 of the central chiplet 300 may execute the workload accordingly.
[0050] Therefore, workloads can be executed in a non-sequential manner, in which certain workloads are buffered until their dependencies are resolved. Thus, to facilitate non-sequential execution of workloads, the reservation table 350 includes a non-sequential buffer, which allows the workload processing chiplets 320 to execute workloads in an order deterministically governed by the resolution of dependencies. Non-sequential execution of workloads is intended to increase the overall processing speed, improve power efficiency, and reduce complexity of the workloads.
[0051] In certain implementations, the workload processing chiplet 320 can deterministically execute workloads as executable elements within each independent pipeline, such that successive workloads in the pipeline are subordinate to the output of preceding workloads within the pipeline. In various examples, the processor 340 and the workload processing chiplet 320 can execute multiple independent workload pipelines in parallel, each containing multiple workloads to be executed as executable elements in a deterministic manner. Each workload pipeline can provide a set of outputs (for example, for other workload pipelines or for processing by the application program 335 to operate the vehicle autonomously). Through the simultaneous execution of responsive workloads in independent pipelines, the application program 335 can autonomously control the vehicle along its travel route.
[0052] Exemplary software structure Figure 4 shows a software structure 400 consisting of a set of executable elements 405 executed by a workload processing chiplet 320, as described herein. In various implementations, the executable elements 405 may be executed by designated hardware components according to a scheduling program 342 and / or a FuSa program 338, as illustrated and described with respect to Figure 3. For example, the FuSa program 338 may monitor communications corresponding to the execution of the executable elements 405 by the set of workload processing chiplets 320 and, upon detecting one or more degradation events such as system overload, processing delay, overheating, or excessive delay, may trigger a degradation of the software structure.
[0053] The software structure 400 shown in Figure 4 can correspond to autonomous driving software used to operate a vehicle autonomously and is provided for illustrative purposes. Each of the executable elements 405 of the software structure 400 can correspond to one or more workload entries in the reservation table 350. Thus, when the dependency of a particular workload entry is resolved, the workload can be executed as an executable element in an independent workload pipeline (e.g., by a designated workload processing chiplet 320). Thus, the collective execution of the executable elements 405 in the software structure 400 can correspond to each of the autonomous driving tasks described herein, such as perception, object detection and classification, scene understanding, ML reasoning, motion prediction, motion planning, and / or autonomous vehicle control tasks.
[0054] As described herein, the FuSa program 338 can monitor the output of each workload pipeline and verify them against each other (for example, verifying consistency between inference executable elements). The FuSa program 338 can further monitor communications within the performance network to determine if an error has occurred, as described below with respect to Figure 6. In addition, the FuSa program 338 can work with the thermal management program 337 to manage the heat generated by the SoC 200, switch between the primary SoC and the backup SoC in the mSoC configuration (see Figure 5), and degrade the execution of the software structure 400 based on a safety assessment of the executable elements 405 and / or the connections between the executable elements 405 within the software structure 400.
[0055] In response to FuSa program 338 triggering system degradation, scheduling program 342 may perform system degradation of the software structure 400 based on the safety evaluation associated with each of the executable elements 405 and / or each connection between the executable elements 405. In certain scenarios, scheduling program 342 may perform degradation by reducing the execution frequency of a selected set of executable elements within the software structure 400 (e.g., executable elements with communication connections that have a lower safety evaluation). As an example, SoC200 may perform scene understanding and reasoning operations when an autonomous vehicle enters a pedestrian-heavy area. The central chiplet 300 and workload processing chiplet 320 may begin to overheat due to the increased computational requirements in the pedestrian-heavy area, potentially causing FuSa program 338 to initiate degradation.
[0056] As provided herein, certain executable elements may depend on the output of other executable elements. For example, a first executable element may include the detection of external dynamic entities within the vicinity of a vehicle (e.g., pedestrians, other vehicles, cyclists, etc.). A second executable element may include predicting the motion of each external dynamic entity. Thus, the second executable element receives the output of the first executable element as input, and these executable elements therefore include connections within the software structure.
[0057] As illustrated herein, the software structure 400 may include safety assessments (e.g., ASIL assessments) of certain executable elements and connections between executable elements 405. Connections between executable elements may correspond to communications and / or dependencies that specified executable elements have with one another within the software structure 400. These relevant safety assessments can instruct the FuSa program 338 and the scheduling program 342 which executable elements, communications, and / or connections between executable elements have priority over other executable elements, communications, and / or connections when system degradation is required (e.g., through throttling to address overheating). The safety assessments may further indicate the importance of communications between executable elements 405 from a safety prioritization standpoint. For example, a connection 410 between two executable elements with an ASIL-D rating may have priority over a connection 415 between two executable elements with an ASIL-B rating when degradation of the autonomous driving system is required. In such an example, when system degradation occurs, communication or connections between executable elements with an ASIL-B rating may degrade or be throttled, while communication between executable elements with an ASIL-D rating may remain robust.
[0058] In further examples, each feasible element or subset of feasible elements may be associated with a safety assessment (e.g., an ASIL assessment). As shown in Figure 4, feasible element 407 may be associated with an ASIL-B assessment and may include relatively unimportant computational tasks, while feasible element 406 may be associated with an ASIL-D assessment and may include preferred computational tasks. For example, feasible element 407 may include the classification of fire hydrant and / or curb color in sensor data for parking purposes. Feasible element 406 may include the prediction of pedestrian motion in the vicinity of a vehicle and is therefore important for the safety of those pedestrians.
[0059] For illustrative purposes, the software structure 400 can be envisioned as the arrangement of nodes in a computation graph corresponding to executable elements 405, and the connections between nodes (e.g., connections 410 and 415) representing dependencies or communications between the executable elements 405. It is intended herein that establishing safety assessments for each executable element 405 and / or connection between executable elements 405 can facilitate adaptive degradation schemes designed to maintain a high safety level for the entire autonomous driving system (e.g., an overall ASIL-D rating).
[0060] In certain implementations, the FuSa program 338, running on the central chiplet 300 of a computing system executing a software structure 400 (e.g., SoC200), can detect when the computing system is experiencing problems such as system overload, processing delays, overheating, or any abnormal occurrence (e.g., heavy rain or snow, lightning, tire failure, braking or steering failure). According to the safety ratings established for node connections between executable elements 405, the FuSa program 338 can cause the central chiplet 300 to hierarchically degrade processes or executable elements 405 corresponding to node connections with lower safety ratings (e.g., ASIL-B rating) while maintaining the robustness of processes or executable elements corresponding to node connections with higher safety ratings (e.g., ASIL-D rating).
[0061] In a further example, degrading the execution of the software structure 400 may include the FuSa program 338 adaptively truncating or morphing the computation graph (e.g., the portion of the software structure 400 being executed by the workload processing chiplet 320). For example, connections between executable elements 405 within the software structure 400 may be adaptively modified so that one or more truncated portions of the computation graph are temporarily ignored. In various embodiments, certain connections between executable elements may be rearranged or otherwise edited so that the output of certain executable elements becomes the input of different executable elements. For example, when the FuSa program 338 adaptively morphs the computation graph, the output of executable element 406 may switch from being the input of executable element 412 to being the input of executable element 411. According to the examples described herein, connections between any number of executable elements within the software structure 400 may be modified during the degradation process.
[0062] As provided herein, degrading the execution of a software structure 400 may include truncating the computation graph or a portion of the software structure 400. As shown in Figure 4, when the FuSa program 338 performs a degradation, one or more executable elements or combinations of executable elements may be excluded from execution, as indicated by the truncation 420 of the software structure 400. In such an example, the workload processing chiplet 320 ignores the truncated executable elements in the truncation 420 until the degradation is reversed.
[0063] Additionally or alternatively, the SoC200 can store multiple software structures and / or computation graphs that can be executed based on the vehicle's level of autonomy. For example, when the vehicle switches from SAE Level 3 to SAE Level 4 autonomy, the SoC200 can switch from executing an SAE Level 3 autonomy software structure to executing an SAE Level 4 software structure. According to embodiments provided herein, the SoC200 can also store multiple software structures and / or compute graphs for degradation purposes. Thus, when the FuSa program 338 triggers degradation, different software structures and / or computation graphs can be executed in the degradation state.
[0064] In certain scenarios, the safety assessment of the connections between executable elements and / or executable elements 405 can be adapted dynamically (e.g., based on the driving scenario). The process, from acquiring sensor data from vehicle sensors (e.g., LIDAR sensors, image sensors, radar sensors, etc.), performing preprocessing of the sensor data (e.g., adjusting the contrast of the acquired images), combining the sensor data (e.g., stitching images, sensor fusion, etc.), and performing inference tasks (e.g., detecting and classifying objects of interest such as other vehicles, pedestrians, traffic signs and signals, etc.) to performing motion prediction, motion planning, and vehicle control tasks, may include a set of safety priorities at any given time.
[0065] In the provided example, a vehicle may approach an extremely pedestrian-heavy area with pedestrians in front of it, but with relatively few external entities behind it. In such a scenario, hardware computing components may experience heavy workloads that could cause processing delays and / or overheating. Furthermore, the safety assessment of connections between executable elements 405 can be dynamically adjusted based on the driving scenario. The scheduling program 342 may detect the driving scenario and prioritize executable elements and connections between executable elements, including pedestrian detection, which may be associated with an ASIL-D safety assessment. In a further example, the scheduling program 342 and / or the FuSa program 338 may dynamically adjust the safety assessment of executable elements and / or executable element connections for less critical computing tasks. As provided herein, degradation may include reducing the inference or execution frequency of certain executable elements, discarding or skipping images, ignoring radar data, and so on.
[0066] Therefore, in certain cases, the scheduling program 342 of the central chiplet 300 can dynamically change the scheduling of executable elements based on the system's performance when running the software structure as a whole. The nodes (executable elements) within the software structure 400 and the node connections between executable elements 405 can include fixed or dynamically adjustable safety ratings (e.g., ASIL ratings) based on the driving scenario. By using the scheduling program 342 within the mailbox component of the central chiplet 300 (e.g., in memory for ASIL-D ratings) to degrade various tasks associated with lower safety ratings, it is intended that effective operation of the autonomous vehicle and a high level of safety of the autonomous driving system as a whole can be maintained in various driving scenarios.
[0067] Multiple System-on-a-Chip Figure 5 is a block diagram showing an example computing system 500 implementing multiple systems-on-a-chip (mSoCs) according to the examples described herein. In various examples, the computing system 500 may include a first SoC 510 having a first memory 515 and a second SoC 520 having a second memory 525, connected by an interconnection 540 (e.g., an ASIL-D evaluation interconnection), which allows each of the first SoC 510 and the second SoC 520 to read each other's memories 515 and 525. During a given session, the first SoC 510 and the second SoC 520 can alternately switch roles between primary and backup SoC. As described herein, the primary SoC can perform a variety of autonomous driving tasks, such as perception, object detection and classification, grid occupancy determination, sensor data fusion and processing, motion prediction (e.g., of dynamic external entities), motion planning, and vehicle control tasks. The backup SoC maintains a set of computing components (e.g., CPU, ML accelerator, and / or memory chiplets) in a low-power state and can continuously or periodically read from the primary SoC's memory.
[0068] For example, if the first SoC510 is the primary SoC and the second SoC520 is the backup SoC, the first SoC510 performs a set of autonomous driving tasks and exposes state information corresponding to these tasks to the first memory 515. The second SoC520 reads the exposed state information from the first memory 515 and continuously checks whether the first SoC510 is operating within nominal thresholds (e.g., temperature threshold, bandwidth and / or memory threshold) and whether the first SoC510 is properly performing the set of autonomous driving tasks. Thus, the second SoC520 performs health monitoring and error management tasks for the first SoC510 and takes over control of the set of autonomous driving tasks when trigger conditions are met. As described herein, trigger conditions may correspond to failures, malfunctions, or other errors suffered by the first SoC510 that may affect the performance of the task set by the first SoC510.
[0069] In various implementations, the second SoC520 may expose state information corresponding to the fact that the computing components are being kept in a standby state (for example, a low-power state in which the second SoC520 maintains preparation for taking over a set of tasks from the first SoC510). In such examples, the first SoC510 may also perform health check monitoring and error management for the second SoC520 by continuously or periodically reading the second SoC520's memory 525 to monitor its state information. For example, if the first SoC510 detects a fault, failure, or other error in the second SoC520, the first SoC510 may trigger the second SoC520 to perform a system reset or reboot.
[0070] In certain examples, the first SoC 510 and the second SoC 520 may each include a functional safety (FuSa) component that performs health monitoring and error management tasks (for example, a FuSa program 338 executed by one or more processors 340 of the central chiplet 300, as illustrated and described with respect to Figure 3). The FuSa component can be kept powered on for each SoC, regardless of whether the SoC operates in primary or backup mode. Thus, the backup SoC may keep other components in a low-power state while the FuSa component is powered up and performing the health monitoring and error management tasks described herein.
[0071] In various embodiments, when the first SoC 510 operates as a primary SoC, the state information exposed in the first memory 515 may correspond to a set of tasks performed by the first SoC 510. For example, the first SoC 510 may expose any information corresponding to the vehicle's surrounding environment (e.g., external entities identified by the first SoC 510, their locations and predicted trajectories, and detected objects such as traffic signals, signs, lane markings, and crosswalks). The state information may further include the operating temperature of the computing components of the first SoC 510, the bandwidth usage and available memory of the chiplets of the first SoC 510, and / or faults or errors in these components, or information indicating faults or errors.
[0072] In a further embodiment, when the second SoC520 operates as a backup SoC, the state information exposed in the second memory 525 may correspond to the state of each computing component of the second SoC520. In particular, these components may operate in a low-power state in which they are ready to take over the set of tasks being performed by the first SoC510. The state information may include whether the components are operating within a nominal temperature and other nominal ranges (e.g., available bandwidth, power, memory, etc.).
[0073] As described throughout this disclosure, the first SoC 510 and the second SoC 520 can switch between operating as primary SoCs and backup SoCs (for example, each time the system 500 is rebooted). For example, in a computing session following a session in which the first SoC 510 operated as the primary SoC and the second SoC 520 operated as the backup SoC, the second SoC 520 may assume the role of the primary SoC and the first SoC 510 may assume the role of the backup SoC. This process of switching roles between the two SoCs is intended to cause the hardware components of each SoC to degrade substantially evenly, thereby extending the overall lifespan of the computing system 500.
[0074] According to the embodiment, the first SoC510 may be powered by a first power source, and the second SoC520 may be powered by a second power source that is independent of or isolated from the first power source. For example, in an electric vehicle, the first power source may comprise a battery pack used to propel the vehicle's electric motor, and the second power source may comprise an auxiliary power source for the vehicle (e.g., a 12-volt battery). In other implementations, the first and second power sources may comprise other types of power sources, such as dedicated batteries for each SoC510, 520, or other power sources that are electrically isolated from or otherwise independent of each other.
[0075] The aim is to improve the safety level (e.g., ASIL rating) of the computing system 500 and the overall autonomous driving system of the vehicle by providing an mSoC configuration for the computing system 500. As described herein, the autonomous driving system may include any number of dual SoC configurations, each capable of performing a set of autonomous driving tasks. In this configuration, the backup SoC dynamically monitors the health of the primary SoC according to a set of functional safety operations, so that when a failure, malfunction, or other error is detected, the backup SoC can immediately power up its components and take over the set of tasks from the primary SoC.
[0076] Functional Safety Accounting Figure 6 is a block diagram showing a performance network and a FuSa network for performing health monitoring, error correction, and system degradation, as illustrated in the examples described herein. In various examples, the FuSa CPU 600 may be located on the central chiplet 300 of the SoC, or on each central chiplet 300 of the mSoC500, as described herein. Each FuSa CPU 600 may execute a FuSa program 602, which may correspond to a FuSa program 338 as illustrated and described with respect to Figure 3. As described herein, execution of FuSa program 602 can cause the FuSa CPU 600 to execute the primary and backup SoC monitoring tasks described with respect to Figure 5, and the FuSa workload in the FuSa pipeline 420 for comparison and verification of independent pipeline outputs described with respect to Figure 4.
[0077] Furthermore, in the example shown in Figure 6, multiple chiplets of the SoC can communicate with each other via a high-bandwidth performance network that includes sets of interconnects (e.g., interconnects 610 and 660) and network hubs (e.g., network hubs 615, 635, and 665). As described herein, the multiple chiplets may include the sensor data input chiplet 310, the central chiplet 300, and the workload processing chiplet 320 in Figure 3, which are represented in Figure 6 by chiplet A 605, chiplet B 655, and any number of additional chiplets (not shown). Furthermore, the cache memories 625, 675 shown in Figure 6 may represent the cache memories associated with the multiple chiplets and / or the cache memory 315 of the central chiplet 300, as shown and described with respect to Figure 3.
[0078] In various examples, raw sensor data, processed sensor data, and various communications between chiplets A 605, B 655, and FuSa CPU 600 can be transmitted over a high-bandwidth performance network including interconnects 610, 660, network hubs 615, 635, 665, and caches 625, 675. For example, if chiplet A 605 contains a sensor data input chiplet, chiplet A 605 can acquire sensor data from various sensors in the vehicle and transmit the sensor data to cache 625 via interconnect 610 and network hub 615. In this example, if chiplet B 655 contains a workload processing chiplet, chiplet B 655 can acquire sensor data from cache 625 via network hubs 615, 635, 665, and interconnect 660 and perform its respective inference workload based on the sensor data.
[0079] In certain implementations, the FuSa CPU 600 can communicate with a high-bandwidth performance network via a performance network-on-chip (NoC) 607 coupled to a network hub 635 through the execution of the FuSa program 602. These communications may include, for example, obtaining output data from an independent pipeline to perform the comparison and verification steps described herein. Communication via the high-bandwidth performance network may further include communication to access the shared memories 515, 525 of each SoC 510, 520 within a plurality of SoCs 500, including a primary SoC and a backup SoC. In such an example, the FuSa CPU 600 of each SoC 510, 520 accesses each other's shared memories 515, 525 to determine whether any fault, failure, or other error has occurred. As described herein, if the backup SoC detects a fault, failure, or error, the backup SoC takes over the primary SoC task (e.g., inference, scene understanding, vehicle control tasks, etc.).
[0080] In some embodiments, interconnects 610 and 660 are used as high-bandwidth data paths for general data purposes to cache memories 625 and 675, while health control modules 620 and 670 and FuSa accounting hubs 630, 640, and 680 are used as high-reliability data paths for transmitting functional safety and scheduler information to the SoC's shared memory. NoC and network interface units (NIUs) on chiplets A 605 and B 655 can be configured to generate error correction code (ECC) data on both the high-bandwidth and high-reliability data paths. Each corresponding NIU on each pairing die has the same ECC configuration, which generates and checks the ECC data to ensure end-to-end error correction coverage.
[0081] According to various embodiments, the FuSa CPU 600 communicates via a FuSa network comprising FuSa accounting hubs 630, 640, 680 and health control modules 620, 670 through a FuSa NoC 609. As provided herein, the FuSa network facilitates communication monitoring and error correction coding techniques. As shown in Figure 6, the FuSa accounting hubs 630, 640, 680 can monitor communications transmitted through each network hub 615, 635, 665 of the high-bandwidth network. Each of the chiplets A 605 and B 655 may communicate with or include health control modules 620, 670, which can transmit ECC data, workload start and end communications, and scheduling information.
[0082] In the case of FuSa network data paths, the NIU can transmit functional safety and scheduler information via health control modules 620 and 670 in two redundant transactions, with the second transaction ordering the bits in reverse order of the first transaction (for example, from bit 31 to 0 on a 32-bit bus). Furthermore, if an error is detected during data transfer between chiplet A 605 and chiplet B 655 over the highly reliable FuSa network, the NIU can reduce the transmission rate to improve reliability.
[0083] In some examples, certain processors of chiplet A 605, chiplet B 655, and / or FuSa CPU 600 may include a transient-tolerant CPU core for executing scheduling program 342 in Figure 3, which schedules workloads belonging to the rapid response program 330, application program 335, thermal management program 337, and / or FuSa program 602. The transient-tolerant CPU core is designed to withstand and recover from transient failures caused by environmental factors such as cosmic rays, power surges, and electromagnetic interference. These failures can cause the CPU to malfunction or produce inaccurate results, potentially leading to system failure or security vulnerabilities. To address these issues, the transient-tolerant CPU core may include various hardware-based fault detection and recovery mechanisms, such as redundant execution units, error correction code (ECC) memory, and register duplication. These mechanisms can detect and correct errors in real time, ensuring that the CPU continues to function correctly even in the presence of transient failures. Furthermore, transient-tolerant CPU cores may include various software-based fault tolerance techniques, such as checkpointing and rollback, to further enhance system reliability and resilience.
[0084] In some embodiments, the health control modules 620, 670 and the FuSa accounting hubs 630, 640, 680 can detect and correct errors in real time, ensuring that the CPU continues to function correctly even in the presence of transient failures. For example, the workload processing chiplet A and the central chiplet can perform error correction checks to verify that processed data has been transmitted and stored in the cache memories 625, 675 completely uncorrupted. For example, for each processed data communication, the workload processing chiplet can generate an error correction code (ECC) using the processed data and transmit the ECC to the central chiplet. The data itself is transmitted along a high-bandwidth performance network between chiplets, while the ECC is transmitted along a high-reliability FuSa network via the FuSa accounting hubs 630, 640, 680. Upon receiving processed data, the central chiplet can generate its own ECC using the processed data, and the FuSa CPU 600 can perform a functional safety call within the central chiplet mailbox to compare the two ECCs and ensure they match, thereby verifying that the data was transmitted correctly.
[0085] As illustrated in the examples provided herein, the FuSa CPU 600 and FuSa program 602 may further monitor communications in the performance network and reliability network for evidence that the system is experiencing system overload, network latency, low bandwidth, and / or overheating. Upon detecting these problems, FuSa program 602 may take any number of mitigation measures, such as switching the primary and backup roles of the SoC in the deployment of the mSoC 500, or initiating system degradation in the manner described throughout this disclosure.
[0086] methodology Figures 7 and 8 are flowcharts illustrating exemplary methods for adapting feasible elements based on safety assessments, using various examples. In the following description of the methods in Figures 7 and 8, reference symbols representing certain feature parts, as described in relation to Figures 1 to 6, may be referenced. Furthermore, the steps described with reference to the flowcharts in Figures 7 and 8 may be performed by the computing system 100, the workload processing chiplet 320 and central chiplet 300 of the SoC 200, and / or the mSoC 500, as illustrated and described in relation to Figures 1 to 6. Moreover, certain steps described with reference to the flowcharts in Figures 7 and 8 may be performed before any other step, simultaneously with any other step, or after any other step, and do not need to be performed in the respective illustrated order. In further implementations, the “computing system” described in relation to Figures 7 and 8 can refer to any computing system and is not limited to onboard computing systems of autonomous or semi-autonomous vehicles. For example, computing systems may include robotic systems, aircraft, marine vehicles, autonomous agricultural equipment, server farms or data centers, personal computing devices, etc.
[0087] Referring to Figure 7, in block 700, the computing system can acquire sensor data from the sensor system. As provided herein, the sensor system may include one or more sensor types such as a LiDAR sensor, an image sensor or camera, a radar sensor, an ultrasonic sensor, a microphone, or a proximity sensor. In further implementations, the sensor system may be installed on an autonomous or semi-autonomous vehicle and may provide a continuous sensor view of the vehicle's surrounding environment. In block 705, based on the sensor data, the computing system may schedule the execution of the executable element 405 (e.g., via a scheduling program 342) according to the software structure 400. As provided herein, the software structure 400 may comprise all the software required to perform the responsive and / or application functions described above, which can be combined to perform all the sensor data acquisition, perception, object detection and classification, scene understanding, ML inference, motion prediction, motion planning, and / or vehicle control tasks required to autonomously operate the vehicle along a travel path.
[0088] In various embodiments, the software structure 400 can define executable elements 405 and connections between them, where relationships or dependencies exist. For example, the output of a first executable element (e.g., an object detection algorithm that identifies any object of interest in sensor fusion data) may be included as input to a second executable element (e.g., an object classification algorithm that classifies each of the detected objects), and its output may be included as input to a third executable element (e.g., a motion prediction algorithm tasked with predicting the motion of each dynamic object classified by the second executable element), and so on. As provided herein, these executable elements can be executed according to a dynamic scheduling and reservation table 350 that identifies when the workload corresponding to each executable element is available (e.g., when the dependency information for the workload has been met or otherwise resolved).
[0089] As provided herein, each connection or subset of connections between executable elements 405 within the software structure 400 may be associated with a safety rating (e.g., an ASIL rating). More critical connections between executable elements may be associated with a higher safety rating than less critical connections. For example, connections between executable element processing data in the forward direction of vehicle movement may be associated with a higher safety rating than connections between executable element processing data in the rear of the vehicle. As another example, connections between executable elements that include the detection of pedestrians or other vulnerable road users (VRUs) may have a higher safety rating than connections between executable elements that include the detection of fire hydrants, curb colors, parking meters, etc.
[0090] In block 710, while the executable element 405 in the software structure 400 is running, the computing system can detect degradation events within the computing system. As provided herein, degradation events can correspond to system overloads, such as overloading of computational tasks or executable elements necessary to safely navigate a given travel path. These scenarios may occur when a vehicle enters a high-process-load area, such as a densely populated urban environment (e.g., in extreme cases, pedestrians, bicycle lanes, external vehicles, traffic signs, etc.) and / or an area with complex traffic and right-of-way rules. In further examples, degradation events can correspond to one or more sensor failures, such as a camera or LIDAR sensor failing and forcing the computing system to rely on other forms of sensor data.
[0091] As another example, a degradation event may correspond to a computing system overheating or beginning to overheat (e.g., due to increased processing, hardware wear, hardware failure, a surge in heat, etc.). In such an example, the thermal management program 337 may perform a set of thermal mitigation tasks, such as enabling heat sinks, cooling fans, radiators, and / or switching between the primary SoC and backup SoC in an embodiment of the mSoC500 computing system. In the case of a surge in heat, the temperature rise can reduce the resistance of wires and other computer components, which can further increase the temperature in a critical feedback cycle. In such a scenario, the FuSa program 338 may detect the surge in heat (e.g., by monitoring the performance network) and initiate degradation in the manner described herein. For example, if the level of heat generated after thermal mitigation measures still exceeds the critical temperature, in block 715, the FuSa program 338 may cause the computing system to selectively degrade the execution of executable elements 405 in the software structure 400 based on a safety assessment of the executable elements in the software structure 400 and / or the connections between the executable elements.
[0092] As described herein, this degradation may include reducing the frequency with which an executable element is executed and which selected devices (e.g., a particular workload processing chiplet) have a relatively low safety rating for that executable element or executable element connection (e.g., ignoring every other image or sensor data iteration), reducing the frequency with which a particular computation task or executable element is executed, ignoring data from certain non-critical sensors, or temporarily preventing the execution of certain executable elements. As further described herein, this degradation may be implemented by the scheduling program 342 using a reservation table 350, which can indicate when certain workloads are executable elements. For example, the scheduling program 342 may flush certain workloads in an unordered buffer of the reservation table 350 without executing them (e.g., data corresponding to the workload may be transferred to an HBM chiplet without being processed by the workload processing chiplet 320), or the reservation table 350 may selectively remove certain workloads associated with executable elements having low safety rating connections. Such methods may be incorporated into the autonomous vehicle computing system to prevent system failure, malfunction, or increased wear due to excessive heat generation, and it is intended that they may provide additional redundancy to the placement of thermal management programs 337 and mSoC500, which may increase the overall ASIL rating of the entire computing system.
[0093] Figure 8 is another flowchart illustrating how executable elements are adapted based on safety assessments, using various examples. Referring to Figure 8, in block 800, the FuSa program 338 of the computing system can detect a degradation event while the computing system executes executable elements 405 in the software structure 400 according to a schedule. In block 805, the computing system can then determine an operating scenario to infer whether the operating scenario is associated with a degradation event. As provided herein, the connections between executable elements and / or executable elements 405 in the software structure 400 can be associated with safety assessments (e.g., ASIL assessments) which can be dynamically adjusted or configured based on the operating scenario.
[0094] In block 810, the computing system can configure or adjust safety assessments of executable elements and / or connections to executable elements within the software structure 400. For example, a driving scenario may include transitions in the driving environment (e.g., from a highway where executable elements processing forward sensor data are preferred to an urban environment where executable elements processing closer proximity data are more important). In a further example, a degradation trigger for initiating degradation may include the changing driving scenario itself (as opposed to, for example, overheating of the computing system). In such an example, an immediate program (which performs inference operations) can trigger the computing system's scheduling program 342 to preemptively initiate the degradation process described herein.
[0095] In various implementations, in block 815, the computing system can selectively degrade the execution of the software structure 400 based on a reconfigured or adjusted safety assessment of the connections between the executable elements 405. In certain examples, the safety assessment of the executable elements and / or the connections between executable elements may be hierarchically adaptable or adjustable. For example, when the system performs a "level 1" degradation, certain executable connections that may not be essential given the operating scenario may still have a high safety assessment.
[0096] In such an example, in decision block 820, the FuSa program 338 of the computing system can determine whether the degradation event has been resolved. If not, the FuSa program 338 can reconfigure the safety assessment and instruct the scheduling program 342 to hierarchically degrade additional viable elements or viable element connections (e.g., "level 2" degradation) until the degradation event is resolved (e.g., the system is cooled below the critical temperature threshold). If the degradation event has been resolved (e.g., through a temperature drop and / or a change in the operating scenario), in block 825, the scheduling program may selectively and / or hierarchically reverse the degradation accordingly.
[0097] It is intended that the examples described herein be extended to the individual elements and concepts described herein, independently of other concepts, ideas, or systems, and that combinations of elements described in any part of this application be included as examples. While examples are described in detail herein with reference to the accompanying drawings, it should be understood that the concepts are not limited to those exact examples. Therefore, many modifications and variations will be apparent to those skilled in the art. Accordingly, the scope of the concepts is intended to be defined by the following claims and their equivalents. Furthermore, specific features described individually or as part of an example are intended to be combined with other features or parts of other examples described individually, even if other features and examples do not refer to those specific features.
Claims
1. A computing system, A sensor data input chiplet for acquiring sensor data from a sensor system, A set of workload processing chiplets, A computing system comprising: a central chiplet having shared memory including a scheduling program for scheduling a set of executable elements to be executed at least partially based on the sensor data, wherein the set of executable elements is contained in a software structure for execution by the set of workload processing chiplets, and each executable element within the software structure and each connection between executable elements is associated with a safety evaluation to facilitate the degradation of the execution of the software structure.
2. The computing system according to claim 1, wherein the safety evaluation of each executable element and each connection between the executable elements includes an Automotive Safety Level (ASIL) evaluation.
3. The computing system according to claim 1, wherein the central chiplet includes a functional safety (FuSa) program that (i) monitors communications corresponding to the execution of the executable elements by the set of workload processing chiplets, and (ii) triggers the degradation when it detects data corresponding to system overload, processing delay, or overheating.
4. The computing system according to claim 3, wherein, in response to the FuSa program triggering the degradation, the scheduling program performs the degradation based on the safety evaluation of each executable element and each connection between the executable elements.
5. The computing system according to claim 4, wherein the degradation by the scheduling program includes reducing the execution frequency of a selected set of executable elements in the software structure.
6. The computing system according to claim 5, wherein the scheduling program reduces the execution frequency of the selected set of executable elements by using a reservation table that identifies when the selected set of executable elements is ready for execution.
7. The computing system according to claim 3, wherein the FuSa program triggers the degradation by performing at least one of the following: (i) morphing the software structure to change one or more connections between the executable elements, or (ii) truncating the software structure to exclude one or more executable elements from execution.
8. The computing system according to claim 1, wherein the sensor data input chiplet, the set of workload processing chiplets, and the central chiplet are included in a Universal Chiplet Interconnect Express (UCIe) system-on-chip (SoC) arrangement.
9. The computing system according to claim 1, comprising an on-board computer for an autonomous vehicle.
10. A non-temporary computer-readable medium for storing instructions, wherein when an instruction is executed by one or more processors of a computing system, the computing system... Using the sensor data input chiplet, acquire sensor data from the sensor system, Scheduling a set of executable elements to be executed at least partially on the sensor data, via the execution of a scheduling program on a central chiplet having shared memory, wherein the set of executable elements is included in a software structure by a set of workload processing chiplets, and each executable element and each connection between executable elements in the software structure is associated with a safety assessment to facilitate degradation of the execution of the software structure, in a non-temporary computer-readable medium.
11. The non-temporary computer-readable medium according to claim 10, wherein the safety evaluation of each executable element and each connection between said executable elements includes an Automotive Safety Level (ASIL) evaluation.
12. The non-temporary computer-readable medium according to claim 10, wherein the central chiplet includes a functional safety (FuSa) program that (i) monitors communications corresponding to the execution of the executable element by the set of workload processing chiplets, and (ii) triggers the degradation when it detects data corresponding to system overload, processing delay, or overheating.
13. The non-temporary computer-readable medium according to claim 12, wherein, in response to the FuSa program triggering the degradation, the scheduling program performs the degradation based on the safety evaluation of each executable element and each connection between the executable elements.
14. The non-temporary computer-readable medium according to claim 13, wherein the degradation by the scheduling program includes reducing the execution frequency of a selected set of executable elements in the software structure.
15. The non-temporary computer-readable medium according to claim 14, wherein the scheduling program reduces the execution frequency of the selected set of executable elements by using a reservation table that identifies when the selected set of executable elements is ready for execution.
16. The non-temporary computer-readable medium according to claim 12, wherein the FuSa program triggers the degradation by performing at least one of the following: (i) morphing the software structure to change one or more connections between the executable elements, or (ii) truncating the software structure to exclude one or more executable elements from execution.
17. The non-temporary computer-readable medium according to claim 10, wherein the sensor data input chiplet, the set of workload processing chiplets, and the central chiplet are included in a Universal Chiplet Interconnect Express (UCIe) system-on-chip (SoC) arrangement.
18. The computing system comprises an on-board computer for an autonomous vehicle, the non-temporary computer-readable medium according to claim 10.
19. A computer implementation method for adapting executable elements, wherein the method is executed by one or more processors. Using the sensor data input chiplet, acquire sensor data from the sensor system, A computer implementation method comprising scheduling a set of executable elements to be executed at least in part on the sensor data, via the execution of a scheduling program on a central chiplet having shared memory, wherein the set of executable elements is included in a software structure by a set of workload processing chiplets, and each executable element and each connection between executable elements in the set of executable elements is associated with a safety evaluation to facilitate the degradation of the execution of the software structure.
20. The method according to claim 19, wherein the safety evaluation of each executable element and each connection between the executable elements includes an Automotive Safety Level (ASIL) evaluation.