Vehicle regulation chip safety management method and system based on core particle architecture and vehicle
By obtaining fault information in the automotive specification chip of the core-particle architecture and determining the fault processing module according to the preset strategy for processing, the problem of redundancy and poor scalability of the automotive specification chip functional safety management system in the prior art is solved, and the effect of simplifying signal routing, improving anti-interference and enhancing scalability is achieved.
Patent Information
- Application Number
- CN202411781290.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-05
- Publication Date
- 2025-05-06
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing automotive chip functional safety management system based on the core-particle architecture has problems such as redundancy, complicated signal routing, poor scalability and susceptibility to interference.
A method of safety management of automotive chips based on core-particle architecture is proposed. When any functional module in the core-particle architecture fails, fault information is obtained, and the corresponding fault processing module is determined according to the preset fault processing allocation strategy, and the fault information is sent to the fault processing module for fault processing to be processed by the fault processing module. This method uses differential lines to transmit signals and determines the corresponding fault handling module for each functional module that fails.
It effectively simplifies signal routing, improves signal anti-interference, enhances the scalability of the automotive specification system, and thus improves the overall performance of the automotive specification system.
Smart Images

Figure CN119938372A_ABST
Abstract
Description
Technical Field
[0001] The present application belongs to the field of vehicle technology, and in particular, relates to a method, system and vehicle for automotive chip safety management based on a chiplet architecture. Background Art
[0002] With the development of the automotive industry, users' functional requirements for vehicles are also increasing. To meet these requirements, the existing automotive chips based on the core-grain architecture integrate multiple functional modules and equip each functional module with a safety management system. In the process of implementing the present invention, the inventors found that the functional safety management system of the existing automotive chips based on the core-grain architecture is redundant, the signal routing is complicated, the scalability is poor, and it is easily interfered. Summary of the invention
[0003] The present application aims to solve at least one of the technical problems existing in the prior art. To this end, the present application proposes a chip safety management method, system and vehicle based on a chip architecture, which can solve the problems of redundant functional safety management systems, complicated signal routing, poor scalability and susceptibility to interference of existing chip architecture chip functions. To achieve the above purpose, the embodiments of the present application provide the following technical solutions:
[0004] In a first aspect, an embodiment of the present application provides a vehicle-standard rail chip safety management method based on a chiplet architecture, comprising:
[0005] When any functional module in the core grain architecture fails, the failure information of the functional module is obtained;
[0006] Determine the corresponding fault handling module according to the preset fault handling allocation strategy;
[0007] The fault information of the functional module is sent to the fault processing module, so that the fault processing module performs fault processing according to the fault information.
[0008] In some embodiments, the above fault information is generated by a functional module where a fault occurs, and the functional module where a fault occurs includes a sub-functional system or a chip.
[0009] In some embodiments, the above fault information is transmitted via a differential signal.
[0010] In some embodiments, the above-mentioned fault information includes the fault type and identification information of the functional module where the fault occurs.
[0011] In some embodiments, the fault processing module includes at least one of a security management system, a CPU system, a startup system, and an off-chip MCU system;
[0012] The above-mentioned fault handling allocation strategy includes at least one of a fault level, a safety level definition and a safety management method.
[0013] In some embodiments, the above-mentioned determining the corresponding fault processing module according to the preset fault processing allocation strategy includes:
[0014] When the fault processing module is normal, the fault processing module processes the faulty functional module according to the fault information.
[0015] In some embodiments, the above-mentioned determining the corresponding fault processing module according to the preset fault processing allocation strategy includes:
[0016] When the safety management system fails, the CPU system and the startup system process the failed functional module based on the fault information.
[0017] In some embodiments, it also includes:
[0018] When the safety management system fails or is shut down and the off-chip MCU system is normal, the off-chip MCU system processes the failed functional module according to the fault information.
[0019] In some embodiments, the fault processing module performs fault processing according to the fault information, including:
[0020] The fault processing module reads the state information of the register of the functional module where the fault occurs, and performs fault processing according to the fault information and the register state information.
[0021] In a second aspect, an embodiment of the present application provides a vehicle-track chip safety management system based on a chiplet architecture, which includes:
[0022] A fault information acquisition module is used to acquire fault information of a functional module when a fault occurs in any functional module in the core grain architecture;
[0023] An allocation module, used to determine a corresponding fault handling module according to a preset fault handling allocation strategy;
[0024] The sending module is used to send the fault information of the functional module to the fault processing module, so that the fault processing module performs fault processing according to the fault information.
[0025] In a third aspect, an embodiment of the present application provides a vehicle, characterized in that it includes the vehicle-track chip safety management system based on the chiplet architecture described in the second aspect.
[0026] The technical solution provided in the embodiment of the present application obtains the fault information of the functional module when any functional module in the core-grain architecture fails, and then determines the corresponding fault processing module according to the preset fault processing allocation strategy, and sends the fault information of the functional module to the fault processing module, so that the fault processing module performs fault processing according to the fault information. Because the above-mentioned automotive chip safety management method based on the core-grain architecture adopts differential line transmission signal in the process of managing the functional modules with faults, and determines the corresponding fault processing module for each functional module with faults, it can not only effectively simplify the signal routing and improve the anti-interference ability of the signal, but also enhance the scalability of the automotive system, thereby improving the overall performance of the automotive system.
[0027] Additional aspects and advantages of the present application will be given in part in the description below, and in part will become apparent from the description below, or will be learned through the practice of the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0028] The above and / or additional aspects and advantages of the present application will become apparent and easily understood from the description of the embodiments in conjunction with the following drawings, in which:
[0029] Figure 1 A schematic diagram of a process flow of a method for managing automotive chip safety based on a chiplet architecture in an embodiment of the present application;
[0030] Figure 2 A schematic diagram of an on-board chip system for autonomous driving based on a chiplet architecture in an embodiment of the present application;
[0031] Figure 3 This is a schematic diagram of a vehicle-regulatory chip based on a bus architecture in an embodiment of the present application;
[0032] Figure 4 This is a schematic diagram of another automotive chip safety management method based on the chiplet architecture in an embodiment of the present application;
[0033] Figure 5 This is a schematic diagram of another automotive chip safety management method based on the chiplet architecture in an embodiment of the present application;
[0034] Figure 6 This is a schematic diagram of another automotive chip safety management method based on the chiplet architecture in an embodiment of the present application;
[0035] Figure 7 This is a schematic diagram of another automotive chip safety management method based on the chiplet architecture in an embodiment of the present application;
[0036] Figure 8 This is a schematic diagram of another automotive chip safety management method based on the chiplet architecture in an embodiment of the present application;
[0037] Fig. 9 This is a schematic diagram of another automotive chip safety management method based on the chiplet architecture in an embodiment of the present application;
[0038] Fig.10 This is a schematic diagram of another automotive chip safety management method based on the chiplet architecture in an embodiment of the present application;
[0039] Fig.11 This is a schematic diagram of another automotive chip safety management method based on the chiplet architecture in an embodiment of the present application;
[0040] Fig.12 This is a schematic diagram of another automotive chip safety management method based on the chiplet architecture in an embodiment of the present application;
[0041] Fig.13 This is a schematic diagram of another automotive chip safety management method based on the chiplet architecture in an embodiment of the present application;
[0042] Fig.14 This is a schematic diagram of the structure of a vehicle-track chip safety management system based on a chiplet architecture in an embodiment of the present application;
[0043] Fig.15 A schematic diagram of the structure of a vehicle provided in an embodiment of the present application. DETAILED DESCRIPTION
[0044] The embodiments of the present application are described in detail below, and examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present application, and cannot be understood as limiting the present application.
[0045] The terms "first", "second", etc. in the specification and claims of the present application are used to distinguish similar objects, and are not used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable under appropriate circumstances, so that the embodiments of the present application can be implemented in an order other than those illustrated or described here, and the objects distinguished by "first", "second", etc. are generally of one type, and the number of objects is not limited. For example, the first object can be one or more. In addition, "and / or" in the specification and claims represents at least one of the connected objects, and the character " / " generally indicates that the objects associated with each other are in an "or" relationship.
[0046] In the description of the present application, “plurality” means two or more.
[0047] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "illustrative embodiments", "examples", "specific examples", or "some examples" means that the specific features, structures, materials, or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic representation of the above terms does not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described may be combined in any one or more embodiments or examples in a suitable manner.
[0048] Because the existing automotive chips based on the chiplet architecture do not have particularly clear functional requirements in the early stages of design, when you want to add or delete a functional module later, you need to re-layout and re-route the input signal lines of the functional safety management system and reallocate the management method. This process is extremely complicated and reduces the scalability of automotive chips.
[0049] In addition, in order to meet the automotive-grade safety level requirements, the above-mentioned automotive chips must have fault detection and fault handling mechanisms in their built-in subsystems, which leads to redundancy in the functional safety management system. The detected fault types also need to be connected to the functional safety management system through complex signal routing for further processing. These routings are not only complicated, but also easily affected by external environmental factors such as temperature, power supply, production process, etc., thus affecting the reliability and safety of the entire automotive chip system.
[0050] Therefore, the embodiment of the present application provides a method for managing automotive chip safety based on the core-grain architecture, which obtains the fault information of the functional module when any functional module in the core-grain architecture fails, and then determines the corresponding fault processing module according to the preset fault processing allocation strategy, and sends the fault information of the functional module to the fault processing module, so that the fault processing module performs fault processing according to the fault information. Because the above-mentioned method for managing automotive chip safety based on the core-grain architecture adopts differential line transmission signals in the process of managing the functional modules that have failed, and determines the corresponding fault processing module for each functional module that has failed, it can not only effectively simplify signal routing and improve the anti-interference ability of the signal, but also enhance the scalability of the automotive system, thereby improving the overall performance of the automotive system. Figure 1 Schematic diagram of a process of a vehicle-grade chip safety management method based on a chiplet architecture in an embodiment of the present application, such as Figure 1 As shown, the following steps are included:
[0051] Step 101: when any functional module in the core-grain architecture fails, obtain the failure information of the functional module;
[0052] In the technical solution provided in the embodiment of the present application, the core architecture is the chiplet architecture, the core idea of which is to split a complex and feature-rich chip die into multiple small, independent corelets with specific functions, and these corelets can be combined together through advanced packaging technology to form a system-level chip; the functional modules therein are used to implement specific functions, and multiple functional modules work together to realize the overall function of the system, and each functional module can correspond to a corelet, such as the corelet with the above-mentioned specific function.
[0053] In an embodiment of the present application, the above-mentioned functional module includes a sub-functional system or chip, wherein the sub-functional system covers multiple aspects. Specifically, the above-mentioned sub-functional system may include a high-bandwidth memory / double data rate system, a direct memory access system, a system control, a peripheral system, an NPU system, a bus system, and a memory system, etc.; the above-mentioned chip includes an on-chip chip and an off-chip chip, wherein the on-chip chip is integrated on a car-regulatory chip of a core-grain architecture, including a data processing chip, such as a CPU chip, a GPU chip, etc., a computing power chip, such as an AI chip, and other chips, etc., wherein the data processing chip is used to process various types of data, for example, the CPU chip is used to execute program instructions, process data, and coordinate system resources; the GPU chip is used to process image data or graphic data; the computing power chip is used to provide computing power, for example, the AI chip is used to provide efficient computing power; other chips refer to chips used to complete specific functions, for example, encryption and decryption chips, audio processing chips, etc.; the off-chip chip can be connected to the car-regulatory chip of the core-grain architecture through a chip-to-chip interface (Die-to-Die, referred to as D2D) to complete the corresponding function.
[0054] For example, in an autonomous driving vehicle chip system based on a chiplet architecture, Figure 2As shown, it includes multiple sub-functional systems, chips, internal interconnection bus matrices, and system management and functional safety management systems, etc. The sub-functional systems include high-bandwidth memory / double data rate system, direct memory access system, system control, peripheral system, NPU system, memory system, high-speed IO system and other systems, etc. The above-mentioned high-bandwidth memory / double data rate (HighBandwidth Memory, HBM for short, Double Data Rate Memory, DDR for short) system, namely HBM / DDR system, is used to ensure high-speed transmission of data between memory and processing unit; Direct Memory Access (DMA for short) system allows hardware devices to directly access main memory without CPU intervention, thereby improving data transmission efficiency and system response speed; System control is used to coordinate the operation of the entire system, ensure the coordinated work between various sub-functional systems, and respond to and process external inputs; Peripheral system is used to provide rich interfaces and extended functions, so that the on-board chip system can communicate and interact with external devices, such as sensors, cameras, etc., so as to realize the perception and monitoring of the vehicle's surrounding environment; Neural Processing Unit (Neural Processing Unit) The main function of the vehicle chip system is to process complex neural network computing tasks, such as image recognition and path planning, and provide powerful intelligent support for autonomous driving. The memory system is used to store and process large amounts of data and information, including map data, vehicle status information, etc., providing reliable data support for autonomous driving decisions. The high-speed IO system is used to provide high-speed data input and output channels to ensure smooth and efficient data exchange between the vehicle chip system and other systems or devices. Other systems, such as the power management system and the heat dissipation system, jointly provide a full range of guarantees for the stable operation of the vehicle chip system. The chips include CPU chips, GPU chips, AI chips and other chips. The internal interconnection bus matrix is the bridge between the above-mentioned sub-functional systems and chips, and is used to efficiently and reliably transmit data and instructions to ensure the coordinated work and efficient operation of the entire autonomous driving vehicle chip system. The system management and functional safety management system is used to monitor and manage the operating status of the entire system to ensure the safety and reliability of the system. Through the safety management system, CPU system, startup system and AXI Stream matrix, the above-mentioned AXI Stream matrix is a data stream processing architecture or technology based on the AXI4-Stream protocol. It uses the switching and scheduling capabilities of the matrix to achieve flexible switching and processing of high-speed data streams between multiple input and output channels. The above-mentioned AXI4-Stream protocol AXI4-Stream protocol is a standard protocol interface, mainly used for high-speed data stream transmission within the chip.In the embodiment of the present application, the above-mentioned fault information is formed due to the failure of the functional module, which includes the fault type and the identification information of the functional module where the failure occurs. The fault types include the following: a fault that can be repaired in transmission data / storage data, a fault that cannot be repaired in transmission data / storage data, an operation timeout fault, an illegal operation fault of a security register, an illegal operation fault of a memory, and other faults, wherein the fault that can be repaired in transmission data / storage data refers to a fault that can be repaired by technical means, such as using data redundancy, checksums or error correction algorithms to restore the correctness of the data; a fault that cannot be repaired in transmission data / storage data refers to a fault that causes a data error to be permanent and cannot be repaired by technical means; an operation timeout fault, when a functional module exceeds a preset time limit when performing an operation, it will be judged as an operation timeout fault; a security register illegal operation fault refers to a fault triggered by illegal operation of a security register, which is a special register used in the system to protect critical data and program security. When these registers are illegally operated, this type of fault will be triggered; a memory illegal operation fault refers to a fault triggered by illegal access or operation to the memory, such as out-of-bounds access to the memory.
[0055] In addition, in this application, users can also customize fault types to meet specific needs in different application scenarios. For example, users can define a "system abnormality" fault type to indicate that the system is abnormal or unstable; define a "system interruption" fault type to indicate that the system is forced to interrupt operation for some reason. This not only enhances the flexibility and scalability of the system, but also makes the description of fault information more accurate and comprehensive.
[0056] The identification information of the above-mentioned faulty functional module is unique and can be used to identify the functional module. In the present application, an 8-bit binary number can be used to represent the identification information of 256 functional modules so that each functional module has unique identification information. This representation method is both efficient and practical, and can not only meet the needs of the current system, but also provide ample space and possibilities for future expansion.
[0057] Step 102: Determine a corresponding fault processing module according to a preset fault processing allocation strategy;
[0058] On the basis of step 101, the preset fault handling allocation strategy is further used to determine the fault handling module responsible for handling the faulty functional module. Specifically, the fault type in the fault information of the functional module is used, combined with the preset fault handling allocation strategy, to determine a fault handling module responsible for handling the faulty functional module. The above-mentioned fault handling module is used to handle and solve the functional module that has a fault, and it includes at least one of a safety management system, a CPU system, a startup system and an off-chip MCU system. The safety management system is used to ensure the information security and functional integrity of the data; the CPU system is used to perform various computing tasks and control operations; the startup system is used to ensure that the computer or device can be successfully started from the shutdown state and enter the loading stage of the operating system or application; the off-chip MCU refers to a high-performance vehicle-mounted controller system with ASIL-D high functional safety, an integrated microprocessor core, memory, input and output interface and other functional modules. 。
[0059] The above-mentioned preset fault handling allocation strategy is used to determine the fault handling module responsible for handling the functional module with faults, which includes at least one of fault level, safety definition and safety management mode. The fault level is an assessment of the severity of the system fault, which is usually divided based on the degree of influence of the fault on the system function, performance and safety, etc. It can also be customized by the user. For different automotive chips, the fault level set will be different. Each type of fault type of the above-mentioned faulty functional module is set with a corresponding fault level, and each fault level is set with a corresponding fault handling module, which is responsible for handling the faulty functional module belonging to the fault level. When the fault handling module responsible for handling the faulty functional module is determined only based on the fault level, it is necessary to first determine the fault type of the faulty functional module, and then determine the fault level to which it belongs according to the fault type, and then determine the corresponding fault handling module according to the fault level. For example, if a peripheral system fails and its fault type is a fault that cannot be repaired when transmitting data / storing data, since the fault level corresponding to the fault type is fault level 1, the fault handling module 1 corresponding to fault level 1 is called to handle the above-mentioned faulty peripheral system.
[0060] The above security levels are used to reflect the system's emphasis on and specific requirements for the security of the functional modules. The security levels of the functional modules can be set according to conventional settings or customized by the user. The security levels of different functional modules will be different, and each security level is provided with a corresponding fault handling module. When determining the fault handling module based only on the security level of the functional module, it is necessary to first determine which specific functional module has failed based on the identification information of the functional module that has failed, and then determine the corresponding fault handling module based on the security level to which the functional module belongs. For example, if the direct access system fails and its security level is security level 2, the fault handling module 2 corresponding to security level 2 is called to handle the direct access system that has failed.
[0061] The above-mentioned security management method refers to the method or means used by the system to manage and monitor the fault handling module to handle the functional module that has a fault. The security management method of the functional module can be set according to the conventional settings or customized by the user. The security management methods of different functional modules will be different, and each security management method is provided with a corresponding fault handling module. When determining the fault handling module based only on the security management method of the functional module, it is necessary to first determine which specific functional module has a fault based on the identification information of the functional module that has a fault, and then determine the corresponding fault handling module based on the security management method to which the functional module belongs. For example, if the direct access system fails and its security management method is security management method 3, the fault handling module 3 corresponding to security management method 3 is called to handle the above-mentioned direct access system that has a fault.
[0062] In the embodiment of the present application, multiple of the above-mentioned fault levels, safety levels and safety management methods can also be used to determine the corresponding fault handling module to handle the functional module where the fault occurs, which can not only improve the pertinence and efficiency of fault handling, but also better ensure the overall safety and stability of the system. For example, by using the fault level and safety level at the same time, by combining the above-mentioned two fault allocation strategies, the corresponding fault handling module is determined to handle the functional module where the fault occurs.
[0063] Step 103: Send the fault information of the functional module to the fault processing module, so that the fault processing module performs fault processing according to the fault information.
[0064] After the fault processing module corresponding to the functional module with the fault is determined in step 102, the fault information of the above functional module is sent to the fault processing module, and then the fault processing module processes the functional module with the fault according to the above fault information. The above processing includes recording the fault details for subsequent analysis and tracking; trying to repair those recoverable transmission data or storage data to reduce data loss; issuing warnings to the driver in time to ensure that the driver can quickly perceive the abnormality of the system; intelligently adjusting system parameters according to the type and severity of the fault to reduce the impact of the fault on the overall performance of the system; when necessary, starting the emergency parking procedure to ensure the safety of personnel and vehicles, etc. These processing measures can deal with the faulty functional modules in an all-round way to ensure the stability and safety of the system.
[0065] As described above, this method can be used to safely manage automotive chips based on the coregrain architecture. Differential lines can be used to transmit signals, and a corresponding fault handling module can be determined for each functional module that fails. This can not only effectively simplify signal routing and improve signal anti-interference, but also enhance the scalability of the automotive system, thereby improving the overall performance of the automotive system.
[0066] Usually, during the signal transmission process, it is often inevitably interfered by various environmental factors, such as small fluctuations in the manufacturing process, changes in the working environment temperature, and instability of the power supply voltage. These interference factors will directly affect the accuracy and integrity of signal transmission, and may cause problems such as loss of transmission data and misjudgment. Therefore, in order to ensure high-quality and stable signal transmission, it is necessary to select the signal transmission method that best suits the current application scenario, such as differential signal transmission.
[0067] In some embodiments, the fault information of the above-mentioned faulty functional module can be transmitted via differential signals. The above-mentioned differential signal is a signal transmission method, which uses two wires, namely a differential pair, to transmit signals. Specifically, the differential signal is realized by sending signals of equal amplitude but opposite polarity on two wires, and the receiving end restores the original signal by comparing the difference between the two signals. By using differential signals to transmit the above-mentioned fault information, the anti-interference ability of the signal can be effectively enhanced, and the reliability and stability of signal transmission can be improved. In an embodiment of the present application, the functional module and the fault handling module can be connected based on the bus architecture to ensure that each sub-functional system or chip has fault detection capability. For example, for an on-board chip system for autonomous driving, when the functional module and the fault handling module are connected based on the bus architecture, the connection method is as follows: Figure 3As shown in the figure, through the information transmission channel of the bus architecture, each sub-functional system, chip and high-speed interface can collect the fault information and corresponding identification information in real time in the form of AXI Stream, that is, in the form of data stream, and obtain the fault queue. Then the above fault queue is transmitted to the collection, arbitration and distribution mechanism. The arbitration mechanism will evaluate and determine which fault signals need to be processed first. Then the distribution mechanism will distribute these fault signals to the corresponding target queues according to the arbitration results. Finally, these target queues will be sent to the fault processing module for further fault processing. Transmitting data in the form of queues in this way can optimize the data transmission process and ensure the integrity and accuracy of the data. For example, the fault type of the NPU system is stored in fault_type[9:0][1:0], and the identification information is stored in m_id[7:0][1:0], where [1:0] is a differential pair; the information stored in the above positions is collected in the form of AXI Stream to form a fault queue. Subsequently, the information in the fault queue is prioritized through the arbitration mechanism and further processed through the distribution mechanism. During the distribution process, the information will be placed in the corresponding position of the target queue, and after being distributed to the corresponding position of the target queue, the target queue Fault queue, which is processed by the corresponding fault processing module, such as the security management system; the above-mentioned fault queue refers to the queue used to store and manage the fault information generated by each functional module; the target queue refers to the fault queue information that needs to be processed by the fault processing module in the end: the collection mechanism is the basis of the entire fault handling process, which is responsible for collecting fault information and identification information from each functional module to ensure that all important fault data can be captured and recorded in time; the arbitration mechanism is used to sort and filter multiple fault information according to the preset priority strategy when they are generated at the same time, so as to ensure that the most important fault information can be processed first, thereby minimizing the system downtime and potential losses; the distribution mechanism is the above-mentioned preset fault distribution strategy, which can distribute the fault signal to the target queue according to the fault level, safety level and safety management method of the functional module, so as to make the fault handling process more efficient and orderly.
[0068] In an embodiment of the present application, the corresponding fault handling module can be determined according to a preset fault handling allocation strategy, and after the fault handling module is determined, further according to the register status information and fault information of the functional module where the fault occurs read by the above-mentioned fault handling module, corresponding measures are taken to handle the functional module where the fault occurs, so as to ensure the stable operation of the entire system. The above-mentioned register status information refers to the status information recorded by the register in the functional module where the fault occurs, for example, the status information of the security register of the functional module, which includes the type and level of the fault, whether the fault data is repairable, the number of times the fault data has been attempted to be repaired, and the specific address associated with the fault data. When determining the corresponding fault handling module according to the preset fault handling allocation strategy to handle the functional module where the fault occurs, the following situations are included:
[0069] In the first case, when all fault processing modules are working normally, each fault processing module works together to process the faulty functional module according to the register status information and fault information. Specifically, in the first case, the result of the fault processing module is determined according to the preset fault processing allocation strategy, as shown in the following table, including several cases:
[0070]
[0071] The first column in the above table is the type of fault handling module, and the first line is the fault type of the functional module where the fault occurs, including 9 fault types. Users can customize the above fault types. For example, fault types 1-5 are defined as repairable faults in transmission data / storage data, unrepairable faults in transmission data / storage data, operation timeout faults, illegal operation faults of security registers, and illegal operation faults of memory. The remaining fault types can be system abnormal faults, system interrupt faults, etc.
[0072] First, when a sub-function system fails and its failure type is failure type 1-5, the safety management system will handle the failed sub-function system. Figure 4As shown, the fault information and identification information of the sub-functional system are collected in the fault queue and transmitted to the AXI-Stream interconnect bus in the form of AXI Stream. After being processed by the arbitration mechanism and the distribution mechanism, they are distributed to the target queue. After the safety management system receives the fault information and identification information of the sub-functional system contained in the target queue, it determines which sub-functional system has failed by the identification information. Then, the safety management system initiates a read command to read the status information of the safety register of the sub-functional system. Finally, the safety management system processes the sub-functional system where the failure occurs based on the read status information and fault information of the above-mentioned safety register. For example, if the fault type of the sub-functional system is a fault that can be repaired by transmission data / storage data, the transmission data / storage data of the above-mentioned sub-functional system is repaired by data repair technology. At this time, the processing of the sub-functional system where the failure occurs must also comply with the requirements of the functional safety specification, that is, from the time when the sub-functional system detects a fault, the delay from the initiation of the fault to the receipt of the fault event by the safety management system is within 10ns.
[0073] The above-mentioned security management system includes two lock-step CPUs, security management and security memory, etc., among which there are usually two lock-step CPUs, namely dual lock-step CPUs, which is a CPU redundancy technology, namely, two identical processors are included in one chip, one as the main processor and the other as the slave processor. The two processors execute the same code and are strictly synchronized. The main processor can access the system memory and output instructions, while the slave processor continuously executes instructions on the bus; security management is used to monitor the operating status of the entire system, including processor execution, memory access and data transmission, etc., and its potential threats are monitored and warned in real time, and corresponding measures are taken when anomalies are found, such as isolating faults, restarting the system, etc., to ensure the security and stability of the system; security memory is a storage device used to store key data and instructions, which can adopt special security mechanisms, such as encryption, access control, etc., to prevent data from being illegally accessed or tampered with.
[0074] Second, when a chip fails and its failure type is failure type 1-5, the safety management system will handle the failed chip. Figure 5As shown, the faulty chip transmits information through the chip-to-chip interface. After the chip-to-chip interface receives the fault information and identification information of the chip, the information is collected in the fault queue and sent to the AXI Stream is transmitted to the AXI-Stream interconnect bus, and after being processed by the arbitration mechanism and the distribution mechanism, it is distributed to the target queue. After the safety management system receives the fault information and identification information of the on-chip chip contained in the target queue, it determines which on-chip chip has failed by the identification information. Then, the safety management system initiates a read command, and the chip-to-chip interface on the safety management system side transmits it to the chip-to-chip interface on the on-chip chip side to read the status information of the security register of the on-chip chip. Finally, the safety management system processes the on-chip chip that has failed based on the read status information and fault information of the above-mentioned security registers. For example, if the fault type of the on-chip chip is a fault that cannot be repaired for transmitting data / storing data, a warning is issued in time to the driver to ensure that the driver can quickly perceive the abnormal situation of the system. At this time, the processing of the on-chip chip that has failed must also comply with the requirements of the functional safety specification, that is, from the time when the on-chip chip detects a fault, the delay from the fault initiation to the receipt of the fault event by the safety management system is within 20ns. Since the chip-to-chip interface latency is within 5-6ns, the status information of the security registers of a failed on-chip chip can be quickly read, helping to improve the overall performance and security of the system.
[0075] Third, when an off-chip chip fails and its failure type is failure type 1-5, the safety management system will handle the failed off-chip chip. Figure 6As shown, the faulty off-chip chip transmits information through the high-speed IO interface. After receiving the fault information and identification information of the off-chip chip from the high-speed IO interface, the information is collected in the fault queue and transmitted to the AXI-Stream interconnect bus in the form of AXI Stream. After being processed by the arbitration mechanism and the distribution mechanism, it is distributed to the target queue. After the safety management system receives the fault information and identification information of the off-chip chip contained in the target queue, it determines which off-chip chip has failed by the identification information. Then, the safety management system initiates a read command, which is transmitted to the off-chip chip by the high-speed IO interface of the safety management system to read the status information of the safety register of the off-chip chip. Finally, the safety management system processes the faulty off-chip chip according to the read status information and fault information of the above-mentioned security register. For example, if the fault type of the off-chip chip is an illegal operation fault of the security register, a warning is issued to the driver in time, and the driver is granted a higher authority to access. At this time, the processing of the faulty off-chip chip must also comply with the requirements of the functional safety specification, that is, from the time when the off-chip chip detects a fault, the delay from the fault initiation to the receipt of the fault event by the safety management system is within 500ns. Since the delay of the high-speed IO interface is within 300-500ns, the status information of the security register of the faulty off-chip chip can be quickly read, which helps to improve the overall performance and security of the system.
[0076] Fourth, when a sub-function system fails and its failure type is failure type 6-9, such as abnormal type failure, system interruption failure, etc., the CPU system will handle the failed sub-function system. Specifically, Figure 7 As shown, the fault information and identification information of the sub-functional system are collected in the fault queue and transmitted to the AXI-Stream interconnect bus in the form of AXI Stream. After being processed by the arbitration mechanism and the distribution mechanism, they are distributed to the target queue. After the CPU system receives the fault information and identification information of the sub-functional system contained in the target queue, it determines which sub-functional system has failed by the identification information. Then, the CPU system initiates a read command to read the status information of the security register of the sub-functional system. Finally, the CPU system processes the sub-functional system that has failed based on the read status information and fault information of the above security register. For example, if the fault type of the sub-functional system is a fault that can be repaired by transmission data / storage data, the transmission data / storage data of the above sub-functional system is repaired by data repair technology. At this time, the processing of the sub-functional system that has failed must also comply with the requirements of the functional safety specification, that is, from the time when the sub-functional system detects a fault, the delay from the fault initiation to the CPU system receiving the fault event is within 10ns.
[0077] The above-mentioned CPU system includes two CPU clusters, interrupt management and memory, wherein the CPU cluster is a collection of a group of interconnected CPU cores or processors in a multi-processor system, and these cores or processors are interconnected in some way, for example, shared cache, shared bus or higher-level interconnection network, so as to achieve higher parallel processing capability and data sharing efficiency; interrupt management is an important function in the CPU system, which allows the CPU to respond to interrupt signals from the outside or inside while processing the current task, thereby switching to another task or executing a specific processing program; memory is a component in the CPU system used to store data and instructions.
[0078] Fifth, when a chip on the chip fails and its failure type is failure type 6-9, such as abnormal type failure, system interrupt failure, etc., the CPU system will handle the failed chip on the chip. Specifically, Figure 8 As shown, the faulty chip transmits information through the chip-to-chip interface. After the chip-to-chip interface receives the fault information and identification information of the chip, the information is collected in the fault queue and sent to the AXI The data is transmitted to the AXI-Stream interconnect bus in the form of a stream, and is distributed to the target queue after being processed by the arbitration mechanism and the distribution mechanism. After the CPU system receives the fault information and identification information of the on-chip chip contained in the target queue, it determines which on-chip chip has failed by the identification information. Then, the CPU system initiates a read command, and the chip-to-chip interface on the CPU system side transmits it to the chip-to-chip interface on the on-chip chip side to read the status information of the security register of the on-chip chip. Finally, the CPU system processes the on-chip chip that has failed based on the read status information and fault information of the above-mentioned security register. For example, if the fault type of the on-chip chip is a fault that cannot be repaired in transmitting data / storing data, a warning is issued to the driver in a timely manner to ensure that the driver can quickly perceive the abnormality of the system. At this time, the processing of the on-chip chip that has failed must also comply with the requirements of the functional safety specification, that is, from the time when the on-chip chip detects a fault, the delay from the fault initiation to the CPU system receiving the fault event is within 20ns. Since the chip-to-chip interface latency is within 5-6ns, the status information of the security registers of a failed on-chip chip can be quickly read, helping to improve the overall performance and security of the system.
[0079] In the second case, when the safety management system fails, the CPU system and the startup system process the faulty functional module according to the register status information and the fault information. Specifically, in the second case, the result of the fault processing module is determined according to the preset fault processing allocation strategy, as shown in the following table, including several cases:
[0080]
[0081] The first column in the above table is the type of fault handling module, and the first line is the fault type of the functional module where the fault occurs, including 9 fault types. Users can customize the above fault types. For example, fault types 1-5 are defined as repairable faults in transmission data / storage data, unrepairable faults in transmission data / storage data, operation timeout faults, illegal operation faults of security registers, and illegal operation faults of memory. The remaining fault types can be system abnormal faults, system interrupt faults, etc.
[0082] First, when a sub-function system fails and its failure type is failure type 1-5, the startup system will process the failed sub-function system. Fig. 9 As shown, the fault information and identification information of the sub-function system and the safety management system are collected in the fault queue and transmitted to the AXI-Stream interconnection bus in the form of AXI Stream. After being processed by the arbitration mechanism and the distribution mechanism, they are distributed to the target queue. After the startup system receives the fault information and identification information of the sub-function system and the safety management system contained in the target queue, it determines which sub-function system has failed by the identification information. Then, the startup system initiates a read command to read the status information of the safety register of the sub-function system and the safety management system. Finally, the startup system processes the sub-function system and the safety management system that have failed according to the status information and fault information of the above-mentioned safety registers. For example, if the fault types of the sub-function system and the safety management system are all faults that can be repaired by transmission data / storage data, the transmission data / storage data of the above-mentioned sub-function system and the safety management system are repaired by data repair technology. At this time, the processing of the sub-function system or safety management system that has failed must also comply with the requirements of the functional safety specification, that is, from the time when the sub-function system and the safety management system detect the fault, from the fault initiation to the start-up system receiving the fault event, the delay is within 10ns.
[0083] The above-mentioned startup system includes two microcontroller units, startup management and memory. The microcontroller unit (MCU) is used to perform various control tasks and data processing. When one of the MCUs fails, the other MCU takes over its tasks to improve the reliability and fault tolerance of the system; the startup management is used to control the entire startup process of the automotive chip, for example, initializing the system's MCU, memory, external devices and other components; the memory is a component in the startup system used to store data and programs.
[0084] Second, when a chip fails and its failure type is failure type 1-5, the startup system will process the failed chip. Fig.10As shown, the faulty chip transmits information through the chip-to-chip interface. When the chip-to-chip interface receives the fault information and identification information of the chip on the chip, and obtains the fault information and identification information of the safety management system, the above information is collected in the fault queue and sent to the AXI The data is transmitted to the AXI-Stream interconnect bus in the form of a stream, and is distributed to the target queue after being processed by the arbitration mechanism and the distribution mechanism. After the startup system receives the fault information and identification information of the on-chip chip and the safety management system contained in the target queue, it determines which on-chip chip has failed by using the identification information. Then, the startup system initiates a read command, which is transmitted to the chip-to-chip interface of the on-chip chip end through the chip-to-chip interface of the startup system end to read the status information of the safety register of the on-chip chip. On the other hand, the status information of the safety register of the safety management system is directly read. Finally, the startup system processes the on-chip chip and the safety management system that have failed based on the read status information and fault information of the above-mentioned safety registers. For example, if the fault type of the on-chip chip and the safety management system is a fault that cannot be repaired when transmitting data / storing data, a warning is issued to the driver in a timely manner to ensure that the driver can quickly perceive the abnormal situation of the system. At this time, the processing of the on-chip chip and the safety management system that have failed must also comply with the requirements of the functional safety specification, that is, from the time when the on-chip chip or the safety management system detects a fault, the delay from the fault initiation to the receipt of the fault event by the startup system is within 20ns. Since the chip-to-chip interface latency is within 5-6ns, the status information of the security registers of a failed on-chip chip can be quickly read, helping to improve the overall performance and security of the system.
[0085] Third, when an off-chip chip fails and its failure type is failure type 1-5, the startup process will be used to process the off-chip chip that failed. Fig.11As shown, the faulty off-chip chip transmits information through the high-speed IO interface. When the fault information and identification information of the on-chip chip are received from the high-speed IO interface, and the fault information and identification information of the safety management system are obtained, the above information is collected in the fault queue and sent to the AXI The information is transmitted to the AXI-Stream interconnect bus in the form of a stream, and is distributed to the target queue after being processed by the arbitration mechanism and the distribution mechanism. After the startup system receives the fault information and identification information of the off-chip chip and the safety management system contained in the target queue, it determines which off-chip chip has failed by the identification information. Then, the safety management system initiates a read command, which is transmitted to the off-chip chip through the high-speed IO interface of the startup system end on the one hand, and the status information of the safety register of the off-chip chip is read. On the other hand, the status information of the safety register of the safety management system is directly read. Finally, the startup system processes the off-chip chip and the safety management system that have failed based on the read status information and fault information of the above-mentioned safety registers. For example, if the fault types of the off-chip chip and the safety management system are both illegal operation faults of the safety register, a warning is issued to the driver in a timely manner, and the driver is granted higher authority to access. At this time, the processing of the off-chip chip and the safety management system that have failed must also comply with the requirements of the functional safety specification, that is, from the time when the off-chip chip or the safety management system detects a fault, the delay from the fault initiation to the start-up system receiving the fault event is within 500ns. Since the delay of the high-speed IO interface is within 300-500ns, the status information of the security register of the faulty off-chip chip can be quickly read, which helps to improve the overall performance and security of the system.
[0086] Fourth, when a sub-functional system or chip fails, and its failure type is failure type 6-9, such as an abnormal type failure, a system interruption failure, etc., the CPU system handles the failed sub-functional system, chip and security management system. The specific process can be referred to the above embodiment and will not be repeated here.
[0087] In the third case, when the safety management system fails or is shut down and the off-chip MCU system is normal, the off-chip MCU system processes the faulty functional module according to the register status information and fault information. Specifically, in the third case, the result of the fault processing module is determined according to the preset fault processing allocation strategy, as shown in the following table, including several cases:
[0088]
[0089] The first column in the above table is the type of fault handling module, and the first line is the fault type of the functional module where the fault occurs, including 9 fault types. Users can customize the above fault types. For example, fault types 1-5 are defined as repairable faults in transmission data / storage data, unrepairable faults in transmission data / storage data, operation timeout faults, illegal operation faults of security registers, and illegal operation faults of memory. The remaining fault types can be system abnormal faults, system interrupt faults, etc.
[0090] First, when a sub-function system fails and its failure type is failure type 1-5, the off-chip MCU system handles the failed sub-function system. Specifically, Fig.12 As shown, the fault information and identification information of the sub-function system and the safety management system are collected in the fault queue and transmitted to the AXI-Stream interconnection bus in the form of AXI Stream. After being processed by the arbitration mechanism and the distribution mechanism, they are distributed to the target queue. Then, the fault information and identification information of the sub-function system and the safety management system contained in the above target queue are sent to the off-chip MCU system through the high-speed IO interface and with the help of input / output pads. After the off-chip MCU system receives the fault information and identification information of the sub-function system and the safety management system contained in the target queue, it is determined which sub-function system has a fault through the identification information. Then, the off-chip MCU system initiates a read command, which is then transmitted via the high-speed IO interface to read the sub-function system and the safety management system. Finally, the off-chip MCU system processes the sub-functional system and safety management system that have failed based on the status information and fault information of the above-mentioned safety registers. For example, if the fault types of the sub-functional system and the safety management system are all faults that can be repaired by transmission data / storage data, the transmission data / storage data of the above-mentioned sub-functional system and safety management system are repaired by data repair technology. At this time, the processing of the sub-functional system and safety management system that have failed must also comply with the functional safety specification requirements, that is, from the time the sub-functional system or safety management system detects a fault, the delay from the fault initiation to the receipt of the fault event by the off-chip MCU system is within 500ns. The above-mentioned input / output pads
[0091] The above-mentioned off-chip MCU system includes a high-speed IO interface, a peripheral interface, security management and a lock-step CPU. The high-speed IO interface is a channel for data transmission between the MCU and external devices, and has the characteristics of high speed, versatility and programmability; the peripheral interface is a channel for connecting the MCU with external devices, such as sensors, actuators, memories, etc.; security management is used to protect the security and integrity of the off-chip MCU system; the system only includes a lock-step CPU to improve the reliability and fault tolerance of the system.
[0092] Second, when an on-chip chip fails and its failure type is failure type 1-5, the off-chip MCU system will handle the failed on-chip chip. Specifically, Fig.13 As shown, the faulty on-chip chip transmits information through the chip-to-chip interface. When the fault information and identification information of the on-chip chip are received from the chip-to-chip interface, and the fault information and identification information of the safety management system are obtained, the above information is collected in the fault queue and transmitted to the AXI-Stream interconnection bus in the form of AXI Stream. After being processed by the arbitration mechanism and the distribution mechanism, it is distributed to the target queue. Then, the fault information and identification information of the sub-functional system and the safety management system contained in the above target queue are sent to the off-chip MCU system through the high-speed IO interface and with the help of input / output pads. After the off-chip MCU system receives the fault information and identification information of the on-chip chip and the safety management system contained in the target queue, it determines which specific on-chip chip has failed through the identification information. Then, the off-chip MCU system initiates a read command, and the read command is transmitted via the high-speed IO interface to read the sub-functional system. and the status information of the safety registers of the safety management system. Finally, the startup system processes the faulty on-chip chip and safety management system based on the status information and fault information of the above-mentioned safety registers read. For example, if the fault type of the on-chip chip and the safety management system is an unrepairable fault in transmitting data / storing data, a warning is promptly issued to the driver to ensure that the driver can quickly perceive the abnormality of the system. At this time, the processing of the faulty on-chip chip and safety management system must also comply with the functional safety specification requirements, that is, from the time the on-chip chip or the safety management system detects a fault, the delay from the fault initiation to the startup system receiving the fault event is within 510ns.
[0093] Third, when a sub-functional system or chip fails, and its failure type is failure type 6-9, such as an abnormal type failure, a system interruption failure, etc., the CPU system handles the failed sub-functional system, chip and security management system. The specific process can be referred to the above embodiment and will not be repeated here.
[0094] In addition, when users completely trust the off-chip MCU system, they can choose to turn off the security management system. In this case, they only need to use the above method to handle the sub-function system, on-chip chip, and off-chip chip that have failed. This setting not only provides users with more operation options and flexibility, but also allows them to optimize the system's operating efficiency and performance according to actual needs, so as to respond to various application scenarios more flexibly.
[0095] With the above Figure 1-Figure 13Corresponding to the vehicle-track chip safety management method based on the core-grain architecture in the illustrated embodiment, the embodiment of the present application provides a corresponding system. Fig.14 FIG. 1 is a schematic diagram of a vehicle-track chip safety management system based on a chiplet architecture in an embodiment of the present application. Fig.14 As shown, the vehicle-track chip safety management system based on the chiplet architecture includes a fault information acquisition module 11, an allocation module 12 and a sending module 13, wherein the fault information acquisition module 11 is used to obtain the fault information of the functional module when any functional module in the chiplet architecture fails; the allocation module 12 is used to determine the corresponding fault processing module according to a preset fault processing allocation strategy; the sending module 13 is used to send the fault information of the functional module to the fault processing module, so that the fault processing module performs fault processing according to the fault information.
[0096] In some embodiments, the above fault information is generated by a functional module where a fault occurs, and the functional module where a fault occurs includes a sub-functional system or a chip.
[0097] In some embodiments, the above fault information is transmitted via a differential signal.
[0098] In some embodiments, the above-mentioned fault information includes the fault type and identification information of the functional module where the fault occurs.
[0099] In some embodiments, the fault processing module includes at least one of a security management system, a CPU system, a startup system, and an off-chip MCU system;
[0100] The above-mentioned fault handling allocation strategy includes at least one of a fault level, a safety level definition and a safety management method.
[0101] In some embodiments, the above-mentioned determining the corresponding fault processing module according to the preset fault processing allocation strategy includes:
[0102] When the fault processing module is normal, the fault processing module processes the faulty functional module according to the fault information.
[0103] In some embodiments, the above-mentioned determining the corresponding fault processing module according to the preset fault processing allocation strategy includes:
[0104] When the safety management system fails, the CPU system and the startup system process the failed functional module based on the fault information.
[0105] In some embodiments, it also includes:
[0106] When the safety management system fails or is shut down and the off-chip MCU system is normal, the off-chip MCU system processes the failed functional module according to the fault information.
[0107] In some embodiments, the fault processing module performs fault processing according to the fault information, including:
[0108] The fault processing module reads the register status information of the functional module where the fault occurs, and performs fault processing according to the fault information and the register status information.
[0109] As mentioned above, the system can manage automotive chips based on the core structure, use differential lines to transmit signals, and determine the corresponding fault handling module for each functional module that fails. This can not only effectively simplify signal routing and improve signal anti-interference, but also enhance the scalability of the automotive system, thereby improving the overall performance of the automotive system.
[0110] The present application also provides a vehicle, such as Fig.15 As shown, the vehicle includes the above Fig.14 The vehicle-track chip safety management system 21 based on the chiplet architecture shown in the figure can transmit fault signals according to differential lines, and use a preset fault processing allocation strategy to determine the fault processing module corresponding to the functional module where the fault occurs, thereby enhancing the anti-interference ability of the signal while enhancing the scalability of the vehicle-regulatory system, thereby improving the overall performance of the vehicle-regulatory system.
[0111] Although the embodiments of the present application have been shown and described, those skilled in the art will appreciate that various changes, modifications, substitutions and variations may be made to the embodiments without departing from the principles and spirit of the present application, and that the scope of the present application is defined by the claims and their equivalents.
Claims
1. A method for managing automotive chip safety based on chiplet architecture, characterized in that: include: When any functional module in the core grain architecture fails, obtaining the failure information of the functional module; Determine the corresponding fault handling module according to the preset fault handling allocation strategy; The fault information of the functional module is sent to the fault processing module, so that the fault processing module performs fault processing according to the fault information.
2. The method according to claim 1, characterized in that The fault information is generated by a functional module where a fault occurs, and the functional module where a fault occurs includes a sub-functional system or a chip.
3. The method according to claim 2, characterized in that The fault information is transmitted via a differential signal. 4 . The method according to claim 1 , wherein the fault information comprises a fault type and identification information of a functional module where the fault occurs.
5. The method according to any one of claims 1 to 4, characterized in that: The fault processing module includes at least one of a safety management system, a CPU system, a startup system, and an off-chip MCU system; The fault handling allocation strategy includes at least one of a fault level, a safety level, and a safety management method.
6. The method according to claim 5, characterized in that The determining of the corresponding fault processing module according to the preset fault processing allocation strategy includes: When the fault processing module is normal, the fault processing module processes the functional module where the fault occurs according to the fault information.
7. The method according to claim 5, characterized in that The determining of the corresponding fault processing module according to the preset fault processing allocation strategy includes: When the safety management system fails, the CPU system and the startup system process the failed functional module according to the failure information.
8. The method according to claim 5, further comprising: When the safety management system fails or is shut down and the off-chip MCU system is normal, the off-chip MCU system processes the failed functional module according to the fault information.
9. The method according to claim 1, characterized in that: The fault processing module performs fault processing according to the fault information, including: The fault processing module reads the register status information of the functional module where the fault occurs, and performs fault processing according to the fault information and the register status information.
10. A vehicle-track chip safety management system based on chiplet architecture, characterized in that: include: A fault information acquisition module, used to acquire fault information of a functional module when a fault occurs in any functional module in the core grain architecture; An allocation module, used to determine a corresponding fault handling module according to a preset fault handling allocation strategy; The sending module is used to send the fault information of the functional module to the fault processing module, so that the fault processing module performs fault processing according to the fault information.
11. A vehicle, characterized in that: Including the vehicle-track chip safety management system based on core-grain architecture as described in claim 10.
Citation Information
Patent Citations
Fault management system for vehicle-specification-level chip function safety
CN110955571A
Functional fault processing system and method thereof
CN114968646A
Fault management method and device for microcontroller
CN116382954A
Fault processing method and device, computer equipment and storage medium
CN117768311A
Processor capable of detecting fault and method of detecting fault of processor core using the same
US20140344619A1