Error detection architecture design method for multi-core safety-critical system
By introducing an asynchronous detection mechanism based on register checkpoint and access error detection in multi-core security critical systems, the problem of insufficient flexibility of error detection mechanism in the existing technology is solved, and more efficient system scheduling and reliability are achieved.
Patent Information
- Application Number
- CN202510100965.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-22
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2045-01-22
AI Technical Summary
The existing error detection mechanism is insufficient in a multi-core processor architecture and cannot efficiently meet the requirements of system reliability and scheduling at the same time.
An error detection architecture design method for multi-core security critical systems is introduced, and an asynchronous detection mechanism based on register checkpoint and memory access error detection is adopted, allowing the inspection core to execute asynchronously on any core outside the main core, supporting preemptive execution and dynamic configuration of core attributes.
It realizes more flexible error detection methods and more efficient system scheduling, reduces the area power consumption overhead caused by over-detection, and improves hardware flexibility and software layer scheduling performance.
Smart Images

Figure CN119937992A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of computer science, and in particular to safety-critical systems, and mainly relates to an error detection architecture design method for multi-core safety-critical systems. Background Art
[0002] In safety-critical systems such as automobiles and aerospace, the processor core is the core component to ensure system reliability and efficient operation. One of the key requirements of these systems is to be able to detect and correct potential hardware failures during task execution to ensure system reliability and stability. At the same time, these systems also need to complete complex task scheduling within a limited time to ensure that all tasks are completed on time to meet strict real-time requirements. However, reliability and schedulability focus on different dimensions of system safety, and they are usually implemented at different stages of processor and system development and at different architectural levels. The schedulability of the system is generally achieved through the scheduling algorithm of the operating system (OS). In order to ensure the reliability of the system, a hardware-level error detection mechanism is usually used. For example, the LockStep technology in the ARMCortex R series processor is a common hardware redundant error detection solution. This technology binds two or more identical processor cores together, executes the same program, and compares the output results at each clock cycle to detect possible errors in the processor.
[0003] However, with the current trend of integrating multiple safety-critical tasks on a shared processor core in safety-critical systems, traditional LockStep faces significant limitations. Due to its own core-bound rigid hardware design, LockStep performs the same level of error detection on all running tasks, ignoring the actual reliability requirements of the tasks, which leads to a waste of error detection capabilities, significant area and power consumption overhead, and challenges to system schedulability. To address these issues, the industry has proposed a LockStep method that supports split-lock (Split-Lock), in which Hybrid Modular Redundancy (HMR) reconfigures the hardware at runtime by explicitly distinguishing the criticality and performance requirements of tasks, thereby reducing resource consumption and providing a certain degree of flexibility. However, these methods still do not get rid of the core-bound hardware design, that is, the check core must perform error checking synchronously with the main core. This synchronous error checking execution cannot be interrupted by high-priority non-checking tasks, resulting in limitations on system schedulability. At the same time, these methods only support static error detection for scheduled tasks, lack the ability to selectively detect errors based on dynamic task requirements, and cannot flexibly adjust the error detection strategy according to the task load and real-time requirements during system runtime, which in turn affects the system's scheduling performance and resource utilization efficiency.
[0004] In summary, the existing error detection mechanism still has obvious problems such as insufficient flexibility and scheduling conflicts in the multi-core processor architecture, making it difficult to fully meet the needs of efficient multi-task scheduling while ensuring system reliability. Summary of the invention
[0005] The present invention discloses a problem that the existing error detection architecture has limited flexibility and cannot simultaneously and efficiently meet the requirements of system reliability and schedulability. The present invention introduces a more flexible hardware error detection architecture to liberate the core bound by the traditional architecture, and the support of the OS to cooperate with the flexible hardware, ultimately realizing an asynchronous, selectable error detection mechanism that can be preempted by non-checking tasks.
[0006] This architecture proposes a design method for error detection architecture for multi-core safety-critical systems; the basis of architecture asynchronous detection is error detection based on register checkpoints (RCPs) and memory access. The application thread running on the main core is divided into several small check segments, which can be re-executed on one or more check cores of the main core to verify the correctness of the main thread. Each check segment contains the start register checkpoints (SCPs), the end register checkpoints (ECPs) and the instructions that need to be re-executed. The check core initializes its own architecture state to SCP before running the check segment, and compares its own architecture state with ECP after the end of the check segment. During the operation, the relevant data of memory access operations, such as the address and data of memory access instructions such as Load / Store, are recorded and forwarded in the main core for the check core to re-execute and verify the correctness of memory access at runtime. If all ECPs and memory access operation data match the original execution, it proves that the main thread is correct. The basic idea of this method is that as long as all data related to RCPs and main thread memory access are recorded and temporarily buffered, the check thread can be executed asynchronously on any core other than the main core. That is, the check core does not need to start re-execution immediately to verify the correctness of the main thread, but to perform other tasks to improve the utilization and flexibility of the check core.
[0007] Any processor core in this architecture can be configured as the main core that runs application threads normally, the check core that performs correctness verification, and the ordinary computing core that does not participate in error detection. This allows threads running on any core to be reproduced and verified on different cores. The verification mode can be configured as one-to-one, one-to-two, or more modes to match task scenarios with different safety-critical requirements.
[0008] At the same time, compared with traditional pure hardware design solutions, the FlexStep architecture introduces OS, and all core attributes are visible to the OS, allowing the OS to dynamically configure core attributes and check modes according to the real-time and reliability requirements of different task loads during operation. In addition, FlexStep supports preemptive execution, allowing any core to be interrupted during operation and prioritize more urgent tasks to ensure that all tasks in the system can be completed before their deadlines.
[0009] In order to achieve the above-mentioned purpose, the technical solution adopted by the present invention adopts a method of software-hardware co-design, including a configurable micro-architecture based on RCPs that supports asynchronous thread-level error detection, a custom RISCV instruction set architecture (ISA) and a control interface provided for the OS scheduling algorithm, and a dedicated inspection thread developed for the inspection core.
[0010] In terms of hardware, we modified the microarchitecture on an open source processor core Rocket and added the necessary functional units to implement asynchronous thread-level error detection. These functional units include the checkpoint management unit (RCPM), the memory access record unit (MAL), and the data cache and channel unit (DBC). RCPM is used to manage register checkpoints and the number of instructions in each check segment, providing the check core with the running instruction boundaries and the architectural state for updating and comparison; MAL tracks all memory access operations and records related data for correctness verification; DBC is used for asynchronous data buffering and establishes data channels between cores through the system interconnect bus.
[0011] In terms of software, we use a customized ISA to abstract the underlying hardware and configuration methods into control flow, and integrate it into the context switching function of the OS layer scheduler, so that the OS can reconfigure the hardware at runtime to support task switching and preemption of different safety criticality levels.
[0012] Beneficial effects:
[0013] 1. The FlexStep architecture proposed in the present invention is a full-stack error detection architecture from the hardware layer microarchitecture to the software layer OS scheduling. Compared with the traditional pure hardware design and core-bound error detection mechanism, FlexStep implements a more flexible detection method and more efficient system scheduling, providing a comprehensive solution to the dual requirements of reliability and real-time performance of current safety-critical systems.
[0014] 2. In terms of hardware, this architecture implements a configurable asynchronous thread-level error detection mechanism for multi-core systems by adding additional functional units and modifying the microarchitecture, using reasonable hardware overhead. This effectively reduces the area and power consumption overhead caused by excessive detection, improves the flexibility of the hardware, and lays the foundation for flexible scheduling at the software layer.
[0015] 3. The introduction of software layer control flow enables the system to perform hardware configuration and task scheduling according to the requirements of different critical tasks, which significantly improves the schedulability of the system compared to traditional solutions. At the same time, the modification of the software layer is constrained to the context switching function of the OS scheduler, which minimizes the design intrusion and system complexity of the software layer, and provides an interface for further development of OS layer scheduling algorithms. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] Figure 1 is a comparison chart of the traditional Lockstep, HMR and FlexStep verification methods; ((a) Traditional LockStep fixed configuration of main core 0 and check core 1; (b) HMR limited flexibility configuration and synchronous verification method; (c) FlexStep runs asynchronous, optional and preemptible verification);
[0017] Figure 2 It is a general picture of the FlexStep architecture at the software and hardware level;
[0018] Figure 3 It is the execution graph of the main thread and the check thread when it is refined to the check segment;
[0019] Figure 4 These are the modifications made by FlexStep at the OS layer; (a) integrating the context switch function of the custom ISA; (b) checking the inspection thread developed by the inspection core). DETAILED DESCRIPTION
[0020] The present invention is further explained below in conjunction with the accompanying drawings and specific embodiments. It should be understood that the following specific embodiments are only used to illustrate the present invention and are not used to limit the scope of the present invention. It should be noted that the words "front", "rear", "left", "right", "upper" and "lower" used in the following description refer to directions in the accompanying drawings, and the words "inner" and "outer" refer to directions toward or away from the geometric center of a specific component, respectively.
[0021] Figure 1This is a comparison chart of the traditional Lockstep, HMR and FlexStep architecture verification methods. The Lockstep architecture requires the check core to participate in the entire process of the check. T3 cannot be placed on the core 1 that could have been released, causing T1 to miss the deadline. In addition, it requires that tasks that do not need to be checked must also be checked, resulting in additional power consumption. The HMR architecture has limited flexibility, allowing the core to execute T3 when it does not execute the T1 check, but due to its predefined fixed check of T2, the system cannot preempt T2 to meet the deadline of T1, which also causes T1 delays. The FlexStep architecture supports asynchronous, selectable and preemptible verification. When the system is running, the key segment of T2 is selected and scheduled to the idle part of the T1 execution cycle for asynchronous verification. At the same time, T1 that needs to meet the deadline first is allowed to preempt its verification, so that all tasks can meet the real-time requirements and critical tasks are fully checked.
[0022] a. Register Checkpoint Management (RCPM) is responsible for managing register checkpoints and instruction counts, providing execution boundaries and snapshots for verification for the check core. It consists of Checkpoint Control (CPC) and Architecture State Snapshot ASS.
[0023] Checkpoint control includes instruction count (Inst.Cnt) and privilege level monitoring (Priv.Mntr) units. For the main core, it generates a new ECP under the following conditions: a) privilege level mode switch occurs; b) instruction count limit is reached (default is 5000), and the number of instructions in the whole segment is recorded. In the whole check segment, the main core sends SCP, memory access information, instruction count (IC) and ECP in sequence (see Figure 3 The check core receives the ECP. For the check core, it starts checking after applying SCP until the instruction count reaches the same as the received IC. It then verifies the ECP using its own architectural state.
[0024] The Architecture State Snapshot (ASS) is responsible for temporarily storing register checkpoints and general architecture states to support replay execution and fast execution state switching. In the main core, before it sends SCP and ECP, ASS temporarily stores them and organizes them into a format suitable for reception by the inspection core before forwarding. In the inspection core, before the inspection thread enters the actual replay execution, it stores the architecture state in ASS instead of main memory, so that the previously stored architecture state can be quickly extracted when the replay execution reaches ECP, so as to return and obtain the suspended SCP in time.
[0025] b. The memory access log (MAL) tracks and records memory accesses for correctness checking. In the main core, the MAL identifies memory access instructions, records the address and data in the processor decoding stage, passes it along the pipeline to the submission stage, and packages it to the DBC when the submission is valid. In the check core, the MAL records information related to memory access instructions in the same way, and compares the information obtained from the DBC when it is valid to check the execution correctness.
[0026] c. Data Buffer and Channel (DBC) manages asynchronous data buffering and communication between cores through the system interconnect. It consists of data buffer FIFO and system interconnect.
[0027] The data buffer FIFO provides a place for the check core to temporarily store the check segment, providing a basis for the proposed asynchronous detection. In the master core, the data buffer can also provide a space for the master core to temporarily store the check segment when multiple master cores access a single slave core, so that only one master core accesses the slave core at the same time and the other master cores temporarily store the generated check segments, thereby avoiding conflicts and arbitration problems involved when multiple master cores access a single slave core at the same time.
[0028] The system interconnect enables inter-core communication between the main core and the check cores. It is implemented as a fully connected MUX-DEMUX network between the data buffer FIFOs of each core. There is a global register that can be modified by custom ISA instructions to generate control signals for the MUX / DEMUX to establish a connection between the main core and one or more check cores.
[0029] At the software level (shown in the green box), FlexStep provides
[0030] d. Customized instruction set architecture (see Table 1) to support the control interface between the hardware microarchitecture and the operating system to achieve
[0031] e. A control flow that can perform context switching between verification tasks and non-verification tasks and perform more flexible scheduling.
[0032] Figure 3 This is the execution graph of the main thread and the check thread when it is refined to the check segment. The legend indicates that a check segment consists of four parts: SCP, recorded memory access related information (including address and data), check segment instruction count and ECP, arranged in the order shown in the figure. Only when entering the kernel state can the OS switch the properties of the master and slave cores, that is, schedule the main thread and the check thread. The main figure shows that the main core and the check core enter the kernel state at regular intervals. In the main core, this will cause a check segment to end early (such as ①); in the slave core, after returning to the user state, the core will continue to check the previous check segment until it encounters the ECP of the check segment (such as ②). The figure also shows the difference in the magnitude of tasks and check segments, where the former is composed of many of the latter.
[0033] Figure 4 It is the modification made by FlexStep at the OS layer, which includes two aspects: (a) integration of the context switch function of the custom ISA, and (b) the development of the checking thread for the checking core.
[0034] The interaction between custom software and hardware in the FlexStep architecture is implemented through self-implemented RISC-V ISA extensions, as shown in Table 1. At each context switch of the OS, that is, when the OS needs to schedule threads so that the main thread or the check thread is scheduled to a processor core, the modified context switch function will call the custom ISA instructions to reconfigure the global registers. The global registers control the custom hardware so that its behavior conforms to the core attribute requirements of the thread scheduled on it by the OS. For example, if a main thread is scheduled to a core, the function will use the G.Configure instruction and M.associate to configure the global registers so that the core's attributes in the global registers are the master core and are associated with the appropriate check core. The changed global registers control the core to behave like a master core, that is, write data to the interconnect instead of taking data from the interconnect and writing it to the buffer unit as when it has the attributes of a slave core; at the same time, the global registers control the interconnect to establish the actual connection between the core and its associated slave cores.
[0035] Table 1 shows all the custom instructions introduced to properly configure the hardware.
[0036]
[0037] Table 1
[0038] From the perspective of the OS layer, FlexStep abstracts the inspection process into an inspection thread, which is scheduled like a normal thread. This requires the special definition of the inspection thread content. In this thread, the core will record its original architecture state into the ASS, and wait until the SCP of the first inspection segment is obtained, then apply its architecture state to itself to start the inspection.
[0039] Table 1 records all the custom instructions introduced to properly configure the hardware. The initial letter G indicates that the instruction is effective for any attribute core, while M and C represent that it only works on the main core or the check core. The design of this extended ISA is based on the following considerations and actual needs: it is necessary for all cores to know and configure the attributes of each core at each context switch; the attributes of the core may change during a context switch, so the execution of the main core and the verification of the slave core must be suspended before the switch, and resumed after re-determining the attributes of all cores; the main core needs to be associated with the corresponding check core; the slave core needs to record its own architectural state when starting the check, and when the check segment arrives, it needs to apply its SCP and start execution from the PC at the beginning of the check segment, and return the comparison results during and at the end of the execution to determine the correctness of the execution.
[0040] The technical means disclosed in the scheme of the present invention are not limited to the technical means disclosed in the above-mentioned implementation mode, but also include technical schemes composed of any combination of the above technical features.
Claims
1. A method for designing an error detection architecture for a multi-core safety-critical system, characterized by: The method of hardware-software co-design includes a configurable micro-architecture based on RCPs that supports asynchronous thread-level error detection, a custom RISCV instruction set architecture and a control interface provided for the OS scheduling algorithm, and a dedicated check thread developed for the check core. The following steps are included Step 1: Microarchitecture modifications were made to the open-source processor core Rocket, and necessary functional units were added to implement asynchronous thread-level error detection; these functional units include the checkpoint management unit RCPM, the memory access record unit MAL, and the data cache and channel unit DBC; MAL tracks all memory access operations and records related data for correctness verification; DBC is used for asynchronous data buffering and establishes data channels between cores through the system interconnect bus; Step 2: Use a customized ISA to abstract the underlying hardware and configuration methods into control flow and integrate them into the context switch function of the OS layer scheduler, so that the OS can reconfigure the hardware at runtime to support task switching and preemption at different safety criticality levels.
2. The method for designing an error detection architecture for a multi-core safety-critical system according to claim 1, characterized in that: In step 1, the checkpoint management unit RCPM is used to manage register checkpoints and the number of instructions in each check segment, provide the check core with running instruction boundaries and architectural states for updating and comparison; it consists of a checkpoint control CPC and an architectural state snapshot ASS.
3. The method for designing an error detection architecture for a multi-core safety-critical system according to claim 1, characterized in that: The checkpoint control cpc contains the instruction count Inst.Cnt and the privilege level monitoring Priv.Mntr unit; For the main core, a new ECP is generated under the following conditions: a) a privileged mode switch occurs; b) the instruction count limit is reached and the number of instructions for the entire segment is recorded; During the whole inspection phase, the main core sends SCP, memory access information, instruction count IC and ECP in sequence for the inspection core to receive; For the check core, it starts checking after applying SCP until the instruction count reaches the same as the received IC; then it uses its own architectural state to verify the ECP.
4. The method for designing an error detection architecture for a multi-core safety-critical system according to claim 1, characterized in that: The architecture state snapshot ASS is responsible for temporarily storing register checkpoints and general architecture states to support replay execution and fast execution state switching; in the main core, before it sends SCP and ECP, ASS temporarily stores them and organizes them into a format suitable for reception by the inspection core before forwarding them; in the inspection core, before the inspection thread enters the actual replay execution, it stores the architecture state in ASS, so that when the replay execution reaches ECP, the previously stored architecture state can be quickly extracted to return and obtain the suspended SCP in time.
5. The method for designing an error detection architecture for a multi-core safety-critical system according to claim 1, characterized in that: The memory access log MAL tracks and records memory accesses for correctness checking; In the main core, MAL identifies memory access instructions, records the device address and data in the processor decoding stage, passes it to the submission stage along the pipeline, and packages it and sends it to DBC when the submission is valid; in the check core, MAL records the relevant information of memory access instructions in the same way, and compares the information with that obtained from DBC when it is effectively submitted to check the execution correctness.
6. The method for designing an error detection architecture for a multi-core safety-critical system according to claim 1, characterized in that: Data Buffering and Channels DBC manages asynchronous data buffering and communication between cores through the system interconnect; It consists of a data buffer FIFO and a system interconnection; in the master core, the data buffer provides space for the master core to temporarily store the check segment when multiple master cores and a single slave core access it, so that only one master core accesses the slave core at the same time while other master cores temporarily store the generated check segments; the system interconnection realizes the inter-core communication between the master core and the check core; it is implemented as a fully connected MUX-DEMUX network between the data buffer FIFOs of each core; There is a global register that is modified by custom ISA instructions to generate control signals for the MUX / DEMUX to establish a connection between the main core and one or more check cores.
7. The method for designing an error detection architecture for a multi-core safety-critical system according to claim 1, characterized in that: The step 2 specifically includes Step 21: FlexStep provides a custom instruction set architecture to support the control interface between the hardware microarchitecture and the operating system, implementing a control flow that can perform context switching between verification tasks and non-verification tasks and more flexible scheduling; Step 22: The modifications made by FlexStep at the OS layer include two aspects: (a) integrating the context switch function of the custom ISA, and (b) developing a check thread for the check core; the interaction between the custom software and hardware in the FlexStep architecture is achieved through the self-implemented risc-vISA extension.
8. The method for designing an error detection architecture for a multi-core safety-critical system according to claim 7, characterized in that: The step 22 specifically includes Step 21: FlexStep provides a custom instruction set architecture to support the control interface between the hardware microarchitecture and the operating system, implementing a control flow that can perform context switching between verification tasks and non-verification tasks and more flexible scheduling; Step 22: The modifications made by FlexStep at the OS layer include two aspects: (a) integrating the context switch function of the custom ISA, and (b) developing a check thread for the check core; The interaction between custom software and hardware in the FlexStep architecture is achieved through self-implemented risc-vISA extensions.
Citation Information
Patent Citations
Multi-core parallel system processing method based on hardware protection
CN107357666A
Function security detection method based on two-out-of-two architecture of multi-core processor
CN114968781A
Multi-core DSP-oriented parallel computing resource self-organizing scheduling method and system
CN116360941A
Framework and method for system bypass monitoring and fault diagnosis
CN117762714A
Neural network accelerator automation design method based on FPGA
CN118690701A