System and method for automatic recovery in a lockstep processor

By receiving the pending instructions of the two processors at the lockstep monitor and detecting mismatch, and performing automatic recovery operations, the service interruption and complexity problems of pipeline error detection and recovery in the lockstep processor in the prior art are solved, and more efficient error handling and system stability are achieved.

CN115698953BActive Publication Date: 2025-06-06MICROCHIP TECHNOLOGY INC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202080100995.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-10-20
Filing Date
2020-12-04
Publication Date
2025-06-06
Estimated Expiration
2040-12-04

Smart Images

  • Figure CN115698953B_ABST
    Figure CN115698953B_ABST
Patent Text Reader

Abstract

A system and method for monitoring a processor operating in lockstep to detect a mismatch in pending pipeline instructions executed by the processor is provided. A lockstep monitor implemented in hardware is provided to detect the mismatch in the pending pipeline instructions executed on the lockstep processor and to initiate an automatic recovery operation at the processor if a mismatch is detected.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0002] This application claims priority to U.S. Provisional Patent Application Serial No. 63 / 030,201 filed on May 26, 2020 and U.S. Non-Provisional Patent Application Serial No. 17 / 075,493 filed on October 20, 2020, the contents of each of which are incorporated herein by reference in their entirety. Background Art

[0003] SEE (Single Event Effect) refers to a class of random events caused by particles of various atmospheric, solar and / or galactic origin that disrupt the operation of solid-state devices. SEE can cause a microprocessor to jump to an incorrect instruction, calculate an incorrect value and / or incorrectly update a register file or data memory. The sensitivity of solid-state devices to SEE is known to increase with altitude and is therefore an important consideration in the aerospace and avionics community. However, with the reduction in silicon geometry and the increase in device density in recent years, the statistical probability of SEE has increased and has become a consideration for deployment at sea level. Single event transient (SET) and single event upset (SEU) are two categories of SEE that can adversely affect logic circuits.

[0004] Mitigation techniques for reducing SEE are known in the art and include dual-modular redundancy (DMR) and triple-modular redundancy (TMR), replication with error detection, checkpointing and recovery, watchdog timer techniques, and error detection and correction (EDAC) protection for memory devices.

[0005] Lockstep dual core is a method of grouping two independent and identical processor cores to achieve dual modular redundancy (DMR), thereby providing a method of detecting errors that occur in one of the two processors. During system startup, the two processors are initialized to the same state, and when they operate in lockstep, they receive the same input and execute the same instruction stream simultaneously. Therefore, during normal operation, the states of the two processors are exactly the same clock by clock. Lockstep processing assumes that an error in either processor will result in a difference between the states of the two processors, which will ultimately manifest as a difference in the output of the processor cores. When an internal error occurs in one of the processor cores, the actions of the core will be different after the internal error, which can cause the processor core to terminate lockstep.

[0006] Lockstep monitors that detect the outputs of dual processors operating in lockstep and signal an error when a difference between the outputs is detected are known in the art. However, in prior art lockstep monitoring systems, the outputs of the lockstep processors are only compared at the system bus, and recovery from a detected error detected on the system bus requires a system reset of both processors. Performing a reset results in an interruption of service and requires increased system design complexity to mitigate the effects of the interruption. While other techniques are known that utilize saved software checkpoints to restore the state of the processors and continue execution, these other techniques require modifications to the software and can reduce processing throughput.

[0007] Therefore, what is needed in the art is an improved system and method for detecting errors in pipeline processing steps executed in a lockstep processor. Additionally, what is needed is an improved method for recovering from a detected mismatch of pending pipeline instructions. Summary of the invention

[0008] In various embodiments, the present invention provides an improved system and method for identifying pipeline processing mismatches in two processors operating in lockstep, and providing automatic recovery operations for the processors when a mismatch is detected.

[0009] In one embodiment, the present invention provides a method, the method comprising: receiving at a lockstep monitor a first pending instruction generated by a first processor executing pipeline instructions, and receiving at the lockstep monitor a second pending instruction generated by a second processor executing pipeline instructions in lockstep with the first processor. The method further comprises: comparing the first pending instruction at the lockstep monitor with the second pending instruction to detect a mismatch, and when a mismatch is detected, performing an automatic recovery operation at the first processor and the second processor to prevent the first processor from executing the first pending instruction and prevent the second processor from executing the second pending instruction.

[0010] The automatic recovery operation at the first processor and the second processor may further include: generating a flush and re-execute interrupt at the lockstep monitor; transmitting the flush and re-execute interrupt to the first processor and the second processor; flushing pipeline instructions in response to the transmitted flush and re-execute interrupt; and re-executing the flushed pipeline instructions in response to the transmitted flush and re-execute interrupt.

[0011] In a particular embodiment, the first pending instruction and the second pending instruction include one or more of: a pending instruction fetch address, a pending write address, and pending write data.

[0012] In another embodiment, the present invention provides a system comprising: a first processor that executes pipeline instructions; a second processor that executes pipeline instructions in lockstep with the first processor; and a lockstep monitor that is coupled to the first processor and the second processor. The lockstep monitor includes a checkpoint circuit system for receiving a first pending instruction from the first processor and receiving a second pending instruction from the second processor. The checkpoint circuit system of the lockstep monitor compares the first pending instruction with the second pending instruction to detect a mismatch. The system further includes an automatic recovery circuit system coupled to the checkpoint circuit system. When a mismatch is detected by the checkpoint circuit system, the automatic recovery circuit system initiates an automatic recovery operation at the first processor and the second processor to prevent the first processor from executing the first pending instruction and prevent the second processor from executing the second pending instruction.

[0013] Thus, the present invention provides a system and method for detecting errors in pipeline processing steps executed in a lockstep processor. Additionally, an improved method for recovering from a detected mismatch of pending pipeline instructions is provided. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] Figure 1 is a block diagram illustrating a lockstep monitor according to one embodiment of the present invention.

[0015] Figure 2 is a block diagram illustrating the integration of a lockstep monitor with two identical processors operating in lockstep, according to one embodiment of the present invention.

[0016] Figure 3 is a block diagram illustrating pipeline processing of one of the lockstep processors and the interaction of a lockstep monitor with the pipeline processing of the processor in accordance with one embodiment of the present invention.

[0017] Figure 4 is a table illustrating an exemplary instruction pipeline for a lockstep processor according to one embodiment of the present invention.

[0018] Figure 5 is a flow chart illustrating a method for detecting mismatches in pending pipeline instructions of a processor operating in lockstep according to one embodiment of the present invention.

[0019] Figure 6 is a flow chart illustrating steps performed during an auto-recovery operation when a mismatch between two lockstep processors is detected, according to one embodiment of the present invention.

[0020] Figure 7is a flow chart illustrating a comparison performed by a lockstep monitor to detect a mismatch in a pending instruction executed in the lockstep monitor according to one embodiment of the present invention. DETAILED DESCRIPTION

[0021] In various embodiments, the present invention provides a system and method for detecting mismatches of pending instructions at various points in the instruction pipelines of identical dual processors operating in lockstep. In addition, the present invention provides a system and method for performing automatic recovery of the dual processors in the event of a mismatch being detected, thereby avoiding a potential loss of lockstep that would require resetting both processors before resuming lockstep processing. The present invention detects SEE failures early in the pipeline and applies appropriate corrections. The present invention does not require software checkpoints to be stored at the dual processors during execution of pipeline instructions, and significantly reduces the likelihood that the processors will need to be reset in the event that lockstep is terminated due to an error in one of the processors.

[0022] refer to Figure 1 , a system 100 for performing mismatch detection and automatic recovery of dual processors operating in lockstep includes a lockstep monitor 105 coupled to a first processor 127 and a second processor 132. It is assumed that the first processor 127 and the second processor 132 are identical and operate in lockstep, as such they execute the same sequence of pipeline instructions on a clock-by-clock basis. In a particular embodiment, the system 100 for performing mismatch detection and automatic recovery of the first processor 127 and the second processor 132 operating in lockstep can be implemented in an integrated circuit, such as a system on a chip (SOC) field programmable gate array (FPGA). The lockstep monitor 105 effectively detects mismatches in pending instructions that may potentially result in a loss of lockstep between the first processor 127 and the second processor 132. In addition, the lockstep monitor 105 effectively performs automatic recovery of the first processor 127 and the second processor 132 when a mismatch is detected, thereby avoiding the need to perform a reset on the first processor 127 and the second processor 132.

[0023] In the following description of the present invention, it is assumed that there is zero clock skew between the first processor 127 and the second processor 132 when they operate in lockstep. However, the implementation can be extended to allow for clock skew between the first processor 127 and the second processor 132. Allowing for clock skew lockstep processing is known in the art, with the tradeoff being increased complexity and reduced throughput.

[0024] The lockstep monitor 105 includes a checkpoint circuit system 110 and an automatic recovery circuit system 115. The checkpoint circuit system 110 further includes a pending instruction fetch address checkpoint circuit system 120, a pending write data and pending write address checkpoint circuit system 125, and a pending system bus checkpoint circuit system 130. The lockstep monitor 105 establishes hardware checkpoints at the first processor 127 and the second processor 132 during the lockstep execution of pipeline instructions to detect a mismatch between a pending instruction to be executed by the first processor 127 and a pending instruction to be executed by the second processor 132. In addition, when a mismatch is detected, the lockstep monitor 105 initiates automatic recovery of the lockstep first processor 127 and the second processor 132.

[0025] In operation of the system 100 of the present invention, when the first processor 127 and the second processor 132 operate in lockstep, the lockstep monitor 105 performs a hardware check on the pending instructions to be executed by each of the first processor 127 and the second processor 132 to determine whether there is a mismatch between the pending instructions. The hardware check may be performed every clock cycle to provide the finest granularity of error detection. Alternatively, the hardware check may be performed at regular predetermined intervals rather than every clock cycle.

[0026] In a specific embodiment, the pending instruction fetch address checkpoint circuit system 120 may receive a first pending instruction fetch address from the first processor 127 and a second pending instruction fetch address from the second processor 132, wherein both the first processor 127 and the second processor 130 execute the same pipeline instruction in lockstep. The pending instruction fetch address checkpoint circuit system 120 compares the received first pending instruction fetch address with the received second pending instruction fetch address to determine whether there is a mismatch between the instructions. When a mismatch is detected, the pending instruction fetch address checkpoint circuit system 120 provides a signal 160 that a mismatch has been detected to the automatic recovery circuit system 115. The automatic recovery circuit system 115 then sends a flush and replay interrupt 166 to the first processor 127 and the second processor 132 to perform an automatic recovery operation. In response to receiving the flush and replay interrupt 166, the first processor 127 and the second processor 132 flush their pipeline instructions and replay the flushed pipeline instructions without resetting the processors 127, 132. As such, in response to the flush and replay interrupt 166 , the auto-recovery operation prevents the first processor 127 from executing the first pending instruction fetch address and prevents the second processor 132 from executing the second pending instruction fetch address in response to receiving the flush and replay interrupt 166 .

[0027] In another embodiment, the pending write data and pending write address checkpoint circuit system 125 may receive the first pending write data or the first pending write address from the first processor 127 and the second pending write data or the second pending write address from the second processor 132, wherein both the first processor 127 and the second processor 132 execute the same pipeline instruction in lockstep. The pending write data and pending write address checkpoint circuit system 125 compares the received first pending write data or the first pending write address with the received second pending write data or the second pending write address, respectively, to determine whether there is a mismatch between the pending write data or the pending write address. When a mismatch is detected, the pending write data and pending write address checkpoint circuit system 125 provides a signal 162 to the automatic recovery circuit system 115 that a mismatch has been detected. The automatic recovery circuit system 115 then sends a flush and replay interrupt 166 to the first processor 127 and the second processor 132 to perform an automatic recovery operation. In response to receiving the flush and replay interrupt 166, the first processor 127 and the second processor 132 flush their pipeline instructions and replay the flushed pipeline instructions without resetting the first processor 127 and the second processor 132. As such, in response to the flush and replay interrupt 166, the automatic recovery operation prevents the first processor 127 from writing the first pending write data or writing to the first pending write address and prevents the second processor 132 from writing the second pending write data or writing to the second pending write address.

[0028] In another embodiment, the pending system bus checkpoint circuit system 130 may receive a first pending instruction from the first processor 127 and a second pending instruction from the second processor 132, wherein both the first processor 127 and the second processor 130 execute the same pipeline instruction in lockstep. The pending system bus checkpoint circuit system 130 compares the received first pending instruction with the received second pending instruction to determine whether there is a mismatch between the pending instructions. When a mismatch is detected, the pending write data and pending write address checkpoint circuit system 125 provides a signal 164 to the automatic recovery circuit system 115 that a mismatch has been detected. The automatic recovery circuit system 115 then sends a flush and replay interrupt 166 to the first processor 127 and the second processor 132 to perform an automatic recovery operation. In response to receiving the flush and replay interrupt 166, the first processor 127 and the second processor 132 flush their pipeline instructions and replay the flushed pipeline instructions without resetting the first processor 127 and the second processor 132. As such, in response to the flush and replay interrupt 166, the auto-recovery operation prevents the first processor 127 from executing the first pending instruction and prevents the second processor 132 from executing the second pending instruction.

[0029] As will be described in more detail below, the lockstep monitor 105 of the present invention is used to detect mismatches between pending instructions at various points in the data pipeline of a dual processor operating in lockstep. The present invention detects SEE failures early in the pipeline and applies appropriate corrections. In addition, the hardware of the lockstep monitor 105 effectively notifies the lockstep first processor 127 and the second processor 132 of the mismatch by sending a flush and re-execute interrupt, which will prevent the first processor 127 and the second processor 132 from executing the pending instructions involved in the detected mismatch. The present invention utilizes existing flush and re-execute interrupts known in computer architecture to resolve branch mispredictions. A microarchitecture that supports branch prediction / misprediction is utilized by instructing the first processor 127 and the second processor 132 to perform a re-execution of one or more instructions executed before the mismatch was detected. As such, recovery from the SEE is automatically performed in hardware using an automatic recovery operation, and no system software changes to the processor are required. Embodiments of the present invention allow for automatic recovery operations to successfully maintain lockstep in the processors with a high probability and avoid system resets that cause undesired service interruptions and increased system design complexity to mitigate the effects of the interruptions. The present invention also eliminates the need to periodically implement software to periodically save the state of the first processor 127 and the second processor 132 without knowing in advance when a failure may occur. Saving the state of the first processor 127 and the second processor 132 to be able to recover to the saved state undesirably reduces throughput and requires significant changes to software code to implement in existing processor-based systems.

[0030] Figure 2 Further shown is the first processor 127 and the second processor 132 operating in lockstep with Figure 1 The interaction between the elements of the checkpoint circuit system 110.

[0031] like Figure 2As shown in FIG. 1 , the first processor 127 includes a corresponding processor core 190, a corresponding register file 205, a corresponding instruction tightly controlled memory (ITCM) 135, a corresponding data tightly controlled memory (DTCM) 145, and a corresponding program counter 155. The second processor 132 includes a corresponding processor core 195, a corresponding register file 207, a corresponding ITCM 140, a corresponding DTCM 150, and a corresponding program counter 160. ITCM and DTCM are generally used in computer architectures with separate storage devices and signal paths for instructions and data. ITCM and DTCM are attached to different elements of the system bus, wherein ITCM is coupled to the instruction bus and is used to store executable instructions, and DTCM is coupled to the data bus and is used to store data. Therefore, the processor core 190 obtains executable instructions from ITCM 135, stores data at DTCM 145, and obtains data from DTCM 145. The processor core 195 obtains executable instructions from ITCM 140, stores data at DTCM 150, and obtains data from DTCM 150. Respective program counters 155, 160 are registers containing the address of the next instruction to be executed in the pipeline and are coupled to ITCMs 135, 140, respectively.

[0032] The pending instruction fetch address checkpoint circuitry 120 provides a first opportunity in the pipeline process to detect a mismatch between pending instructions executed on two lockstep processors 127, 132. The pending instruction includes a pending instruction fetch address specified for the ITCM of the processor. As shown, the pending instruction fetch address checkpoint circuitry 120 is coupled between the program counter 155 and the ITCM 135 of the first processor 127, and is coupled between the program counter 160 and the ITCM 140 of the second processor 132. As such, the pending instruction fetch address checkpoint circuitry 120 receives a first pending instruction fetch address specified for the ITCM 135 of the first processor 127 from the program counter 155, and receives a second pending instruction fetch address specified for the ITCM 140 of the second processor 132 from the program counter 160. For each of the first processor 127 and the second processor 132, the pending instruction fetch address identifies the ITCM address that will be used to fetch the next instruction to be executed by that processor in the pipeline. The pending instruction fetch address checkpoint circuitry 120 compares the first pending instruction fetch address with the second pending instruction fetch address to detect a mismatch. Since the first processor 127 and the second processor 132 operate in lockstep, the first pending instruction fetch address and the second pending instruction fetch address should match in the absence of an error. If the pending instruction fetch address checkpoint circuitry 120 detects a mismatch, the pending instruction fetch address checkpoint circuitry 120 provides a signal 160 to the automatic recovery circuitry 115, and in response to the provided signal 160, the automatic recovery circuitry 115 sends a flush and re-execute interrupt 166 to the first processor 127 and the second processor 132 to prevent the first pending instruction fetch address from being executed at the first processor 127 and to prevent the second pending instruction fetch from being executed at the second processor 132. The pipeline instruction is flushed from the pipeline and re-executed on both the first processor 127 and the second processor 132. By identifying differences between pending instruction fetch addresses prior to execution by the ITCMs 135 , 140 , the lockstep monitor 105 provides recovery and continued lockstep operation between the first processor 127 and the second processor 132 without requiring a time consuming reset of the first processor 127 and the second processor 132 .

[0033] The second opportunity for the lockstep monitor 105 to detect mismatches in pending instructions at the first processor 127 and the second processor 132 is provided by the pending write data and pending write address checkpoint circuitry 125. In response to monitoring the bus to identify an execution instruction designated for the processor cores 190, 192, the pending write data and pending write address checkpoint circuitry 125 receives a first pending write data or a first pending write address designated for the register file 205 or DTCM 145 of the first processor 127 from the first processor core 190. In addition, the pending write data and pending write address checkpoint circuitry 125 additionally receives a second pending write data or a second pending write address designated for the register file 207 or DTCM 150 of the second processor 132 from the second processor core 195. The pending write data is data to be written to the corresponding register file 205, 207 or the corresponding DTCM 145, 150 in the next executed pipeline instruction. The pending write address identifies the address of the corresponding register file 205, 207 or the corresponding DTCM 145, 150 to which the write data will be written in the pipeline instruction to be executed next. The pending write data and pending write address checkpoint circuit system 125 compares the first pending write data and / or the first pending write address with the second pending write data and / or the second pending write address to detect a mismatch. Since the first processor 127 and the second processor 132 operate in lockstep, the first pending write data and the second pending write data should match in the absence of an error. If the first pending write data does not match the second pending write data, a mismatch is detected. Additionally or alternatively, if the first pending write address does not match the second pending write address, a mismatch is detected. If the pending write data and pending write address checkpoint circuitry 120 detects a mismatch, the pending write data and pending write address checkpoint circuitry 125 sends a signal 162 to the automatic recovery circuitry 115 indicating that an automatic recovery operation should be initiated for the first processor 127 and the second processor 132, and in response to the signal 162, the automatic recovery circuitry 115 sends a flush and re-execute interrupt 166 to the first processor 127 and the second processor 132 to prevent the first pending write data or the first pending write address from being executed at the first processor 127 and to prevent the second pending write data or the second pending write address from being executed at the second processor 132. The pipeline instructions are flushed from the pipeline and re-executed on both the first processor 127 and the second processor 132. By identifying differences between pending write data and / or pending write addresses specified for register files 205 , 207 or DTCMs 145 , 150 , lockstep monitor 105 provides recovery and continued lockstep operation between first processor 127 and second processor 132 without requiring a time consuming reset of processors 127 , 132 .

[0034] A third opportunity for the lockstep monitor 105 to detect a mismatch in the pending instructions at the first processor 127 and the second processor 132 is provided by the pending system bus checkpoint circuitry 130. The pending system bus checkpoint circuitry 130 receives a first pending instruction designated for the system bus 250 from the first processor 127. The pending system bus checkpoint circuitry 130 also receives a second pending instruction designated for the system bus 250 from the second processor 132. The pending system bus checkpoint circuitry 130 compares the first pending instruction with the second pending instruction to detect a mismatch. Since the first processor 127 and the second processor 132 operate in lockstep, the first pending instruction should match the second pending instruction in the absence of an error. If the first pending instruction does not match the second pending instruction, a mismatch is detected. If the pending system bus checkpoint circuitry 130 detects a mismatch, the pending system bus checkpoint circuitry 130 sends a signal 164 to the automatic recovery circuitry 115 indicating that an automatic recovery operation should be initiated for the first processor 127 and the second processor 132. In response to the signal 164, the automatic recovery circuit 115 sends a flush and replay interrupt 166 to the first processor 127 and the second processor 132 to prevent the first pending instruction from being executed at the first processor 127 and to prevent the second pending instruction from being executed at the second processor 132. The pipeline instructions are flushed from the pipeline and re-executed on both processors 127, 132. By identifying differences between pending instructions prior to execution on the processors 127, 132, the lockstep monitor 105 provides recovery and continued lockstep operation between the processors 127, 132 without requiring a time-consuming reset of the processors 127, 132.

[0035] In the present invention, when a mismatch of pending instructions is detected, the automatic recovery circuit system 115 sends a clear and re-execute interrupt 166 to the first processor 127 and the second processor 132, and then submits an update (i.e., write) to the DTCM, register, or system bus. In addition, if the mismatch continues to exist after the automatic recovery operation, the processor core is stopped, the system bus is isolated, and a single event functional interrupt (SEFI) is triggered by the automatic recovery circuit system 115 to reset the processor. Specifically, the automatic recovery circuit system 115 transmits the SEFI to the processors 127, 132 to indicate that the automatic recovery operation cannot resolve the mismatch, and the automatic recovery circuit system 115 provides a safety monitoring function as a fail-safe mechanism to resolve the continued existence of the mismatch, which is an infrequent occurrence.

[0036] In addition, when a mismatch of pending instructions is detected, it is assumed that the pending instruction at one of the processors is correct, while the pending instruction at the other processor is incorrect. As such, after performing an automatic recovery operation, the pending instruction at only one of the processors should be different from the pending instruction before the automatic recovery operation. In order to address this infrequent occurrence, where the pending instructions at the two processors are different after the automatic recovery operation, the previous state of the hardware checkpoint caused by the mismatch is saved at the lockstep monitor 105 and compared with the current state of the hardware checkpoint after attempting automatic recovery. Therefore, if the current state on the two processor cores is different from the previous state, SEFI is triggered, even if the mismatch does not continue to exist after the automatic recovery.

[0037] Figure 3 An exemplary processor pipeline executing in the first processor 127 and the relationship between the first processor 127 and the lockstep monitor 105 are shown. Although the operation of the first processor 127 is described in detail below, it should be understood that the second processor 132 operates in lockstep with the first processor 132 and the lockstep monitor 105 interacts with both the first processor 127 and the second processor 132 to detect pending instruction mismatches.

[0038] As is well known in the art, the single-cycle data path of each processor in the first processor 127 and the second processor 130 is split into five functional units separated by control buffers. The instruction fetch (IF) functional unit is separated from the instruction decode (ID) functional unit by the IF-ID buffer 315. The ID functional unit is separated from the execution state (EX) functional unit by the ID-EX buffer 320. The EX functional unit is separated from the memory cycle (MEM) functional unit by the EX-MEM buffer 325, and the MEM functional unit is separated from the write back (WB) functional unit by the MEM-WB buffer 330. The buffers 315, 320, 325, 330 store the results of the previous stage so that the results can be used in the next clock cycle. The control lines 317, 322, 327, 332 are used to route data to functional elements in the circuit system, such as the register file 205 and the DTCM 145. Arithmetic logic units (ALUs) 340 , 342 , 344 and control signal logic 346 , 348 , 349 also help maintain pipeline processing at the first processor 127 .

[0039] When the processor operates in lockstep, the logic state of the core is tracked and kept in sync. The complete logic state of the core can be divided into persistent components such as command and status registers, program counters, and register files, and transient components such as execution pipelines and data memory. The system and method of the present invention tracks the flow of instructions through the pipeline of a processor operating in lockstep. Early detection of pending instruction mismatches in the pipeline prevents failures from propagating forward in the pipeline and provides an opportunity for automatic recovery without explicit intervention in software.

[0040] like Figure 3 As shown in FIG, the program counter 155 of the first processor 127 is coupled to the ITCM 135 of the first processor 127. During pipeline execution, the program counter 155 provides the pending instruction fetch address 205 to the ITCM 135. After receiving the pending instruction fetch address 205 at the read address of the ITCM 135, the ITCM 135 may respond by providing the instruction corresponding to the pending instruction fetch address 205 to the instruction fetch IF-ID buffer 315.

[0041] The instructions may be decoded and the results provided as inputs 350, 352 to the register file 205. In response, the register file 205 may provide outputs 345, 356 to the ID-EXE buffer 320. After execution in the EX functional units including, but not limited to, the ALUs 340, 342 and the control signal logic 348, the results may be provided to the EXE-MEM buffer 325. The resulting data 230 may then be written to the DTCM 145 at the specified write address 225. The data stored at the DTCM 145 is buffered at the MEM-WB buffer 330 and subsequently provided to the register file 205 at the output of the control signal logic 340 at the specified write address 210 provided by the MEM-WB buffer 330 as write-back data 215.

[0042] In the present invention, the lockstep monitor 105 provides hardware checkpoints at various locations in the processing pipeline. The pending instruction fetch address checkpoint circuitry 120 provides hardware checkpoints for the pending instruction fetch address 205 specified for the ITCM 135. The pending write data and pending write address checkpoint circuitry 125 provide hardware checkpoints for the pending write back data 215 and the pending specified write address 210 specified for the register file 205. The pending write data and pending write address checkpoint circuitry 125 also provide hardware checkpoints for the specified write address 225 and result 230 specified for the DTCM 145. The pending system bus checkpoint circuitry 130 provides hardware checkpoints for pending instructions 390 from the system bus 250.

[0043] Although only Figure 31 shows the processing pipeline for the first processor 127, but as previously described, the lockstep monitor 105 is also coupled to the second processor 132 to provide a reference Figure 3 The same hardware checkpoints described in connection with the first processor 127. As such, it is determined that the pending instruction fetch address checkpoint circuitry 120 receives pending instruction fetch addresses specified for the respective ITCMs 135, 140 from the processing pipelines of both the first processor 127 and the second processor 132. The pending instruction fetch address checkpoint circuitry 120 compares the received instruction fetch addresses from the processors to detect mismatches. The pending write data and pending write address checkpoint circuitry 125 receives pending write data and pending write addresses specified for the respective register files 205, 207 or the respective DTCMs 145, 150 from both the first processor and the second processor. The pending write data and pending write address checkpoint circuitry 125 then compares the received pending write data and / or pending write addresses from the processors to detect mismatches. The pending system bus checkpoint circuitry 130 receives pending instructions specified for the system bus 250 from both the first processor 127 and the second processor 132. The pending system bus checkpoint circuitry 130 then compares the received pending instructions from the processor to detect mismatches.

[0044] In a particular embodiment, depending on when a pending write data mismatch is detected in the pipeline, the comparison may be performed while the current contents 220 of the register file 205 are saved in holding registers of the pending write data and pending write address checkpoint circuitry 125. If a mismatch is detected, the register file 205 may be restored to the saved contents 220 to reduce delays in performing automatic recovery operations.

[0045] The result of the comparison at the hardware checkpoint may then be used to perform an automatic recovery operation of the first processor 127 and the second processor 132. Specifically, after a mismatch is detected at the pending instruction fetch address checkpoint circuitry 120, a signal 160 is sent to notify the automatic recovery circuitry 115 to send a flush and replay interrupt 166 to the first processor 127 and the second processor 132 to initiate an automatic recovery operation. After a mismatch is detected at the pending write data and pending write address checkpoint circuitry 125, a signal 162 is sent to notify the automatic recovery circuitry 115 to send a flush and replay interrupt 166 to the first processor 127 and the second processor 132 to initiate an automatic recovery operation. After a mismatch is detected at the pending system bus checkpoint circuitry 130, a signal 164 is sent to notify the automatic recovery circuitry 115 to send a flush and replay interrupt 166 to the first processor 127 and the second processor 132 to initiate an automatic recovery operation.

[0046] Figure 4 Provided is a table showing the execution of pipelines in processors 127, 132 operating in lockstep. Pipelines improve efficiency by dividing instructions into a fixed number of steps, and each step is implemented as a pipeline segment. As shown, up to five instructions can be inflight in the pipeline. For example, during clock cycle 0, executable instruction fetches IF(0), and during clock cycle 1, the instruction fetched during clock cycle 0 can be decoded as ID(0), and the next instruction can be obtained as IF(1). It is thus concluded that at clock cycle 4, five instructions are in the pipeline, wherein instruction fetches IF(0) is now in the write-back stage WB(0), instruction fetches IF(1) is in the memory stage MEM(1), instruction fetches IF(2) is in the execution state EXE(2), instruction fetches IF(3) is in the instruction decoding stage ID(3), and instruction fetches IF(4) is currently being fetched. Known pipelines include dependency and conflict checks to ensure correct pipeline instruction execution flow. For simplicity, conditional branches and branch predictions are not shown in this table. The key principle of the present invention for performing automatic recovery is to identify any SEE induced faults early in the execution flow of the pipeline before committing state changes to the corresponding ITCM 135 , 140 , the corresponding DTCM 145 , 150 or the corresponding register files 205 , 207 .

[0047] In the Figure 4 Table with Figure 2 In a related exemplary embodiment, if the pending instruction fetch address, IF(3) (indicated at 300), specified by the ITCM 135 for the first processor 127 during clock cycle 3 does not match the pending instruction fetch address IF(3) specified by the ITCM 140 for the second processor 132 during clock cycle 3, the pending instruction fetch address checkpoint circuitry 120 may detect the mismatch in the pipeline at 400. If a mismatch is detected, the pending instruction fetch address is not executed at the ITCMs 135, 140, and pipeline instructions that would have been executed in the remaining clock cycles are flushed and re-executed.

[0048] In the Figure 4 Table with Figure 2In another example embodiment of the related, if the pending write-back data or pending specified write address WB(3) specified by the register file 205 for the first processor 127 during clock cycle 7 does not match the pending write-back data or pending specified write address WB(3) specified by the register file 207 for the second processor 132 during clock cycle 7, the pending write data and pending address checkpoint circuitry 125 may detect a mismatch in the pipeline at 310. If a mismatch is detected, the pending write-back data or pending specified write address is not executed at the register files 205, 207, and pipeline instructions that would otherwise be executed in the remaining clock cycles are flushed and re-executed.

[0049] Figure 5 is a flow chart illustrating a method 500 for automatic recovery in a processor operating in lockstep according to one embodiment of the present invention. At operation 505, the method begins by receiving, at a lockstep monitor, a first pending instruction generated by a first processor executing pipeline instructions. Figure 1 , the lockstep monitor 105 is coupled to the first processor 127 to receive a first pending instruction generated by the first processor 127 .

[0050] At operation 510, the method continues by receiving, at a lockstep monitor, a second pending instruction generated by a second processor executing pipeline instructions in lockstep with the first processor. Figure 1 , the lockstep monitor 105 is coupled to the second processor 132 to receive a second pending instruction generated by the second processor 132 .

[0051] At operation 515, the method continues by comparing the first pending instruction at the lockstep monitor with the second pending instruction to detect a mismatch. Figure 1 The lockstep monitor 105 includes a circuit system for comparing a first pending instruction with a second pending instruction.

[0052] At operation 520 , when a mismatch is detected, the method ends by performing an automatic recovery operation at the first processor and the second processor to prevent the first processor from executing the first pending instruction and to prevent the second processor from executing the second pending instruction.

[0053] Figure 6 is a flow chart describing the automatic recovery operation 520 of the present invention in more detail. At operation 600, the method includes storing the first pending instruction and the second pending instruction that caused the mismatch at the lockstep monitor. Figure 1 , the lockstep monitor 105 may include a memory and associated circuitry for storing the first pending instruction and the second pending instruction.

[0054] At operation 605, the method continues by generating a flush and replay interrupt at the lockstep monitor 605 and transmitting the flush and replay interrupt to the first processor 127 and the second processor 132 at operation 610. Figure 1 , the lockstep monitor 105 may include an automatic recovery circuitry 115 for generating a flush and replay interrupt 166 and transmitting the flush and replay interrupt 166 to the first processor 127 and the second processor 132 .

[0055] At operation 615, the method continues by flushing the pipeline instructions and re-executing the pipeline instructions 615. Figure 1 , after receiving the flush and replay interrupt 166 from the lockstep monitor 105 , the first processor 127 and the second processor 132 proceed by flushing their pipeline instructions and replaying the pipeline instructions to recover from the detected mismatch. Figure 4 Exemplary pipeline instructions executed by the first processor 127 and the second processor 132 are shown in FIG.

[0056] After the automatic recovery operation is performed, Figure 5 The comparison process is repeated at operation 505. At operation 620, if the mismatch continues to exist after the comparison is repeated, the method continues at operation 625 by performing a reset of the first processor and the second processor. Figure 1 , if the mismatch continues to exist after automatic recovery has been performed, the lockstep monitor 105 initiates a reset of the first processor 127 and the second processor 132.

[0057] Alternatively, if the mismatch does not continue to exist at operation 620, the method continues at operation 630 by comparing the stored first pending instruction that caused the mismatch with the current first pending instruction, and comparing the stored second pending instruction that caused the mismatch with the current second pending instruction. Figure 1 The lock-step monitor 105 compares the instruction stored in the memory with the instruction Figure 5 The currently pending instruction is compared with the one produced by the most recent comparison operation.

[0058] The method continues at operation 635. If both the stored first pending instruction and the second pending instruction are different from the corresponding current first pending instruction and the second pending instruction, then at operation 625, a reset operation is performed at the first processor and the second processor. Alternatively, if only one of the stored first pending instruction and the second pending instruction is the same as the current first pending instruction and the second pending instruction, then no reset operation is performed and the method continues. Figure 5The method continues with operation 505. By storing the pending instruction that caused the mismatch and then comparing the stored pending instruction to the current pending instruction, the method prevents the situation where there is no mismatch between the current pending instructions at the processors, but neither processor has a current pending instruction that matches a previous pending instruction.

[0059] Figure 7 Yes Figure 5 700 depicts a more detailed description of the comparison operation of operation 515 .

[0060] At operation 705, the method begins by comparing first pending write data specified for a register file of a first processor with second pending write data specified for a register file of a second processor. Figure 2 The pending write data and pending write address checkpoint circuitry 125 of the lockstep monitor compares the first pending write data specified by the register file 205 for the first processor 127 with the second pending write data specified by the register file 207 for the second processor 132 .

[0061] At operation 710, the method continues by comparing a first pending write address specified by a register file of a first processor with a second pending write address specified for a register file of a second processor. Figure 2 The pending write data and pending write address checkpoint circuitry 125 of the lockstep monitor compares the first pending write address specified by the register file 205 for the first processor 127 with the second pending write address specified by the register file 207 for the second processor 132 .

[0062] At operation 715, the method continues by comparing the first pending write data specified by the DTCM for the first processor with the second pending write data specified by the DTCM for the second processor. Figure 2 The pending write data and pending write address checkpoint circuitry 125 of the lockstep monitor compares the first pending write data specified by the DTCM 145 for the first processor 127 with the second pending write data specified by the DTCM 150 for the second processor 132 .

[0063] At operation 720, the method continues by comparing the first pending write address specified by the DTCM for the first processor with the second pending write address specified by the DTCM for the second processor. Figure 2The pending write data and pending write address checkpoint circuitry 125 of the lockstep monitor compares the first pending write address specified by the DTCM 145 for the first processor 127 with the second pending write address specified by the DTCM 150 for the second processor 132 .

[0064] At operation 725, the method continues by comparing the first pending instruction fetch address specified by the ITCM for the first processor with the second pending instruction fetch address specified by the ITCM for the second processor. Figure 2 The pending instruction fetch address checkpoint circuitry 120 of the lockstep monitor compares the first pending instruction fetch address specified by the ITCM 135 for the first processor 127 with the second pending instruction fetch address specified by the ITCM 140 for the second processor 132 .

[0065] At operation 730, the method continues by comparing the first pending instruction designated for the system bus with the second pending instruction designated for the system bus. Figure 2 , the pending system bus checkpoint circuitry 130 of the lockstep monitor compares the first pending instruction designated for the system bus 250 with the second pending instruction designated for the system bus 250 .

[0066] After the comparison operations 705, 710, 715, 720, 725, and 730 are completed, if a mismatch is detected at operation 735, the method proceeds to Figure 5 Step 520 of , wherein an automatic recovery operation of the processor is initiated. Alternatively, if no mismatch is detected at operation 735 , the method continues back to operation 705 , and the comparison operation is repeated until a mismatch requiring an automatic recovery operation is detected.

[0067] As such, the present invention provides an improved system and method for detecting mismatches in pipeline instructions of a processor operating in lockstep, which mismatch can cause the processor to terminate lockstep. The present invention also provides an improved system and method for performing automatic recovery of the processor without the need for a time-consuming and complex reset operation.

[0068] In one embodiment, portions of the lockstep circuitry may be implemented in an integrated circuit on a single semiconductor die. Alternatively, the integrated circuit may include multiple semiconductor dies electrically coupled together, such as a multi-chip module packaged in a single integrated circuit package.

[0069] In various embodiments, portions of the system of the present invention may be implemented in a field programmable gate array (FPGA) or an application specific integrated circuit (ASIC). It will be appreciated by those skilled in the art that the various functions of the circuit elements may also be implemented as processing steps in a software program. Such software may be used, for example, in a digital signal processor, a network processor, a microcontroller, or a general purpose computer.

[0070] Unless specifically stated otherwise as apparent from the discussion, it should be understood that throughout this specification, discussions utilizing terms such as "receiving," "determining," "generating," "limiting," "sending," "counting," "classifying," and the like may refer to actions and processes of a computer system or similar electronic computing device that manipulates and transforms data represented as physical (electronic) quantities within the computer system's registers and memories into other data similarly represented as physical quantities within the computer system's memories or registers or other such information storage, transmission, or display devices.

[0071] The present invention may be embodied on various computing platforms that perform actions in response to software-based instructions. The following provides a pre-foundation of information technology that can be used to implement the present invention.

[0072] The method of the present invention can be stored on a computer-readable medium, which can be a computer-readable signal medium or a computer-readable storage medium. A computer-readable storage medium can be, for example, but not limited to, an electronic, magnetic, optical, electromagnetic, infrared or semiconductor system, device or equipment, or any suitable combination of the foregoing. A more specific example (non-exhaustive list) of a computer-readable storage medium will include the following: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing. In the context of this document, a computer-readable storage medium can be any non-transient tangible medium that can contain or store a program for use by an instruction execution system, a device or equipment or used in combination with them.

[0073] A computer readable signal medium may include, for example, a data signal propagated in baseband or as part of a carrier wave, in which a computer readable program code is embodied. Such propagated signals may take any of a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer readable signal medium may be any computer readable medium that is not a computer readable storage medium, which may convey, propagate, or transmit a program for use by or in conjunction with an instruction execution system, device, or apparatus. However, as indicated above, due to the circuit statutory subject matter limitations, the claims of the present invention as software products are those embodied in non-transient software media such as computer hard drives, flash-RAM, optical disks, and the like.

[0074] Program code embodied on a computer readable medium may be transmitted using any appropriate medium, including but not limited to wireless, wireline, fiber optic cable, radio frequency, etc., or any suitable combination of the foregoing. Computer program code for performing operations of various aspects of the present invention may be written using any combination of one or more programming languages, including object-oriented programming languages ​​such as Java, C#, C++, Visual Basic, and conventional procedural programming languages ​​such as the "C" programming language or similar programming languages.

[0075] Hereinafter, various aspects of the present invention will be described with reference to the flow chart and / or block diagram of the method, device (system) and computer program product according to the embodiment of the present invention.It should be understood that each frame of the flow chart and / or block diagram and the combination of the frames in the flow chart and / or block diagram can be realized by computer program instructions.These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer or other programmable data processing device to produce a machine, so that the instruction executed via the processor of the computer or other programmable data processing device produces a device for realizing the function / action specified in one or more flow charts and / or block diagram frames.

[0076] These computer program instructions may also be stored in a computer-readable medium that can direct a computer, processor, or other programmable data processing device or other apparatus to function in a specific manner, so that the instructions stored in the computer-readable medium produce an article of manufacture including instructions for implementing the functions / actions specified in one or more flowcharts and / or block diagrams.

[0077] The computer program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device so that a series of operational steps are performed on the computer, other programmable apparatus, or other device to produce a computer-implemented process, whereby the instructions executed on the computer or other programmable apparatus provide a process for implementing the functions / actions specified in one or more flowcharts and / or block diagram blocks.

[0078] In addition, in order to discuss and understand the embodiments of the present invention, it should be understood that various terms are used by those skilled in the art to describe technology and methods. In addition, in this specification, in order to explain, many specific details are set forth in order to provide a thorough understanding of the present invention. However, it is obvious to those of ordinary skill in the art that the present invention can be put into practice without these specific details. In some cases, in order to avoid blurring the present invention, well-known structures and equipment are shown in block diagram form rather than in detail. These embodiments are described in full detail so that those of ordinary skill in the art can put into practice the present invention, and it should be understood that other embodiments can be utilized without departing from the scope of the present invention, and logical, mechanical, electrical and other changes can be made.

Claims

1. A method for automatic recovery in a lockstep processor, the method include: receiving, at a lockstep monitor, a first pending instruction generated by a first processor executing pipeline instructions, wherein the first pending instruction is to be executed by the first processor; receiving, at the lockstep monitor, a second pending instruction generated by a second processor executing the pipeline instructions in lockstep with the first processor, wherein the second pending instruction is to be executed by the second processor; comparing the first pending instruction at the lockstep monitor with the second pending instruction to detect a mismatch; as well as When a mismatch is detected, automatic recovery operations are performed at the first processor and the second processor to prevent the first processor from executing the first pending instruction and to prevent the second processor from executing the second pending instruction. 2 . The method of claim 1 , wherein the first pending instruction comprises one or more of: a pending instruction fetch address, a pending write address, and pending write data. 3 . The method of claim 1 , wherein the second pending instruction comprises one or more of: a pending instruction fetch address, a pending write address, and pending write data.

4. The method of claim 1 , wherein the first pending instruction specifies for one of: a register file of the first processor, a data tightly controlled memory (DTCM) of the first processor, an instruction tightly controlled memory (ITCM) of the first processor, and a system bus of the first processor.

5. The method of claim 1, wherein the second pending instruction specifies for one of: a register file of the second processor, a DTCM of the second processor, an ITCM of the second processor, and a system bus of the second processor.

6. The method of claim 1 , wherein the first pending instruction comprises a first pending instruction fetch address specified for the ITCM of the first processor, and the second pending instruction comprises a second pending instruction fetch address specified for the ITCM of the second processor, and wherein the comparison is used to detect a mismatch between the first pending instruction fetch address and the second pending instruction fetch address, and wherein performing the automatic recovery operation at the first processor comprises preventing the first processor from fetching instructions from the first pending instruction address at the ITCM of the first processor and preventing the second processor from fetching instructions from the second pending instruction address of the ITCM of the second processor.

7. The method of claim 1 , wherein the first pending instruction comprises a first pending write address specified for the DTCM of the first processor, and the second pending instruction comprises a second pending write address specified for the DTCM of the second processor, and wherein the comparison is used to detect a mismatch between the first pending write address and the second pending write address, and wherein performing the auto-recovery operation at the first processor comprises preventing the first processor from writing to the first pending write address at the DTCM of the first processor and preventing the second processor from writing to the second pending write address of the DTCM of the second processor.

8. The method of claim 1 , wherein the first pending instruction includes first pending write data specified for the DTCM of the first processor and the second pending instruction includes second pending write data specified for the DTCM of the second processor, and wherein the comparison is used to detect a mismatch between the first pending write data and the second pending write data, and wherein performing the auto-recovery operation at the first processor includes preventing the first processor from writing the first pending write data to the DTCM of the first processor and preventing the second processor from writing the second pending write data to the DTCM of the second processor.

9. The method of claim 1 , wherein the first pending instruction includes a first pending write address specified for a register file of the first processor, and the second pending instruction includes a second pending write address specified for a register file of the second processor, and wherein the comparison is used to detect a mismatch between the first pending write address and the second pending write address, and wherein performing the auto-recovery operation at the first processor includes preventing the first processor from writing to the first pending write address of the register file of the first processor and preventing the second processor from writing to the second pending write address of the register file of the second processor.

10. The method of claim 1 , wherein the first pending instruction includes first pending write data specified for a register file of the first processor and the second pending instruction includes second pending write data specified for a register file of the second processor, and wherein the comparison is used to detect a mismatch between the first pending write data and the second pending write data, and wherein performing the auto-recovery operation at the first processor includes preventing the first processor from writing the first pending write data to the register file of the first processor and preventing the second processor from writing the second pending write data to the register file of the second processor.

11. The method of claim 1 , wherein the first pending instruction is specified for a system bus coupled to the first processor and the second pending instruction is specified for the system bus, and wherein the comparison is used to detect a mismatch between the first pending instruction and the second pending instruction, and wherein performing the auto-recovery operation at the first processor comprises preventing the first processor from executing the first pending instruction and preventing the second processor from executing the second pending instruction.

12. The method of claim 1, wherein performing the automatic recovery operation at the first processor and the second processor further comprises: include: generating a flush and replay interrupt at the lockstep monitor; transmitting the flush and replay interrupts to the first processor and the second processor; flushing the pipeline instructions in response to the transmitted flush and replay interrupt; as well as The flushed pipeline instructions are re-executed in response to the transmitted flush and re-execute interrupt.

13. The method of claim 1, further comprising, after performing the auto-recovery operation at the first processor and the second processor, if the mismatch continues to exist, performing a reset of the first processor and the second processor.

14. The method according to claim 1, further comprising: include: storing the first pending instruction and the second pending instruction causing the mismatch at the lockstep monitor; After performing the auto-recovery operation at the first processor and the second processor, comparing the stored first pending instruction that caused the mismatch with a current first pending instruction, and comparing the stored second pending instruction that caused the mismatch with a current second pending instruction; as well as If the stored first pending instruction is different from the current first pending instruction and the stored second pending instruction is different from the current second pending instruction, a reset of the first processor and the second processor is performed.

15. A system for automatic recovery in a lockstep processor, the system include: a first processor, the first processor executing pipeline instructions; a second processor, the second processor executing the pipeline instructions in lockstep with the first processor; A lock-step monitor coupled to the first processor and the second processor, the lock-step monitor comprising: a checkpoint circuitry to receive a first pending instruction from the first processor and a second pending instruction from the second processor, the checkpoint circuitry to compare the first pending instruction to the second pending instruction to detect a mismatch, wherein the first pending instruction is to be executed by the first processor and the second pending instruction is to be executed by the second processor; and an automatic recovery circuit system coupled to the checkpoint circuit system, and when a mismatch is detected by the checkpoint circuit system, the automatic recovery circuit system is used to initiate an automatic recovery operation at the first processor and the second processor to prevent the first processor from executing the first pending instruction and to prevent the second processor from executing the second pending instruction.

16. The system of claim 15, wherein the first pending instruction comprises one or more of the following: a first pending instruction fetch address, a first pending write address, and first pending write data, and the second pending instruction comprises one or more of the following: a second pending instruction fetch address, a second pending write address, and second pending write data.

17. The system of claim 16, wherein the first processor comprises a register file, an instruction tightly controlled memory (ITCM), and a data tightly controlled memory (DTCM), and the second processor comprises a register file, an ITCM, and a DTCM, and wherein the first processor and the second processor are coupled to a system bus.

18. The system of claim 17, wherein the checkpoint circuitry further include: a pending instruction fetch address checkpoint circuitry configured to receive a first pending instruction fetch address specified by the ITCM for the first processor, receive a second pending instruction fetch address specified by the ITCM for the second processor, and compare the first pending instruction fetch address with the second pending instruction fetch address to detect a mismatch; a pending write data and pending write address checkpoint circuitry configured to: receive first pending write data specified by the register or the DTCM for a first processor, receive second pending write data specified by the register or the DTCM for a second processor, and compare the first pending write data with the second pending write data to detect a mismatch; The pending write data and pending write address checkpoint circuitry is further configured to: receive a first pending write address specified by the register or the DTCM for a first processor, receive a second pending write address specified by the register or the DTCM for a second processor, and compare the first pending write address to the second pending write address to detect a mismatch; and A pending system bus checkpoint circuit system is configured to receive a first pending instruction specified for the system bus, receive a second pending instruction specified for the system bus, and compare the first pending instruction with the second pending instruction to detect a mismatch.

19. The system of claim 15, wherein the automatic recovery circuitry is further configured to generate a flush and replay interrupt and transmit the flush and replay interrupt to the first processor and the second processor.

20. The system of claim 19, wherein the first processor and the second processor are configured to: receiving said flush and replay interrupt from said automatic recovery circuitry; flushing the pipeline instructions in response to the interrupt; and The pipeline instruction is re-executed.

Citation Information

Patent Citations

  • Error recovery for intra-core lockstep mode

    CN111164578A

  • Multiprocessor with pair-wise high reliability mode, and method therefore

    US20020073357A1