Delayed Lock Step Processor Fault Recovery
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Fault-tolerant computing systems face challenges in detecting and recovering from processing faults, such as hardware errors or environmental factors, which can lead to system state corruption and prolonged recovery times.
Innovation Solution
A fault-tolerant computing system is designed with a secondary processor executing in delayed lock step with a primary processor, using comparators in data and writeback paths to detect faults before writing invalid data, and incorporating triple module redundancy in store and writeback paths to ensure data integrity, allowing for aborting execution and preventing system state corruption.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If traditional fault detection methods are used, then system safety is maintained, but recovery time is prolonged and system state corruption occurs
Solution Approach 1:
The patent applies preliminary action by delaying the writeback operation until after fault detection is complete. The writeback path includes delay stages that hold data until the comparator verifies data integrity, preventing corrupted data from being written to memory before the fault is detected. This preliminary verification action enables fast recovery by avoiding system state corruption in the first place.
2Reliability
If delayed lock step execution is implemented, then fault detection capability is improved, but device complexity increases
Solution Approach 1:
The patent uses copying by maintaining a secondary processor that replicates the primary processor's execution. The secondary processor executes the same instructions from a common program store, and its results are compared with the primary processor's results. This copying approach provides fault detection capability while keeping each individual processor unit relatively simple, avoiding the need for complex fault detection circuitry within a single processor.
Solution Approach 2:
The patent introduces an intermediary comparator component that mediates between the primary and secondary processors. The comparator receives results from both processors and determines whether they match, providing fault detection without requiring complex integration between the processors. This intermediary approach simplifies the overall system architecture by separating the comparison function from the processing functions.
3Reliability
If triple module redundancy is used in store data and writeback paths, then data integrity is ensured, but device complexity and resource usage increase
Solution Approach 1:
The patent applies segmentation by dividing the redundancy implementation into distinct segments: the primary processor path, the secondary processor path, and the comparator. Instead of implementing complex TMR within each processor, the redundancy is segmented across parallel execution paths. This segmentation allows each segment to remain relatively simple while achieving overall data integrity through the comparison of results across segments.
Data Source
AI summary
An exemplary fault-tolerant computing system comprises a secondary processor configured to execute in delayed lock step with a primary processor from a common program store, comparators in the store data and writeback paths to detect a fault based on comparing primary and secondary processor states, and a writeback path delay permitting aborting execution when a fault is detected, before writeback of invalid data. The secondary processor execution and the primary processor store data and writeback may be delayed a predetermined number of cycles, permitting fault detection before writing invalid data. Store data and writeback paths may include triple module redundancy configured to pass only majority data through the store data and writeback path delay stages. Some implementations may forward data from the store data path delay stages to the writeback stage or memory if the load data address matches the address of data in a store data path delay stage.


