Delayed Lock Step Processor Fault Recovery

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Fault-tolerant computing systems face challenges in detecting and recovering from processing faults, such as hardware errors or environmental factors, which can lead to system state corruption and prolonged recovery times.

Innovation Solution

A fault-tolerant computing system is designed with a secondary processor executing in delayed lock step with a primary processor, using comparators in data and writeback paths to detect faults before writing invalid data, and incorporating triple module redundancy in store and writeback paths to ensure data integrity, allowing for aborting execution and preventing system state corruption.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If traditional fault detection methods are used, then system safety is maintained, but recovery time is prolonged and system state corruption occurs

Engineering Contradiction:
Improvefault detection reliabilityVSAvoidrecovery time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent applies preliminary action by delaying the writeback operation until after fault detection is complete. The writeback path includes delay stages that hold data until the comparator verifies data integrity, preventing corrupted data from being written to memory before the fault is detected. This preliminary verification action enables fast recovery by avoiding system state corruption in the first place.

Inventive Principle:
Principle #10Preliminary action

2Reliability

If delayed lock step execution is implemented, then fault detection capability is improved, but device complexity increases

Engineering Contradiction:
Improvefault detection capabilityVSAvoidprocessor system complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent uses copying by maintaining a secondary processor that replicates the primary processor's execution. The secondary processor executes the same instructions from a common program store, and its results are compared with the primary processor's results. This copying approach provides fault detection capability while keeping each individual processor unit relatively simple, avoiding the need for complex fault detection circuitry within a single processor.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent introduces an intermediary comparator component that mediates between the primary and secondary processors. The comparator receives results from both processors and determines whether they match, providing fault detection without requiring complex integration between the processors. This intermediary approach simplifies the overall system architecture by separating the comparison function from the processing functions.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Reliability

If triple module redundancy is used in store data and writeback paths, then data integrity is ensured, but device complexity and resource usage increase

Engineering Contradiction:
Improvedata integrityVSAvoidredundancy circuit complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent applies segmentation by dividing the redundancy implementation into distinct segments: the primary processor path, the secondary processor path, and the comparator. Instead of implementing complex TMR within each processor, the redundancy is segmented across parallel execution paths. This segmentation allows each segment to remain relatively simple while achieving overall data integrity through the comparison of results across segments.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS11928475B2Fast recovery for dual core lock step
Publication Date: 2024.03.12 CEREMORPHIC INC
  • US11928475B2 patent drawing
  • US11928475B2 patent drawing
  • US11928475B2 patent drawing

AI summary

An exemplary fault-tolerant computing system comprises a secondary processor configured to execute in delayed lock step with a primary processor from a common program store, comparators in the store data and writeback paths to detect a fault based on comparing primary and secondary processor states, and a writeback path delay permitting aborting execution when a fault is detected, before writeback of invalid data. The secondary processor execution and the primary processor store data and writeback may be delayed a predetermined number of cycles, permitting fault detection before writing invalid data. Store data and writeback paths may include triple module redundancy configured to pass only majority data through the store data and writeback path delay stages. Some implementations may forward data from the store data path delay stages to the writeback stage or memory if the load data address matches the address of data in a store data path delay stage.