Graceful Degradation in Redundant Processing Systems
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
In safety-critical applications, once an error is detected in processor lockstep and/or multi-processor voting algorithms, there is no mechanism to allow the system to continue functioning with reduced error detection capabilities, leading to a complete failure rather than graceful degradation.
Innovation Solution
Implementing a redundancy protocol that adjusts based on detected errors, aggregating errors across multiple cores, and disabling misbehaving cores to transition from triple-core redundancy to dual-core or single-core lockstep modes, using error correction codes and failure tracking to ensure continued operation with progressively decreased safety margins.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If processor lockstep and multi-processor voting algorithms are employed to detect program flow execution issues, then error detection capability is improved, but system functionality stops completely once an error is detected
Solution Approach 1:
The system dynamically transitions between different redundancy modes (triple-modular redundancy, dual-processor lockstep, single-processor operation) based on detected error conditions. This allows the error detection capability and system configuration to adapt dynamically, maintaining functionality while adjusting safety margins as errors are detected and processed.
Solution Approach 2:
The system changes operational parameters by transitioning between discrete redundancy states. When errors are detected, the system parameter changes from high-redundancy mode to lower-redundancy modes, progressively adjusting the level of error detection and processing capability while maintaining continuous operation.
2Reliability
If redundant processing elements are used to ensure safety, then reliability is improved, but device complexity increases
Solution Approach 1:
The redundant processing system is segmented into discrete operational modes (triple-modular redundancy mode, dual-processor lockstep mode, single-processor mode). Each mode represents a distinct configuration that can be independently activated, allowing the system to provide high reliability when needed while reducing complexity when error conditions permit.
Solution Approach 2:
The processing elements serve multiple functions depending on the operational mode. The same physical processors can operate in full redundant configuration for maximum safety, transition to reduced redundancy for continued operation, or function as individual standalone units, making the hardware universally applicable across different safety requirements.
Data Source
AI summary
An apparatus and method for redundant data processing with graceful degrading functionality. For example, one embodiment of an apparatus comprises: three processing elements operable in a first redundancy mode, the three processing elements to execute a same sequence of instructions to produce three corresponding results; detection circuitry to detect when any one processing element of the three processing elements produces a different result from the other two processing elements of the three processing elements; tracking circuitry to associate an error with the one processing element when it produces the different result from the other two processing elements, wherein if an error threshold is reached for the one processing element, the other two processing elements are to operate in a second redundancy mode excluding the one processing element.


