Graceful Degradation in Redundant Processing Systems

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

In safety-critical applications, once an error is detected in processor lockstep and/or multi-processor voting algorithms, there is no mechanism to allow the system to continue functioning with reduced error detection capabilities, leading to a complete failure rather than graceful degradation.

Innovation Solution

Implementing a redundancy protocol that adjusts based on detected errors, aggregating errors across multiple cores, and disabling misbehaving cores to transition from triple-core redundancy to dual-core or single-core lockstep modes, using error correction codes and failure tracking to ensure continued operation with progressively decreased safety margins.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If processor lockstep and multi-processor voting algorithms are employed to detect program flow execution issues, then error detection capability is improved, but system functionality stops completely once an error is detected

Engineering Contradiction:
Improveerror detection capabilityVSAvoidsystem continuity
Core Design Contradiction:
Measurement precisionVSReliability

Solution Approach 1:

The system dynamically transitions between different redundancy modes (triple-modular redundancy, dual-processor lockstep, single-processor operation) based on detected error conditions. This allows the error detection capability and system configuration to adapt dynamically, maintaining functionality while adjusting safety margins as errors are detected and processed.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The system changes operational parameters by transitioning between discrete redundancy states. When errors are detected, the system parameter changes from high-redundancy mode to lower-redundancy modes, progressively adjusting the level of error detection and processing capability while maintaining continuous operation.

Inventive Principle:
Principle #35Parameter changes

2Reliability

If redundant processing elements are used to ensure safety, then reliability is improved, but device complexity increases

Engineering Contradiction:
Improvesafety critical operationVSAvoidredundant processing structure
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The redundant processing system is segmented into discrete operational modes (triple-modular redundancy mode, dual-processor lockstep mode, single-processor mode). Each mode represents a distinct configuration that can be independently activated, allowing the system to provide high reliability when needed while reducing complexity when error conditions permit.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The processing elements serve multiple functions depending on the operational mode. The same physical processors can operate in full redundant configuration for maximum safety, transition to reduced redundancy for continued operation, or function as individual standalone units, making the hardware universally applicable across different safety requirements.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS20250004892A1Apparatus and method for graceful degradation of redundant processing
Publication Date: 2025.01.02 ALTERA CORP
  • US20250004892A1 patent drawing
  • US20250004892A1 patent drawing
  • US20250004892A1 patent drawing

AI summary

An apparatus and method for redundant data processing with graceful degrading functionality. For example, one embodiment of an apparatus comprises: three processing elements operable in a first redundancy mode, the three processing elements to execute a same sequence of instructions to produce three corresponding results; detection circuitry to detect when any one processing element of the three processing elements produces a different result from the other two processing elements of the three processing elements; tracking circuitry to associate an error with the one processing element when it produces the different result from the other two processing elements, wherein if an error threshold is reached for the one processing element, the other two processing elements are to operate in a second redundancy mode excluding the one processing element.