Processing Circuitry High Resilience Mode for Fault Tolerance
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Data processing systems are vulnerable to faults caused by external events like radiation, leading to errors that consume significant processing time and resources, and may require invasive procedures such as full system reboots, impacting availability.
Innovation Solution
Implementing a high resilience mode of operation for processing circuitry by modifying the usage of components such as fetch circuitry, buffer structures, and error detection schemes to reduce the likelihood of faults resulting in errors, while executing critical code sequences, thereby enhancing fault tolerance without significant performance impact.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If error detection and handling mechanisms are implemented, then system reliability is improved, but processing time and resource consumption increase
Solution Approach 1:
The system dynamically switches between default mode and high resilience mode based on the criticality of the code sequence being executed. This allows the error detection and handling mechanisms to be activated only when necessary, rather than continuously, thereby reducing overall processing time while maintaining reliability for critical operations.
Solution Approach 2:
The high resilience mode is applied selectively to specific critical code sequences rather than uniformly across all code. This localized application ensures that error protection is provided where it is most needed while avoiding the performance overhead in non-critical sections.
2Reliability
If invasive error handling procedures such as full system reboot are performed, then system reliability is restored, but system availability deteriorates
Solution Approach 1:
The system performs preliminary error detection and handling during the execution of critical code sequences in high resilience mode. By detecting and handling errors early, before they propagate through the system, the need for invasive procedures like full system reboots is avoided, thereby maintaining system availability.
Solution Approach 2:
The high resilience mode acts as a protective cushion before errors can cause system failure. By modifying component usage and enabling enhanced error detection beforehand during critical operations, the system prevents errors from escalating to a point where invasive recovery procedures are necessary.
3Reliability
If component usage is modified to increase fault resilience, then reliability is improved, but device complexity increases
Solution Approach 1:
The modification of component usage is dynamic rather than static. The system adjusts component behavior based on the execution context (critical vs. non-critical code sequences), allowing the same hardware to operate in different modes without requiring additional physical components or permanent structural changes.
Data Source
AI summary
Aspects of the present disclosure relate to an apparatus comprising processing circuitry to execute a plurality of code sequences, and configuration storage to store mode control data for the processing circuitry. When the processing circuitry is executing one of said plurality of code sequences, the mode control data is set so as to identify a high resilience mode of operation of the processing circuitry where usage of one or more components of the processing circuitry is modified so as to increase resilience of the processing circuitry to faults relative to a default mode of operation of the processing circuitry.


