Ground State Checking for Distributed System Robustness
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Large real-time distributed computer systems face increased disturbances due to high complexity and transient hardware errors, leading to non-deterministic behavior and reduced reliability, necessitating a method to ensure sensible system behavior despite errors.
Innovation Solution
Implementing a method where a processing component periodically sends a Ground State (GS) message to a ground-state checking component, which corrects errors and restarts the processing component using a corrected GS message at the next restart time, leveraging Fault Containment Units and resilience principles to maintain system robustness.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If the number of components is increased to handle complex real-time tasks, then system functionality and processing capability are improved, but transient hardware errors and system reliability deteriorate
Solution Approach 1:
The system performs preliminary actions by periodically saving ground state information before errors can occur. The ground state checking component continuously monitors and stores valid system states, enabling recovery without losing functionality when transient errors occur in complex systems with many components.
Solution Approach 2:
The invention creates a copy of the essential system state (ground state) that can be used for recovery. By maintaining a validated ground state message that can be restored, the system preserves its functional capability even when transient errors affect individual components, thus resolving the reliability issue while maintaining adaptability.
2Quantity of substance
If transistor size is reduced to increase component density, then system integration is improved, but transient hardware errors increase
Solution Approach 1:
The ground state checking component provides continuous feedback validation of system states. By checking incoming ground state messages in both value domain and time domain, the system detects and corrects transient errors that occur in miniaturized components, maintaining reliability despite increased component density.
Solution Approach 2:
The system performs preliminary validation of ground state messages before they can cause system-wide errors. By checking messages in advance in value and time domains, transient errors in miniaturized components are caught and corrected before they can propagate, allowing high component density without sacrificing reliability.
3Measurement precision
If Ground State messages are checked in both value domain and time domain, then error detection capability is improved, but processing overhead increases
Solution Approach 1:
The ground state checking component performs partial checks by validating only critical aspects of ground state messages in value and time domains. This selective validation approach provides sufficient error detection capability without requiring exhaustive analysis of all message parameters, thus balancing detection precision with acceptable processing overhead.
4Reliability
If periodic restarts are implemented to correct errors, then system robustness is improved, but system productivity decreases
Solution Approach 1:
The system implements periodic ground state checking and validation at defined restart times. This periodic action ensures robustness by regularly validating system states and enabling recovery when errors occur, while the structured timing minimizes disruption to overall system productivity by concentrating restart operations at predetermined intervals.
Data Source
Figure 1~2
AI summary
The invention relates to a method for increasing the robustness of a distributed computer system, comprising a number of components (110, 120, 130), wherein each component (110, 120, 130) can transmit messages to the other components via a communication system (100). According to the invention, at least one of the components (110, 120, 130) is a processing component (HO), and at least one of the components (110, 120, 130) is a ground state checking component (120), wherein the at least one processing component (110) periodically, at a periodically recurring restart time, transmits a ground state (GS) message, which comprises a ground state of the processing component (110) relevant immediately before the time of transmission, to the at least one ground state checking component (120), and wherein the ground state checking component (120) checks the value range and time range of the incoming ground state message, and wherein in the event a fault is detected in the ground state message the ground state checking component (120) corrects the fault in the ground state and before the next restart time transmits the corrected ground state in a corrected ground state message to that processing component (110) from which the faulty ground stage message originated, and wherein upon receipt of the corrected GS message by the processing component (110) said processing component (110) at the next restart time performs a restart, using the corrected ground state present in the corrected GS message.