Fault-Tolerant Decision Architecture for Byzantine Error Containment
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Cyber-physical systems, such as autonomous vehicles, face critical challenges in detecting and managing Byzantine errors caused by hardware failures, software design errors, or intrusions, which can lead to catastrophic failures due to the complexity of real-time computer systems and the impossibility of eliminating all design errors in complex software systems.
Innovation Solution
A distributed real-time computer system architecture comprising multiple largely independent subsystems with diverse software and hardware, a Fault-Tolerant Decision Subsystem running simple software on fault-tolerant hardware, and a time server for synchronization, which ensures detection and containment of Byzantine errors through fault containment units and critical event handling mechanisms.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Extent of automation
If complex software systems are used for autonomous control, then functionality and automation capability are improved, but reliability deteriorates due to undetected design errors and Byzantine faults
Solution Approach 1:
The system is divided into multiple independent subsystems (normal processing subsystem, monitor subsystem, critical event handling subsystem), each with its own fault containment unit. This segmentation allows the system to isolate and detect Byzantine faults in individual subsystems without compromising the entire system, thereby maintaining high automation capability while improving reliability.
Solution Approach 2:
A monitor subsystem is introduced as an intermediary that independently verifies the correctness of the normal processing subsystem. The monitor subsystem checks whether target values and environmental models are consistent, detecting Byzantine faults before they affect system operation. This intermediary mechanism enables complex autonomous control while ensuring reliability through continuous verification.
2Device complexity
If hardware failures and software defects are allowed in complex systems, then system complexity and functionality are improved, but safety deteriorates due to catastrophic failures
Solution Approach 1:
The system implements beforehand cushioning by pre-establishing fault containment units and monitor subsystems that detect and contain Byzantine faults before they can cause catastrophic failures. The monitor subsystem continuously checks for inconsistencies in target values and environmental models, providing a safety buffer that allows complex systems to operate with hardware and software components that may fail, without resulting in catastrophic outcomes.
3Ease of operation
If intrusion detection mechanisms are bypassed, then system accessibility and ease of operation are improved, but security deteriorates due to Byzantine faults from intrusions
Solution Approach 1:
The monitor subsystem implements continuous feedback by independently calculating environmental models and comparing them with those from the normal processing subsystem. This feedback mechanism detects Byzantine faults caused by intrusions that bypass traditional intrusion detection mechanisms. The system maintains ease of operation by allowing direct access while ensuring security integrity through this independent verification feedback loop.
Data Source
Figure 1
Figure 2
Figure 3
AI summary
The invention lies in the field of computer technology. It describes the architecture of a safe automation system and a method for the safe autonomous operation of a technical device, e.g., a technical system such as a robot or a vehicle, in particular a motor vehicle. The architecture disclosed here solves the problem that any possible Byzantine fault in one of the complex subsystems of a distributed real-time computer system, regardless of whether the fault was triggered by a random hardware failure, a design flaw in the software, or an intrusion, must be detected and controlled in such a way that no safety-relevant incident occurs. The architecture consists of four largely independent subsystems, which are hierarchically arranged and each form an encapsulated Fault Containment Unit (FCU).At the top of the hierarchy is a safe subsystem, the Fault-Tolerant Decision Subsystem, which executes simple software on fault-tolerant hardware. The other three subsystems are unsafe because they contain complex software executed on non-fault-tolerant hardware. Experience has shown that it is difficult to find all design flaws in a complex software system and prevent intrusions. Due to the redundancy and diversity inherent in the architecture, any error—even a Byzantine one—in an unsafe subsystem is masked to such an extent that no safety-critical failure can occur.