Non-Voting Redundant Controller Fault Tolerance
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing redundant computing and control systems face challenges in maintaining fault tolerance when both the master and hot standby components fail, particularly in dual modular redundancy (DMR) architectures, and the need for voting systems can lead to issues when voters fail.
Innovation Solution
A non-voting redundant system (NVR) is introduced, which includes three microcontrollers where one remains active, and upon detection of faults in both the primary and secondary controllers, the third 'smart reserve' controller is activated, switching the state of the faulty controllers to idle or shutdown for repair, eliminating the need for voting systems.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If DMR architecture with master and hot standby configuration is used, then fault tolerance for single fault is achieved, but complete system failure occurs when both master and hot standby fail
Solution Approach 1:
The system segments the redundant configuration into three distinct controller modules (primary, secondary, and tertiary controllers), each capable of independent operation. This segmentation allows the system to distribute the fault tolerance function across multiple independent units, enabling the tertiary controller to take over when both primary and secondary controllers fail, thus resolving the limitation of DMR architecture.
Solution Approach 2:
The system implements preliminary action by pre-configuring the tertiary controller in a standby state with all necessary operational capabilities before any fault occurs. The controller modules are pre-programmed with fault detection and takeover logic, allowing seamless transition of control without service interruption when failures occur, eliminating the need for complex voting systems.
2Reliability
If TMR architecture with majority voting system is used, then fault detection capability is improved, but system failure occurs when voters fail or cannot reach majority consensus
Solution Approach 1:
The invention extracts and removes the voting system component from the redundant controller architecture. Instead of using a separate voting mechanism to determine system state, each controller module independently monitors its own operational status and automatically initiates takeover procedures when faults are detected, eliminating the complexity and potential failure points of voting systems while maintaining fault detection capability.
Solution Approach 2:
Each controller module is equipped with self-service capabilities including autonomous fault detection, self-diagnosis, and automatic takeover initiation. The controllers monitor their own operational parameters and can independently determine when they or other controllers have failed, eliminating the need for external voting mechanisms to assess system health.
3Reliability
If voting system is implemented to detect failures, then fault masking capability is improved, but system reliability decreases when voters fail
Solution Approach 1:
The voting mechanism is completely extracted from the system architecture. Instead of using voters to mask faults and determine system state, the invention employs direct fault detection by each controller module and automatic takeover protocols, eliminating the intermediate voting layer that could itself fail and compromising overall system reliability.
Solution Approach 2:
The system implements direct feedback loops where each controller continuously monitors its own operational status and the status of other controllers. When a controller detects a fault in itself or another controller, it immediately initiates appropriate corrective action including self-isolation or takeover, providing reliable fault masking without the complexity of voting systems.
Data Source
AI summary
Systems and methods for resolving fault detection in a control system is provided. The system includes an I/O module operably connected to a first, second, and third microcontroller for transmitting data. The first microcontroller is in an active state, i.e., in control, while the remaining controllers are in an idle state. The system further includes an event generator for generating an event indicative of a fault occurrence, and a means for detecting a fault event. The system also includes a means for reassigning a controller, wherein upon detection of a fault event in both the first and second controllers, the means for reassigning a controller changes the state of the third controller to active, leaving the remaining controllers idle or in a shutdown state, thereby effectively assigning control from the first controller to the third controller.


