Distributed Redundancy Management for Digital Control Systems
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional digital control systems face challenges in rapidly recovering from soft faults, particularly in aerospace and other critical applications, where simultaneous failures of processing units can lead to system degradation or failure, and existing redundancy methods require synchronization and may not ensure quick recovery.
Innovation Solution
A distributed redundancy management system that allows for rapid recovery of processing units and actuator control units through asynchronous operation, using command blending, equalization, and internal/external monitoring to isolate faults and restore system resources without synchronization, enabling quick fault detection and recovery within a single computing frame.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If conventional redundancy methods with multiple processing units are used, then system reliability is improved, but system complexity increases and recovery time is extended due to synchronization requirements
Solution Approach 1:
The system segments redundancy management into independent processing units, each capable of autonomous fault detection and recovery. Each unit maintains its own state variables and can operate independently, eliminating the need for complex centralized synchronization while preserving reliability through distributed redundancy.
Solution Approach 2:
The system performs preliminary actions by pre-storing state variables in protected memory areas before faults occur. When a soft fault is detected, the system can immediately restore from these pre-stored states without requiring complex recovery procedures, thereby improving reliability while maintaining simplicity.
2Reliability
If dissimilar computational redundancy is implemented, then generic fault prevention is improved, but device complexity and operational difficulty increase
Solution Approach 1:
The system applies local quality by implementing dissimilar computational redundancy only where needed for generic fault prevention, while maintaining uniformity in the core recovery mechanism. Each processing unit can use different computational approaches locally, but the overall system maintains a unified simple recovery protocol through protected memory restoration.
3Measurement precision
If synchronized operation of redundant elements is required, then measurement precision of system state is improved, but speed of recovery is reduced
Solution Approach 1:
The system performs preliminary action by continuously maintaining accurate state variables in protected memory areas during normal operation. When a fault occurs, recovery is achieved by simply restoring from these pre-maintained states, eliminating the need for time-consuming synchronized re-measurement while preserving monitoring precision through the protected memory mechanism.
4Reliability
If fault isolation and recovery procedures are implemented, then system reliability is improved, but loss of time during recovery occurs
Solution Approach 1:
The system performs preliminary action by continuously updating and protecting state variables in memory areas designated for fault recovery. When a soft fault is detected, the system can immediately restore from these pre-prepared states within a single computing frame, minimizing recovery time while maintaining reliable fault isolation capability.
Solution Approach 2:
The system skips complex recovery procedures by directly restoring from protected memory states when faults are detected. This rushing through of the recovery process eliminates time-consuming diagnostic and restoration steps, achieving rapid recovery within a single computing frame while maintaining fault isolation through the protected memory mechanism.
Data Source
AI summary
A method and system for redundancy management is provided for a distributed and recoverable digital control system. The method uses unique redundancy management techniques to achieve recovery and restoration of redundant elements to full operation in an asynchronous environment. The system includes a first computing unit comprising a pair of redundant computational lanes for generating redundant control commands. One or more internal monitors detect data errors in the control commands, and provide a recovery trigger to the first computing unit. A second redundant computing unit provides the same features as the first computing unit. A first actuator control unit is configured to provide blending and monitoring of the control commands from the first and second computing units, and to provide a recovery trigger to each of the first and second computing units. A second actuator control unit provides the same features as the first actuator control unit.


