Redundant Control Computer System with Switching Matrix for Fault Tolerance
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing dual-redundant processor systems face challenges in maintaining high availability and efficiency, as they often require high redundancy levels, leading to idle processors and increased costs, especially in safety-critical applications like automotive braking systems, where non-redundant components are common and fault tolerance is essential.
Innovation Solution
A control computer system with at least two processor pairs, each with two redundant processors, comparison units for error detection, and a switching matrix that allows selective access to memory and peripheral units, enabling error-free processor pairs to take over functions from faulty ones, while maintaining secure connection with non-redundant components.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If triple modular redundancy (TMR) of processors and shared memory is used, then fault tolerance is improved, but system cost and complexity increase significantly
Solution Approach 1:
The system divides processors into multiple pairs, where each pair operates independently. This segmentation allows the system to achieve fault tolerance through pairwise comparison rather than requiring full TMR of all processors, reducing overall system complexity while maintaining reliability.
Solution Approach 2:
The system dynamically switches between different processor pairs based on error detection. When a lockstep error is detected in one pair, the system can activate another processor pair, providing adaptive fault tolerance without requiring all processors to be constantly redundant, thus reducing complexity.
2Reliability
If multiple processor pairs are dedicated to single tasks, then system reliability is improved, but processor utilization efficiency deteriorates
Solution Approach 1:
Processor pairs can be dynamically assigned to different tasks based on system needs and error states. A processor pair that has completed its current task can be reassigned to another task, allowing processors to serve multiple functions rather than being permanently dedicated, thus improving utilization efficiency while maintaining reliability through the availability of multiple pairs.
Solution Approach 2:
When a lockstep error occurs in a processor pair, the system can discard the current task execution in that pair and recover by transferring the task to another processor pair. This allows the affected processor pair to be reset and reused for future tasks, improving overall processor utilization while maintaining system availability.
3Reliability
If processor pairs are switched upon error detection, then system availability is improved, but the number of idle processors increases
Solution Approach 1:
The system maintains continuous useful action by having multiple processor pairs ready to execute tasks. When one pair encounters an error, another pair can immediately take over without interruption to critical functions, ensuring system availability while keeping all processor pairs potentially productive rather than having dedicated idle backups.
Solution Approach 2:
The system changes the operational state of processors dynamically based on error conditions. Processors transition between active, standby, and recovered states rather than remaining permanently idle, optimizing the balance between availability and resource utilization by adjusting the number of active processors based on actual system needs.
Data Source
Figure 1
Figure 2
Figure 3
AI summary
The invention relates to a control computer system. The control computer system comprises at least two modules (1001, 1002, 1003, 1004) configured to be redundant to each other; at least one comparator unit (1011, 1012) for monitoring the synchronization state of the at least two redundant modules (1001, 1002, 1003, 1004) and for detecting a synchronization fault; and at least one peripheral unit (1030, 1031,..., 1038). The control computer system further comprises at least one switch matrix (1013) set up for allowing or blocking access to the at least two redundant modules (1001, 1002, 1003, 1004) or access by the at least two redundant modules to the peripheral unit (1030, 1031,..., 1038). A fault handling unit (1080) is set up for receiving signals of the at least one comparator unit (1011, 1012) and to actuate the at least one switching matrix (1013) in order to optionally completely or selectively prevent access to the at least two redundant modules or access by the at least two redundant modules to the peripheral unit.