Automation Controller Fault Tolerance With Selective Redundancy
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Fault-tolerant computing systems require significant hardware resources, leading to high costs, as triple modular redundant (TMR) and dual-dual configurations necessitate three to four times the computational capacity of non-fault-tolerant systems, making them costly and resource-intensive for applications like robotic control where reduced capacity operation may be acceptable.
Innovation Solution
A reduced-cost fault-tolerant system for automation controllers is implemented by dividing computing phases into fail operational and fail degraded modes, allowing for efficient allocation of computational resources, where critical phases have full redundancy, while less critical phases can operate with reduced redundancy, thereby minimizing hardware requirements.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If triple modular redundant (TMR) or dual-dual configurations are used to achieve fault tolerance, then system reliability is improved, but hardware resources and computational capacity increase by three to four times
Solution Approach 1:
The patent segments the computing phases into distinct categories: safety-critical phases requiring full fault tolerance, phases allowing reduced capacity operation, and non-critical phases. This segmentation enables selective application of fault tolerance mechanisms only where necessary, rather than uniformly across all computing operations, thereby reducing overall hardware requirements while maintaining reliability for critical functions.
Solution Approach 2:
The patent applies local quality by implementing different fault tolerance strategies tailored to specific computing phases. Safety-critical phases receive full TMR or dual-dual protection, while other phases use reduced redundancy or no redundancy. This localized approach ensures that hardware resources are concentrated where they are most needed (in critical phases) rather than being uniformly distributed, resolving the contradiction between reliability and resource consumption.
2Reliability
If full redundant configurations (TMR or dual-dual) are implemented, then fault detection and continued operation are ensured, but system cost and computational bandwidth requirements increase significantly
Solution Approach 1:
The patent divides computing tasks into segments based on their criticality level. Safety-critical segments are processed with full redundant configurations for comprehensive fault detection, while non-critical segments use simpler processing without full redundancy. This segmentation allows the system to maintain fault detection capability for essential operations while avoiding the prohibitive costs of applying full redundancy to all operations.
Solution Approach 2:
The patent applies partial action by implementing fault tolerance only to the extent necessary for each computing phase. Rather than applying excessive redundancy uniformly, the system applies partial redundancy selectively - full redundancy for safety-critical phases, reduced redundancy for other phases, and no redundancy for non-critical phases. This resolves the contradiction by avoiding excessive hardware investment while maintaining adequate fault detection where required.
Data Source
AI summary
Fault tolerance for an automation controller for a machine is provided. A first portion of phases of the automation controller may be processed with fail operational protection, in which a failure of one of the computers used for the first portion still permits full operational functionality in the machine. The remaining portion of the phases are processed with fail degraded protection, in which a failure of a computer used for the remaining portion permits continued operation but with one or more constraints, as compared to the fail operational portions.


