Fault-Tolerant Automation System With Autonomous Restart
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing automation systems face high maintenance costs and complexity, especially after errors occur, with troubleshooting often requiring specialist intervention and being costly, which is a challenge in the context of the Internet of Things (IoT) applications where reliability and maintainability are critical.
Innovation Solution
A fault-tolerant automation system with fail-silent Fault Containment Units (FCUs) that autonomously restart and monitor message traffic to identify and replace faulty components, allowing a layperson to identify and replace faulty units using indicators like control lamps, ensuring continuous operation with minimal maintenance.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If traditional automation systems are used with centralized control architecture, then system functionality is achieved, but maintenance costs increase and system reliability decreases when errors occur
Solution Approach 1:
The patent divides the centralized control system into multiple autonomous control units (active and passive) that operate independently. Each unit can function autonomously, and the system segments the control functionality across multiple nodes rather than relying on a single centralized controller, thereby improving reliability while maintaining manageable complexity through modular design
Solution Approach 2:
The patent changes the operational state parameters of the control units by introducing fail-silent modes where units can transition between active, passive, and error states. This parameter change allows the system to maintain reliability by having units switch states based on operational conditions without increasing maintenance complexity
2Productivity
If fail-silent Fault Containment Units are implemented with autonomous restart capability, then system availability improves, but the complexity of fault detection and identification increases
Solution Approach 1:
The patent implements feedback mechanisms where control units exchange status information and monitor each other's operational states. The active and passive units provide feedback about system health, enabling automatic fault detection and identification without increasing overall system complexity, as the feedback is integrated into the existing communication protocol
Solution Approach 2:
The control units possess self-service capabilities through autonomous restart functionality and self-diagnosis. When a fault occurs, the unit can automatically attempt recovery and identify its own status, reducing the burden on external monitoring systems and maintaining simple fault detection mechanisms while improving availability
3Reliability
If redundant control units are deployed with automatic failover, then system reliability improves, but the cost of the control system increases
Solution Approach 1:
The patent makes control units universal by designing them to perform multiple functions - they can operate as either active or passive units depending on system needs. This multi-functionality allows the same hardware design to serve different roles, reducing the need for specialized redundant components and lowering overall system costs while maintaining reliability through role flexibility
Data Source
Figure 1~2
AI summary
The invention relates to a fault-tolerant, serviceable automation system comprising two central computers, a process peripheral area and gateway computers, wherein the central computers and the gateway computers are fail-silent FCUs and represent autonomous exchange units, and the central computers and the gateway computers exchangetime-controlled state messages via communication channels, and wherein each gateway computer sets up the connection to the process peripheral area associated with the gateway computer and stores the current state of the process peripheral area associated with the gateway computer, and wherein one central computer takes on the role of an active central computer and another central computer takes on the role of a passive central computer, and wherein the active central computer exercises control over the gateway computers, and wherein the active central computer, preferably periodically, sends a sign-of-life message to the passive central computer, and wherein the passive central computer acknowledges the arrival of a sign-of-life message from the active central computer in a periodic sign-of-life message and monitors it using a timeout, and wherein the passive central computer, given absence of these sign-of-life messages after the timeout, performs the role of the active central computer , and wherein the failed, previously active central computer autonomously attempts a restart and, following a successful restart, observes the message traffic within a cluster, which cluster comprises the central computers and the gateway computers, in order to learn the current state of the cluster, and wherein it performs the role of the passive central computer and communicates to the now active central computer, by means of, preferably periodic, sign-of-life messages, that it is performing the role of the passive central computer, and wherein, if the restart is not successful, the failed central computer indicates the permanent error by means of a display means.