Redundant Automation Switchover Delay for Transient Error Recovery

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing redundant automation systems face complete failure due to transient errors, where synchronization issues lead to premature switching of the reserve subsystem into troubleshooting mode, resulting in prolonged downtime as both subsystems fail to maintain process control.

Innovation Solution

Implementing a time delay before initiating troubleshooting in the reserve subsystem to allow it to assume process control and avoid complete system failure, with the master subsystem transferring internal data and updating the reserve subsystem to determine error causes, thereby preventing sudden system shutdowns.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Loss of time

If the reserve subsystem immediately initiates troubleshooting after loss of synchronization, then the error can be localized quickly, but the automation system may fail completely during troubleshooting

Engineering Contradiction:
Improvetime to localize errorVSAvoidsystem availability
Core Design Contradiction:
Loss of timeVSReliability

Solution Approach 1:

The reserve subsystem delays initiating troubleshooting by a predefined time period after detecting loss of synchronization. This preliminary waiting period allows the reserve subsystem to first attempt to assume master function and determine if the master subsystem has failed. Only after this preliminary action does the reserve initiate troubleshooting, preventing complete system failure while still enabling error localization.

Inventive Principle:
Principle #10Preliminary action

2Reliability

If the reserve subsystem waits with a time delay before troubleshooting, then complete system failure is avoided, but the time to localize error increases

Engineering Contradiction:
Improvesystem availabilityVSAvoidtime to localize error
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The predefined time delay serves as a preliminary action period during which the reserve subsystem monitors whether the master subsystem recovers or fails. This structured waiting period balances system reliability with error localization time by allowing the reserve to assume control if needed while still initiating troubleshooting within a bounded time frame.

Inventive Principle:
Principle #10Preliminary action

3Stability of the object's composition

If the reserve subsystem continuously synchronizes with the master, then data consistency is maintained, but transient errors can propagate to both subsystems

Engineering Contradiction:
Improvedata synchronizationVSAvoidsystem resilience to transient errors
Core Design Contradiction:
Stability of the object's compositionVSReliability

Solution Approach 1:

The system applies preliminary anti-action by having the reserve subsystem delay troubleshooting initiation and attempt to assume master function before full synchronization failure occurs. This prevents the propagation of transient errors to both subsystems by allowing the reserve to operate independently if the master fails, rather than forcing continuous synchronization that could spread errors.

Inventive Principle:
Principle #9Preliminary anti-action

Data Source

PatentUS11262745B2Method for operating a redundant automation system to increase availability of the automation system
Publication Date: 2022.03.01 SIEMENS AG
  • US11262745B2 patent drawing
  • US11262745B2 patent drawing
  • US11262745B2 patent drawing

AI summary

A method for operating a redundant automation system having a plurality of subsystems, wherein one subsystem of the plurality of subsystems operates as a master and assumes process control and the other subsystem operates as a reserve during redundant operation, where measures are provided by which the availability of the redundant automation system is increased, and where regardless of whether transient errors occur on the subsystem of the plurality of subsystems operating as the master or on the subsystem operating as the reserve, a total failure of the automation system is largely avoided.