Redundant Controller Failover via Network Reference Point
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional heartbeat-based failure detection in redundant controller setups can lead to a dual-primary condition due to network problems, resulting in inconsistent system states, downtime, or material damage, as it cannot reliably distinguish between network failures and primary controller failures.
Innovation Solution
A decentralized arrangement of control devices that uses a Network Reference Point (NRP) to differentiate between network failures and primary controller failures, ensuring that only one control device assumes the primary role by verifying the NRP's responsiveness, thereby reducing the risk of dual-primary conditions and ensuring consistency.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If heartbeat-based failure detection is used to monitor primary controller status, then failure detection capability is improved, but the risk of dual-primary condition increases due to inability to distinguish network failures from controller failures
Solution Approach 1:
The patent introduces a Network Reference Point (NRP) as an intermediary entity that mediates between the backup controller and the network. The NRP provides reference information about network status that allows the backup controller to distinguish between actual primary controller failures and network failures, thereby preventing false failover decisions that lead to dual-primary conditions.
Solution Approach 2:
The patent implements a feedback mechanism where the backup controller continuously monitors network status through the NRP and adjusts its failover decisions based on this feedback. The NRP provides real-time information about network connectivity, allowing the backup controller to make informed decisions about whether to initiate failover, thus reducing the risk of dual-primary conditions while maintaining reliable failure detection.
2Device complexity
If conventional heartbeat monitoring is implemented, then system simplicity is maintained, but system consistency is compromised when network partitions occur
Solution Approach 1:
The NRP acts as an intermediary that provides network status information to the backup controller without requiring complex changes to the overall system architecture. This allows the system to maintain relative simplicity while adding the capability to detect network partitions and prevent inconsistent dual-primary states through informed failover decisions.
3Productivity
If failover decision is made quickly upon heartbeat loss, then availability is improved, but consistency is lost due to potential network failures being misinterpreted as controller failures
Solution Approach 1:
The patent implements preliminary action by having the backup controller continuously monitor network status through the NRP before a failover event occurs. This preliminary monitoring establishes a baseline of network health that can be referenced during failover decisions, allowing the system to quickly respond to actual failures while avoiding premature failover due to network issues.
Solution Approach 2:
The feedback mechanism from the NRP provides real-time network status information that influences failover timing. The backup controller uses this feedback to determine whether network conditions justify immediate failover, thereby achieving both quick response to genuine failures and accurate differentiation from network-induced false alarms.
Data Source
Figure 1A~1B
Figure 1C~1D
Figure 1E~1F
AI summary
A control device (220a) for use with at least one further control device (220b) in controlling an industrial system (210), to which the control device and further control device are connected via a data network (240). On the one hand, the control device is operable to function as primary controller, wherein it feeds control signals (110) to the industrial system. On the other hand, the control device is operable to function as a backup controller, wherein it routinely performs a failure detection on the primary controller via the data network, and transforms into primary controller in reaction to a positive failure detection. The transformation from backup controller into primary controller is conditional upon verifying that a node (230) in the data network appointed as network reference point, NRP, responds to a call from the backup controller. A malfunctioning NRP can be replaced with a different NRP candidate at runtime.