Redundant Controller Failover Using NRP Reachability Checks
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional heartbeat-based failure detection in redundant controller systems cannot reliably distinguish between network failures and primary controller failures, leading to a dual-primary condition that can result in inconsistent system states and potential downtime or material damage.
Innovation Solution
Implement a decentralized redundant control system with Network Reference Point (NRP)-guided failure detection, where the backup controller verifies the reachability of an NRP to determine if the primary controller has failed, using a time-limited lease mechanism to ensure consistent failover decisions.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If conventional heartbeat-based failure detection is used, then the backup controller can detect primary controller failures, but network failures cannot be distinguished from primary controller failures, leading to dual-primary conditions
Solution Approach 1:
The patent introduces a Network Reference Point (NRP) as an intermediary node that mediates between the backup controller and the network infrastructure. The NRP responds to test messages from the backup controller to confirm network reachability, allowing the backup to distinguish between network failures (NRP unresponsive) and primary controller failures (NRP responsive but primary unresponsive). This intermediary mechanism resolves the ambiguity that causes dual-primary conditions.
2Speed
If the backup controller assumes primary role immediately upon heartbeat timeout, then failover speed is improved, but the risk of dual-primary condition increases due to network partitioning
Solution Approach 1:
The patent implements a preliminary action by requiring the backup controller to send a test message to the NRP before assuming the primary role. This preliminary check verifies network reachability and prevents premature failover in network partitioning scenarios. The backup controller only becomes primary if both the heartbeat timeout occurs and the NRP is reachable, ensuring system consistency while maintaining relatively fast failover.
3Reliability
If redundant network paths are implemented, then the probability of network problems is reduced, but cannot be eliminated entirely, leaving dual-primary risk non-zero
Solution Approach 1:
The NRP acts as a common reference point that both primary and backup controllers can reach through the network. Even with redundant network paths, the NRP provides a single point of verification for network health. The backup controller uses the NRP response to make an informed failover decision, eliminating the dual-primary risk that persists even with redundant networks.
Data Source
AI summary
A control device is used with at least one further control device in controlling an industrial system to which the control device and further control device are connected via a data network. The control device functions as primary controller when it feeds control signals to the industrial system, and functions function as a backup controller when it routinely performs a failure detection on the primary controller via the data network, and transforms into the primary controller in reaction to a positive failure detection. The backup controller transforms into the primary controller only when a network reference point, NRP, responds to a call from the backup controller, wherein the NRP is a node in the data network which connects the primary controller and backup controller to the industrial system. A malfunctioning NRP can be replaced at runtime.


