Secure failover between redundant controllers
By introducing a network reference point (NRP) to guide fault detection, the problem of not being able to reliably distinguish between network faults and master controller faults in redundant controller systems is solved, achieving higher system consistency and reliability and reducing the risk of dual-master conditions.
Patent Information
- Application Number
- CN202480022158.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-09-11
- Filing Date
- 2024-02-13
- Publication Date
- 2025-10-31
AI Technical Summary
In existing redundant controller systems, heartbeat-based fault detection methods cannot reliably distinguish between network faults and master controller faults, resulting in a high risk of dual-master conditions and affecting the consistency and availability of industrial systems.
A Network Reference Point (NRP) is introduced to guide fault detection. By verifying the reachability and lease mechanism of the NRP, network faults are distinguished from master controller faults, ensuring the reliability and consistency of fault switching.
It effectively reduces the risk of dual-master conditions, improves the reliability and consistency of redundant controller systems in the event of network failures, and ensures the stable operation of industrial systems.
Smart Images

Figure CN120883196A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of redundant controller setups having an active controller (primary controller) and one or more additional controllers (backup controllers) that monitor the active controller and are prepared to replace the primary controller upon detecting a failure. In particular, a failover scheme is proposed that reduces exposure to known highly disruptive dual-primary conditions. Background Technology
[0002] Control systems are increasingly network-oriented. Control applications are envisioned to operate across a wide range of targets, from today's embedded controllers to industrial PCs and edge devices, and even in the cloud. Therefore, control system software should be as hardware-independent as possible to broaden the range of deployment options. The preferred communication method is switched Ethernet, and communication-related functions (such as redundancy) should preferably rely solely on standard Ethernet networks.
[0003] A backup controller redundancy solution consists of an active primary controller and at least one passive backup controller. The backup redundancy solution requires the backup controller to detect failures in the primary controller, thus necessitating fault detection. Fault detection should be hardware-independent and rely solely on the network solution to maximize deployment alternatives. One possible fault detection solution is heartbeat signal detection. Heartbeat fault detection involves the supervised controller periodically sending messages to the supervising node. In the redundancy controller context, the supervised controller is the primary controller. The supervising node, acting as a backup in the redundancy controller context, expects to observe heartbeat messages within a known timeframe. Detecting a heartbeat absence exceeding a threshold / timeout causes the supervising backup controller to assume the heartbeat absence is due to a primary controller failure. EP3933596A1 discloses a solution following these ideas.
[0004] However, the absence of an observable heartbeat on the backup controller could also be caused by network issues. Redundant networks typically reduce the probability of network problems, but it will never be zero. Therefore, using traditional heartbeat-based fault detection, the probability of the system facing a dual-master situation due to network problems is zero.
[0005] Figure 16AThe deployment is illustrated, with one control device 220a (shown as a distributed control node DCN) acting as the master controller and another redundant control device 220b acting as a backup controller. The master controller controls the industrial system comprised of components 210a, 210b, and 210c (Fieldbus Communication Interface FCI) by feeding control signals to it via a network infrastructure including multiple switches 230. The master controller also sends a heartbeat signal, which the backup controller monitors. When two indicated network failures F1 and F2 occur, the heartbeat signal from the master controller can no longer reach the backup controller, which may mistakenly assume the role of the master controller. Components 210a, 210b, and 210c will then receive inconsistent control signals from the two control devices 220a and 220b. Due to this dual-master condition, the industrial system may eventually be in an inconsistent state, leading to downtime or material damage.
[0006] The goal is to avoid dual-master conditions as much as possible. Both US7986618B2 and US8082340B2 propose methods for distinguishing between link and node failures, but these solutions make assumptions about the network that cannot be generalized to all industrial-related settings.
[0007] US20060056285A1 discloses a redundant host pair runtime configuration for a process control network environment. The active partner of the failover host pair operates on a first machine communicatively connected to the main network, and hosts a set of executing application components. The standby partner of the failover host pair operates on a second machine communicatively connected to the main network. The standby partner checks the heartbeat of the activity engine via the main network and via the active partner's dedicated redundant message channel, RMC. Failover is triggered when the standby partner does not receive a heartbeat from the active partner via RMC, provided that the standby partner and at least one other node in the main network are unreachable from the activity engine node via the main network. Otherwise, if the standby partner is still reachable, it is assumed that RMC is down, and the active partner enters a ready state. Summary of the Invention
[0008] One objective of this disclosure is to provide a distributed arrangement of redundant control devices to reduce the risk of two control devices simultaneously assuming the role of master controller (dual master condition). Another objective is to provide a control device for use with at least one other control device when controlling an industrial system via a data network to reduce exposure to dual master conditions. A further objective is to provide a control device capable of making consistent failover decisions (i.e., becoming the master controller) within a finite timeframe.
[0009] At least some of these objectives are achieved by the invention as defined in the independent claims. The dependent claims relate to advantageous embodiments.
[0010] In a first aspect of this disclosure, a control device is provided for use with at least one additional control device when controlling an industrial system (or a general-purpose technical system), the control device and the one or more additional control devices being connected to the industrial system via a data network. The control device is operable to: act as a master controller, wherein the control device feeds control signals to the industrial system; and act as a backup controller, wherein the control device routinely performs fault detection on the master controller via the data network, and switches to master controller in response to a positive fault detection. According to the first aspect, the switching from backup controller to master controller depends on whether a verification network reference point (NRP) responds to a call from the backup controller, wherein the NRP is a node in the data network connecting the master controller and the backup controller to the industrial system. (Alternatively, the switching from backup controller to master controller depends on whether the network node acting as the verification NRP responds to a call from the backup controller.)
[0011] In a second aspect of this disclosure, an arrangement of control devices having the behavior according to the first aspect is provided. Due to this behavior, the role of the main controller will be performed by one of the control devices at a time, or, in special cases, not by any control device.
[0012] In a third aspect, a computer program is provided, comprising instructions for causing a computer or, in particular, a control device to behave in accordance with the first aspect. The computer program can be stored or distributed on a data carrier.
[0013] Finally, in a fourth aspect, a method is provided for operating an arrangement of redundant control devices when controlling an industrial system, wherein the control devices and other control devices are connected to the industrial system via a data network. The method includes: feeding control signals to the industrial system from a first control device currently acting as the master controller; routinely performing fault detection on the master controller from a second control device currently acting as a backup controller; in response to a positive fault detection, making a call to a network node acting as a network reference point (NRP) from the second control device; and deciding to switch to master controller only when a response from the NRP is received from the second control device.
[0014] The aspects of this disclosure described above are based on the inventors' understanding that a key reason for the dual-master condition is that, from the perspective of the backup controller, network failures cannot be reliably distinguished from failures in the primary controller itself. Any network node (existing or added) designated for this purpose, such as the NRP, aims to provide an indication of the current health status of the network, which allows for the reliable resolution of ambiguity regarding whether an observed failure is related to the network or the primary controller. In other words, based on the presence or absence of a response from the NRP, the backup controller distinguishes between failures in the primary controller and failures in the data network. In particular, the reachability of the NRP can provide information about whether the network is affected by failures that divide the network into components. This information helps determine whether failover (or equivalent switching or transformation) is sufficient.
[0015] In some embodiments, fault detection is based on timestamped heartbeat signals. In some cases, this eliminates the need to call the NRP, allowing for transition decisions to be made in a shorter time. This is particularly useful in data networks operating implementations that do not support protocols with bounded response times such as ping.
[0016] In other embodiments, fault detection is mediated by the NRP and is based on whether the primary controller has a valid lease with the NRP. More specifically, when acting as the primary controller, the control device initiates a time-limited lease with the NRP and continues to renew the lease while the control device is acting as the primary controller. The control device is also configured such that when acting as a backup controller, it performs fault detection on the primary controller by querying the NRP whether the NRP has a lease with the primary controller. The lease can be time-limited, such that the lease lasts for a pre-configured duration. T L After the expiration date, this means that the main controller preferably ensures that the NRP expires. T L The next lease renewal request is processed within a time unit to achieve uninterrupted leases. One anticipated advantage of these embodiments is that NRP-mediated fault detection adds additional robustness to the aforementioned undesirable dual-master condition. Indeed, even if the master controller is functioning normally, some alternative fault detection methods can return positive results (i.e., they detect obvious faults in the master controller) due to faults in the data network. When NRP is used for the dual purpose of detecting faults in both the master controller and the data network, this situation occurs in very rare cases, or is completely eliminated, according to embodiments with NRP-mediated fault detection.
[0017] In embodiments with NRP-mediated fault detection, preferably, the primary controller and backup controller use a common message format for lease initiation / renewal and lease querying. The message format indicates the sender. NRP interprets the lease differently depending on whether it is valid, and if valid, whether the sender is the primary controller.
[0018] In some embodiments, the control device is configured to replace a faulty NRP with a different NRP candidate during runtime. This enhances the reliability of the control device arrangement. In particular, the control device arrangement will maintain functionality for a longer average time and / or may be more resilient to disturbances.
[0019] As used herein, "data carrier" can be a transient data carrier, such as modulated electromagnetic waves or light waves, or a non-transient data carrier. Non-transient data carriers include both volatile and non-volatile memories, such as permanent and non-permanent storage media of magnetic, optical, or solid-state types. Still within the scope of "data carrier," such memories can be fixed in place or portable.
[0020] Generally, unless otherwise expressly defined herein, all terms used in the claims shall be interpreted in accordance with their ordinary meaning in the technical field. Unless otherwise expressly stated, all references to “an element, device, assembly, component, step, etc.” shall be interpreted as referring to at least one instance of an element, device, assembly, component, step, etc. Unless expressly stated otherwise, the steps of any method disclosed herein need not be performed in the exact order disclosed. Attached Figure Description
[0021] Various aspects and embodiments will now be described by way of example with reference to the accompanying drawings, in which: Figures 1A to 1G It is a sequence diagram illustrating the data exchange between control equipment, industrial systems, and network nodes designated as NRPs; Figure 2 and Figure 3 The arrangement of two redundant control devices according to the prior art is shown; Figures 4 to 15 The arrangement of the two control devices according to this disclosure is shown to respond to various situations; Figure 16A and Figure 16B The arrangement of two control devices connected to an industrial system via a network infrastructure is shown; and Figure 17 This is a state diagram illustrating the behavior of the control device according to this disclosure. Detailed Implementation
[0022] Various aspects of this disclosure will now be described more fully with reference to the accompanying drawings, which illustrate certain embodiments of the invention. However, these aspects may be embodied in many different forms and should not be construed as limiting; rather, these embodiments are provided by way of example so that this disclosure will be thorough and complete, and will fully convey the scope of all aspects of the invention to those skilled in the art. Throughout the description, the same numerals denote the same elements. introduction
[0023] The Consistency, Availability, and Partition Tolerance (CAP) theorem states that, in the case of partitioning, a distributed system can be either available or consistent, but never both. See Gilbert and Lynch's "Brewer's Conjecture and the Feasibility of Consistent, Available, Partition-Tolerant Web Services" (…). SIGACT News vol. 33, no. 2 (June 2002), pp. 51–59, doi:10.1145 / 564585.564601) and Brewer's "CAP twelve years later: How the 'rules' have changed" Computer (, vol. 45, no. 2 (February 2012), pp. 23–29, doi: 10.1109 / MC.2012.37). In the context of distributed control systems and controller redundancy, this means that in the case of partitioning, the dual-master case provides availability because both master controllers provide output values. (As used herein, "partitioning" refers to a network being undesirably split into separate components due to a failure in the network infrastructure.) However, because partitioning causes the master controllers to become unsynchronized, they are likely to be scattered into different states, thus providing inconsistent outputs. In other words, in the context of redundant controllers, conventional heartbeat-based fault detection provides availability in the case of partitioning, but not consistency.
[0024] Conversely, assuming that the DCNs forming the redundant pair will become passive in another uncertain state, they can be expected to maintain consistency because if the controller does not provide any value, the input / output interfaces will use pre-configured default values (outputs are set to predetermined values (OSP)).
[0025] The Network Reference Point Failure Detection (NRP FD) solution disclosed in this paper minimizes the unavoidable availability mitigation caused by the CAP theorem while maintaining consistency by prioritizing consistency over availability. Network Reference Point (NRP) Selection
[0026] A network reference point (NRP) can be any device in a data network that has no common causes of failure shared with the monitored devices, apart from the data network itself. For example, the managed switch closest to the main controller that meets the requirement of no common causes of failure can be designated as an NRP. This managed switch is independent of the main controller. More generally, a device in the data network is suitable to be designated as an NRP if it is independent of the main controller and has no common causes of failure with it. Optionally, the NRP should be independent of the control device used as the main controller, i.e., independent of processing resources that execute computer-readable instructions that implement the functions of the main controller.
[0027] Because networks are typically redundant, each network can have at least one potential NRP. Figure 16B ( Figure 16A In the annotated version, a set of potential NRPs is represented as NRP candidates. (In other examples, switches 230b and 230e can be additional NRP candidates.) NRP candidates can overlap, meaning the primary controller and backup controller can propose the same NRP candidate. This occurs when there is only one switch between the primary and backup controllers. It is the primary controller that proposes the NRP and is proposed by the NRP candidate. Regardless of which controller proposes the NRP candidate, the specified NRP will be common to all controllers in data network 240 configured to control industrial system 210. (If another industrial system is controlled by a similar redundant control device arrangement, its NRP is specified independently. This is also valid if the redundant control devices are connected to the same data network 240.)
[0028] In some embodiments, an NRP is a (individually addressable) node in a data network, jointly designated as an NRP by the control devices. In some embodiments, a particular control device is configured to propose a node in the data network to one or more other control devices for designation as an NRP, the node having no known common cause of failure with the particular control device. In some embodiments, additionally or alternatively, a particular control device is configured to propose a node in the data network located on a path from itself to at least one other control device for designation as an NRP to one or more other control devices. A sequence in which one control device proposes an NRP candidate, and all other control devices accept the NRP candidate as the NRP, is understood as the joint designation of the NRP. Heartbeat-based NRP boot failure detection – Overview
[0029] The NRP-guided fault detection algorithm (NRP FD) proposed in this paper can be described as a heartbeat-based fault detection method that utilizes NRP in addition to the heartbeat. In order to become and maintain the master controller, the control device 220 must be able to reach the NRP. Therefore, it is beneficial to use an NRP that is already part of the existing infrastructure, such as a switch. Playing the role of the NRP typically does not require any new functionality from the node; responding to ping messages is already part of the network protocols supported by most commercial network devices.
[0030] The core of NRP FD is described below, starting with the primary controller's perspective and then moving to the backup controller's view. The same thing will be described in more detail in the next section.
[0031] The primary controller selects an NRP from its NRP candidates. The primary controller sends a heartbeat to the backup controller. The heartbeat message may contain the IP address of the NRP, as shown in Table 1. The iteration / loop count can be replaced with a timestamp, and it is equivalent to a timestamp because the relationship between the iteration / loop count and clock time can be calculated. The master controller also monitors the NRP; if it cannot reach the NRP, it proposes a change. If the change request fails, the master controller will relinquish its master controller role by transforming into a backup controller.
[0032] The backup controller detects that it receives a heartbeat signal from the primary controller. If the heartbeat monitoring determines a timeout, the backup controller checks if the NRP is reachable; if the NRP is reachable, the backup controller becomes the primary controller. NRP reachability testing can be performed using ICMP ping (such as ICMP echo). ICMP ping has no hard real-time guarantee. This lack of a hard real-time guarantee can be mitigated by interpreting simultaneous silence as a primary failure rather than two simultaneous network failures (which is impossible). In this case, simultaneous silence would cause the backup controller to directly resume the primary role without testing NRP reachability. The drawback of interpreting simultaneous failures as primary failures is that it may incorrectly interpret actual simultaneous network failures as primary failures. Another approach to addressing the lack of a ping with real-time guarantees in the protocol is to configure the network node acting as the NRP accordingly; this may increase the overall implementation cost and / or reduce the number of nodes that meet the criteria for being an NRP, but this can be the best option if high responsiveness (fast failover) is required.
[0033] In this disclosure, the term NRP Ping is used for NRP reachability checks. In cases where only one network remains (alternative term: only one network path), there is no other way for the backup controller to indicate this except by using NRP Ping to become the primary controller. However, this is optional and may depend on failover time requirements.
[0034] NRP FD can be advantageously implemented in data networks where some nodes, particularly switches, support hard real-time NRP Ping (ping with bounded response time). Currently, this is unavailable in all relevant types of industrial control networks. Hard real-time low-latency NRP Ping could reduce NRP Ping test times using ICMP Ping from 5-20 ms to less than one millisecond. Link supervision mechanisms similar to Bidirectional Forwarding Detection (BFD) can also leverage this real-time NRP Ping to provide a real-time link failure mechanism. Real-time link failure detection mechanisms can be used in operational technology (OT) networks deployed for purposes other than redundancy; one envisioned use case is network supervision. Heartbeat-based NRP-guided fault detection – State machine description
[0035] NRP FD such as Figure 17 The state machine diagram is shown in [the diagram]. Figure 17 In this context, the abbreviation "tmo" indicates timeout, and the variable "NW" indicates the number of different networks (network paths) through which the backup controller monitors the heartbeat signal.
[0036] The primary states are Startup-Initialization / Wait 1710, Backup Controller-Supervisor 1720, and Master Controller-Supervised 1730, with the initial substate indicated by 1712. Control devices operating in any of these states exhibit deterministic behavior. A control device in the Startup-Initialization / Wait state 1710 cannot be characterized as a master controller or a backup controller, but is typically distinguished as one based on received signals, timeouts, or a combination thereof.
[0037] Startup - Initialization / Waiting State 1710 The process begins by waiting for an NRP candidate (state 1714). After waiting for an NRP candidate, the system waits for confirmation of becoming the master controller or a heartbeat from another master controller (state 1716). If confirmation of becoming the master controller is received, the control device will decide to act as the master controller. The master confirmation is issued by the operator, system owner, or other person monitoring the system. Under normal circumstances, each industrial system to be controlled should have at most one master controller. Upon assuming the master role, an NRP is selected from the NRP candidates.
[0038] If a heartbeat is observed from the existing master controller, the control device decides to act as the backup controller. The backup controller announces its existence to the master controller and will not become the takeover backup controller until it sees its own IP address from the master controller. The master controller confirms its awareness of the backup controller's existence by sending the backup controller's IP address in the heartbeat message; see the example message format in Table 1.
[0039] The pseudocode in Table 2 summarizes this behavior.
[0040] The pseudocode in Table 3 summarizes advanced NRP candidate monitoring / detection. This behavior is common to all states 1710, 1720, and 1730. More precisely, it involves monitoring NRP candidates and maintaining a list of reachable NRP candidates (…). reachableCandidates The set of NRPs. NRP selection can be performed by selecting an NRP from this set. reachableCandidates The set can be updated periodically, for example, several times per minute. In implementations utilizing multiple networks (network paths), each network has NRP candidates. Function GetCandForNw () retrieves NRP candidates for a specific network, and PingNRP () Call NRP; it can be implemented as ICMP Ping or real-time ping.
[0041] In Backup Controller - Monitoring Status 1720 The control equipment monitors the main controller. For an industrial system to be controlled, there can be multiple backup controllers. The pseudocode in Table 4 summarizes the behavior of the backup controllers. In the implementation based on this pseudocode, the backup controller continuously reports its existence to the primary controller because the primary controller needs to know if the backup controller exists in case the primary controller cannot reach the NRP. Function ChkHbStsOnAllNw This function checks the heartbeat status on all networks (network paths) expected to have heartbeats. If the function finds that a timeout has occurred on all networks (network paths) simultaneously, it returns 0. tmoAllSimul If the heartbeat times out on all networks (network paths), but not simultaneously, return 0. tmoAllNotSimul Finally, if the backup controller detects a heartbeat on some (but not all) networks (network paths), the function returns. tmoSomeNotAll This indicates that a partition has occurred.
[0042] In state 1726, check the heartbeat status of each network (each network path). Assume the heartbeat signal times out on all networks (network paths) connecting the master and backup controllers, and times out simultaneously on two or more networks. In this case, you can directly enter the master state without performing an NRP test, as the silence is likely due to a failure of the master controller, not a simultaneous network failure. However, if the NRP Ping is considered to meet the required real-time properties, this execution path can be skipped, relying solely on the NRP Ping. This is the preferred solution and the only solution to completely eliminate the dual-master risk. Using simultaneous timeouts reduces the risk of the aforementioned dual-master condition, but not necessarily to zero.
[0043] State 1724 corresponds to the discovery that the heartbeat is lost on some, but not all, networks (network paths). Then, NRP reachability is tested (state 1722), because silence can indicate a network outage between the backup controller and the NRP itself. If the backup controller finds the NRP unreachable, it can choose to request a replacement NRP from the controller.
[0044] Control devices in the master controller-supervised state 1730 can have behavior consistent with the pseudocode in Table 5.
[0045] Initially, in state 1732, a heartbeat signal is sent in each iteration / loop. The heartbeat signal may contain messages such as declaring the IP address of the NRP; see Table 1.
[0046] Next, NRP reachability is tested, and if NRP reachability is achieved, execution continues.
[0047] However, if the NRP is unreachable, and if a backup is available, the master controller proposes to switch the NRP (state 1734). If no backup exists, the master controller will simply switch the NRP, assuming it has a available NRP candidate reachable on another network. If a backup exists, the master controller sends a request to the backup to change the NRP and waits for a response within a finite time. If the change is positively acknowledged, the master controller will change the NRP. This change affects all controllers in the network. Otherwise, if a negative response or no response is received, the master controller can no longer be the master controller and re-enters the startup-initialization / wait state 1710. NRP-mediated fault detection
[0048] The NRP-mediated fault detection algorithm proposed in this paper makes dual use of NRP: NRP reachability is used as a proxy for evaluating the operational status of the data network, and NRP is used to deliver information from the primary controller to (multiple) backup controllers for fault detection.
[0049] In order to become and remain the primary controller, control device 220 must be able to reach the NRP. Therefore, it is advantageous to use an NRP that is part of existing infrastructure, such as a switch. The NRP role typically does not require the node to provide any new functionality. For an incoming lease query from the backup controller, the NRP should preferably respond to whether the primary controller has an active lease, or it can respond to the time when the primary controller renewed the most recent lease, so that the backup controller can infer whether the lease is still valid.
[0050] The master controller selects an NRP from its NRP candidates. The master controller initiates a time-limited lease with the NRP and continues to renew the lease during its role as master controller. It should be noted that the role of the master controller includes feeding control signals 110 to the industrial system 210.
[0051] In the context of data networks, a “lease” in this public sense is understood as a contract that grants its holder specific rights for a limited period. A lease is not a legal contract, but rather a metaphor for the rules governing the interactions of entities connected to a data network. A lease can alternatively be described as a token designating rights, or as a lock on designated rights with a timeout. Designated rights refer to the authority (responsibility) of acting as a master controller, including feeding control signals 110 to industrial system 210.
[0052] For example, a lease can be represented as an NRP state, such as the value of a (binary) state variable, specifically the pre-configured duration corresponding to the duration of the time-limited lease. T L The lease is valid as long as the timer is running, and expires when the timer expires. T L The lease expires after a certain number of time units. Alternatively, an internal variable representing the identity of the control device 220 currently acting as the master controller can be maintained in the NRP, and this internal variable is configured to expose the identity in response to a lease query.
[0053] Alternatively, lease-related bookkeeping can be implemented as an internal variable of the NRP, which stores the time of the primary controller's most recent ping. In some implementations, the NRP can maintain an internal clock or access a network time source upon receiving a ping from the primary controller and store that time. In other implementations, the current time can be pre-specified in the primary controller's ping, where the NRP should accept that time as valid and store it. According to this alternative, the NRP responds to subsequent lease queries (calls) from the backup controller by declaring the stored time. This allows the backup controller to calculate the elapsed time since the stored time and compare it with... T L Compare this. Another alternative is that NRP calculates the time elapsed since the storage time and compares it with... T L The comparison is performed. If the NRP also stores the identifier of the primary controller holding the lease, it can determine whether a subsequent call originates from the primary controller or the backup controller. The NRP can then process a call with a uniform format (e.g., a ping message) as a lease renewal request or a lease query based on the initiator.
[0054] After the primary controller initiates a lease for the first time, it can renew the lease by sending a New Initiation Request (NRP) (or a dedicated renewal request). To avoid time intervals between consecutive leases—which, in the worst case, could trigger a backup controller to incorrectly conclude that the primary controller has failed—the primary controller should preferably ensure that the NRP is maintained within a specified timeframe. T L The main controller can process its next lease renewal request within a time unit. Therefore, the main controller can operate for a duration of... T L -Δ units of internal timers, where Δ corresponds at least to the round-trip time between the master controller and the NRP in the data network, possibly with safety margins and additional processing time. The master controller monitors the NRP by repeatedly requesting lease renewals; if it cannot reach the NRP, it proposes a change. If the change request fails, the master controller relinquishes its role by transforming into a backup controller. If one or more backup controllers accept (implicitly or explicitly) the proposed change for the new NRP, the master controller initiates a lease with the new NRP. Optionally, the master controller can store a reference to the new NRP in the old NRP (e.g., by writing the identity of the new NRP into the memory of the old NRP with write access), so that any other backup controller querying the old NRP is aware of the NRP change and redirects accordingly.
[0055] The backup controller repeatedly monitors whether the primary controller has an active lease with the NRP. The backup controller performs this monitoring by sending a lease query to the NRP. If the NRP responds negatively to the lease query (or the storage time exceeds the past...), the backup controller will detect the active lease. T L If the backup controller becomes the primary controller (i.e., there are multiple units), then the backup controller will become the primary controller. In fact, the backup controller is now aware of NRP reachability and will not attribute negative responses to network failures. It should be remembered that in NRP-guided heartbeat-based fault detection, the backup controller must perform a separate NRP reachability check before becoming the primary controller.
[0056] Alternatively, NRP-mediated fault detection can be implemented in data networks where some nodes, particularly switches, support hard real-time NRP Ping (Ping with bounded response time) or hard real-time responses to incoming queries. Refer to the observations and discussions on hard real-time capabilities in the earlier sections of this disclosure.
[0057] After studying this disclosure, those skilled in the art will be able to implement the NRP-mediated fault detection method using conventional components and instructions. In particular, if a state machine-based implementation is required, those skilled in the art will be able to refer to Tables 2-5 and... Figure 17 Make the necessary modifications to form the state machine description of the NRP-mediated fault detection method. Implementation
[0058] NRP equipment As mentioned earlier, the preferred NRP (Network Provider Relationship) is a device already included in the data network. A switch, as part of the network infrastructure, is such an example. Choosing a different device as the NRP (i.e., a device that is not part of the infrastructure connecting the primary and backup controllers) often has a negative impact on reliability / availability, because this additional device also needs to function properly for the redundancy to be available.
[0059] If the equipment also includes a duration indicating the time-limited lease... T L If a timer is used, then the device is suitable for use as an NRP in embodiments with NRP-mediated fault detection. Such a timer should be started each time a lease is initiated or renewed so that when the timer expires, i.e., on [date missing], [further details missing]. T L After a certain number of units, the lease is considered to have expired. Furthermore, devices without timers can also be used as NRPs; it is sufficient for the NRP to store the time when the latest lease was initiated (which can be declared by the primary controller) and declare this time in response to a lease query from the backup controller. The backup controller can then execute the calculation of the elapsed time since the storage time and the pre-configured duration. TL A comparison.
[0060] NRP reachability test / NRP Ping Today's managed commercial off-the-shelf (COTS) switches typically support Ping messages compliant with the Internet Control Message Protocol (ICMP) specification. ICMP Ping can be used as NRP Ping. However, as mentioned above, ICMP Ping does not offer a hard real-time guarantee. Empirical testing shows that response times can range from less than milliseconds to several milliseconds, such as up to 5ms. To achieve lower latency and a hard real-time guarantee, adding hard real-time 'ping functionality' support to existing switches can be considered. Real-time ping guarantees low latency and bounded response times.
[0061] For embodiments with NRP-mediated fault detection, a ping message can trigger at least two distinct responses from the NRP, indicating a state of 'the primary controller has an active lease' (P) and 'the primary controller has no lease' (N). The distinct responses can be represented as a single-format message, which may employ two different values or two different message formats. Alternatively, as described above, the NRP can respond by specifying the time the latest lease was initiated / renewed, allowing the backup controller to detect faults by a pre-configured duration. T L A comparison is used to determine whether the lease is still valid. Advantageously, a unified message format is used for lease initiation / renewal (by the primary controller) and lease inquiries (by the backup controller); in practice, a message format indicating the message sender is used. More precisely, - If the NRP does not have an active lease when it receives the message, it will grant the lease to the sender; - If the NRP has a lease with another control device, it sends a response to the sender to indicate that the master controller has a valid lease; and - If NRP has an active lease with the message sender, it confirms that the lease has been renewed for a new time period. T L . Using a unified message format for lease initiation / renewal and lease inquiries can avoid certain conflict situations. For example, when it has been agreed that the first control device will take over the role of master controller from the second control device, it can prevent the second control device from unintentionally acquiring a lease with the NRP before the first control device sends its lease request.
[0062] NRP candidate discovery and selectionNRP FD can use predefined configurations to discover NRP candidates. Another dynamic approach is to use the Link Layer Discovery Protocol (LLDP). Most managed switches support LLDP to announce their presence to neighboring devices. With this information, the Distributed Control Nodes (DCNs) forming redundant pairs can map the topology between them and dynamically select their NRP candidates. Example
[0063] Figures 1A to 1D This is a sequence diagram illustrating data exchange between control device 220, industrial system 210, and network nodes NRP designated as NRPs in several example scenarios. It should be understood that these entities are interconnected via at least one data network 240 (preferably a wired data network). Data network 240 may have a mesh topology. Furthermore, it is assumed that control devices 220a, 220b operate in a decentralized manner, i.e., without any assistance from a centrally coordinating entity. Figures 1A to 1D An example of heartbeat-based fault detection involving the backup controller.
[0064] exist Figure 1A In this configuration, the first control device 220a serves as the main controller, and the second control device 220b serves as the backup controller. As described above, the initial assignment of roles can initially be established via confirmation from the main controller issued by the operator. For simplicity, this example and other examples described herein refer to a binary (1oo2) setup, where one is the main controller and the other is the backup controller; however, it is understood that the proposed solution can be readily adapted to deployments with multiple backup controllers. Thus, the main controller 220a feeds control signals 110 to the industrial system and issues a heartbeat message 112, which can be monitored by the secondary controller 220b. Receiving the heartbeat message 112 within a predetermined time interval (e.g., within a pre-configured time period T1 after the previous heartbeat message) allows the backup controller 220b to make a negative fault determination 114 (N). However, if no heartbeat message is received after T2 units, where T2 > T1, the backup controller 220b makes a positive fault determination 114 (P).
[0065] Then, the second control device, acting as backup controller 220b, makes a call 116 to the NRP, such as a ping. If it receives a response 118, it can safely conclude that the expected absence of heartbeat message 112 is not due to an error in data network 240, but is more likely due to a failure of the first control device 220a. Therefore, the second control device switches from backup controller to master controller (step 120). After the switch, the second control device 220b will feed control signals 110 to the industrial system and send heartbeat message 112 every T1 time unit. If the first control device 220a starts operating again, it can monitor heartbeat message 112 and use it as the basis for recurring failure detection of the second control device 220b.
[0066] The effect of these final steps is that the condition for the backup controller to become the primary controller is verifying whether the Network Reference Point (NRP) responds to a call from the backup controller, where the NRP is the node in the data network that connects the primary and backup controllers to the industrial system. (In other words, the condition for the backup controller to become the primary controller is verifying whether the network node acting as the NRP responds to call 116 from the backup controller.) In other words, the backup controller is configured not to become the primary controller if no response to 118 is detected, unless there are special circumstances; conversely, it remains the backup controller even without a response to 118. Typically, becoming the primary controller does not depend on the content of the response to 118; this means that any network node capable of responding to any type of ping or ping-like message is eligible to become the NRP, a low requirement that greatly simplifies implementation.
[0067] An example of an exception could be that the backup controller detects a simultaneous interruption of timestamped heartbeat message streams 112 on at least two different paths of data network 240. This detection is possible, particularly when the master controller 220a sends heartbeat signals as timestamped broadcast / multicast messages (as shown in Table 1); detection is also possible if the master controller 220a sends two timestamped unicast message streams to the backup controller via different routes. If the backup controller detects a simultaneous interruption of timestamped heartbeat message streams 112 on at least two different paths of data network 240, it can take over as the master controller regardless of whether a response is received from the NRP; in effect, the backup controller waits for an NRP response until the timeout period ends.
[0068] In an embodiment where multiple control devices acting as backup controllers supervise a control device acting as a primary controller, the backup controllers may each issue a call 116 to the NRP. An equivalent basic configuration of the active control devices in data network 240 may result in the backup controller that first receives a response 118 to its call 116 becoming the primary controller, but each of the remaining backup controllers will remain a backup controller. In an implementation where call 116 is a message type specifically defined for this purpose (e.g., a dedicated type of ping), the NRP may be configured with a blocking period such that it will not respond to a second, third, or subsequent call 116 after responding to the first call 116 until the blocking period expires. This can limit the risk of dual-primary conditions. Alternatively, the NRP may be configured to respond to each incoming call 116 but indicate in its response 118 whether it has recently responded to another call 116 (e.g., a period of time equivalent to the blocking period); this would allow a backup controller receiving a response 118 with such an indication to voluntarily avoid becoming the primary controller.
[0069] exist Figure 1A In the illustrated embodiment, backup controller 220b routinely performs fault detection by listening to a periodic heartbeat signal. Note that in implementations where the heartbeat signal is used for other purposes, the listening period of backup controller 220b may be shorter than the period of the heartbeat signal. In other embodiments, backup controller 220b may perform fault detection by calling master controller 220a on a periodic, scheduled, or event-triggered basis (e.g., status polling).
[0070] Figure 1B The maintenance of the NRP by the main controller 220a and the procedure for specifying the replacement NRP are illustrated. For simplicity, Figure 1B This only implicitly includes message exchange between the main controller 220a and the industrial system 210; for a more complete description, please refer to [reference needed]. Figure 1D Initially, network node NRP1 is used as the NRP. Therefore, control device 220a, acting as the master controller, routinely verifies NRP1's response 118 to call 116. If response 118 is received, it makes a positive verification decision 122(P); otherwise, it makes a negative verification decision 122(N). In the illustrated embodiment, new calls 116 are sent at a period of T3 time units; alternatively, calls 116 may follow a pre-configured schedule, or calls 116 may be triggered by observable events by control device 220a, acting as the master controller. Figure 1BIn the process, since no response 118 is received after T4 time units (timeout period), a negative verification decision 122 (N) is made. Then, the primary controller 220a sends a proposal 124 to the backup controller 220b to replace the NRP (here: NRP2). After verifying that it has received a response to the proposed replacement NRP, the backup controller 220b may send an acceptance 126 to the primary controller 220a. This corresponds to designating NRP2 as the new NRP 128. Therefore, the primary controller 220a will redirect its routine calls 116 to NRP2. Furthermore, if the backup controller 220b later performs a positive fault detection on the primary controller 220a, the backup controller will check whether NRP2 is responsive before the backup controller 220b becomes the primary controller. In some embodiments, the backup controller routinely checks whether the NRP is responsive to a call from the backup controller, i.e., not only when it performs a positive fault detection on the primary controller.
[0071] Assuming that if the backup controller 220b does not accept the proposal to use NRP2 as a replacement NRP (i.e., by not sending an acceptance within a pre-agreed delay), the master controller 220a may propose 124 different replacement NRPs.
[0072] Figure 1C The scene in the original and Figure 1B The scenario is the same. However, within a time period of T5 time units exceeding the pre-configured timeout duration, the backup controller 220b does not send any acceptance of the replacement NRP proposal 124 to the primary controller 220a. It should be understood that the primary controller 220a has no other NRP candidates available to propose. In this case, the configuration of the control devices will result in the first control device 220a changing from the primary controller 130 to the backup controller and the second control device 220b changing from the backup controller 132 to the primary controller. Therefore, the second control device 220b will send a call 116 to NPR1 and routinely verify 122 whether NPR1 sends a response 118. The second control device 220b will also be responsible for controlling the industrial system 210, although this is in Figure 1C There was no explicit statement from China.
[0073] During the interval between the transition 130 from the main controller to the backup controller in the first control device 220a and the transition 132 from the backup controller to the main controller in the second control device 220b, both control devices 220a and 220b will function as backup controllers. This dual-backup condition is not harmful, especially when the industrial system 210 (or its input / output interfaces) is configured to apply predetermined control signal values (output set to the predetermined value OSP) in the absence of external control signals. The OSP signal will guide the industrial system 210 into a safe state, although this may not necessarily be useful or effective.
[0074] exist Figure 1D Initially, the first control device 220a is used as the main controller, the second control device 220b is used as the backup controller, and NRP1 has been designated as the NRP. In the illustrated embodiment, the period T1 of fault detection 114 and the period T3 of NRP response verification are equal. One cycle of ordered operation is shown, in which the main controller feeds control signal 110 to industrial system 210, feeds heartbeat message 112 to backup controller 220b, and feeds call 116 to NRP to verify whether NRP is operating. This will cause backup controller 220b to make a negative fault detection decision 114(N), and it will cause NRP to send response 118, which allows main controller 220a to conclude that NRP is operating, i.e., decision 122(P).
[0075] However, in the next cycle, NRP does not respond to call 116 from master controller 220a within T4 time units. Master controller 220a makes a negative authentication decision 122(N) and proposes node NRP2 (124) as a replacement NRP to backup controller 220b. Backup controller 220b implicitly rejects the proposal by not accepting it for T5 time units. As a result, first control device 220a will switch from master controller 130 to backup controller and thus stop sending heartbeat messages 112. Then, after NRP1, whose authentication is still designated as NRP, responds to call 116, second control device 220b makes a positive failure decision 114(P) and switches 132 to master controller.
[0076] In the next cycle, the second control device 220b feeds control signal 110 to industrial system 210, feeds heartbeat message 112 to the first control device 220a (which has failed or is used as backup controller 220b), and feeds call 116 to verify whether NRP1 is operating. If NRP1 responds 118 to call 116, the second control device 220b can infer 122(P) that NRP1 is a valid NRP and no replacement needs to be proposed before further notification.
[0077] Figures 1E to 1G (This relates to an embodiment with NRP-mediated fault detection) is a sequence diagram illustrating the data exchange between control device 220, industrial system 210, and network nodes designated as NRPs in several example scenarios. These entities are interconnected via at least one data network 240 (preferably a wired data network). It is assumed that control devices 220a, 220b operate in a decentralized manner, i.e., without any assistance from a centrally coordinating entity.
[0078] exist Figure 1EInitially, the first control device 220a is used as the main controller, the second control device 220b is used as the backup controller, and NRP1 has been designated as the NRP. The initial designation of roles can initially be established through confirmation from the main controller issued by the operator. Therefore, the main controller 220a feeds control signals 110 to the industrial system while ensuring it has a valid lease with the NRP.
[0079] The main controller 220a sends a lease request 116-L (call) to the NRP and receives a lease confirmation 118-A (response) from the NRP. When the main controller 220a receives the lease confirmation 118-A, it knows that it has a lease Ls1 with the NRP. Lease Ls1 has a pre-configured duration known to the main controller 220a. T L .
[0080] When lease Ls1 is active, each lease query 116-Q (call) sent by backup controller 220b will cause NRP to return a positive lease response 118-R(P) (response). This allows backup controller 220b to conclude that master controller 220a is functioning correctly. Furthermore, when lease Ls1 is active, master controller 220a feeds control signals 110 to industrial system 210. As mentioned above, the messages lease request 116-L and lease query 116-Q can have a uniform format or two different formats; the option with a uniform format seems safer, although NRP can more easily handle messages with two different formats.
[0081] Figure 1E An example is shown where the first lease Ls1 is renewed immediately upon expiration, i.e., as a result, the main controller 220a sends another lease request 116-L (call) to the NRP and receives a lease confirmation 118-A (response). A second lease (with a duration equal to the first lease Ls1) is then renewed. Figure 1E In this context, Ls2 is used as the denoting element. Optionally, the lease request 116-L that initiates the first lease Ls1 can be the same as the subsequent lease request 116-L that renews it; otherwise, dedicated lease initiation and lease renewal request messages can be defined and used.
[0082] In addition, Figure 1EIn the example, when the second lease Ls2 expires (exp), it will not be renewed. When the backup controller 220b sends a lease query 116 Q, it will receive a negative lease response 118-R(N). Based on its configured behavior, the control device 220b concludes that the control device 220a (currently the primary controller) has failed and that it has acquired a lease Ls3 with the NPR, which allows it to become the primary controller. The control device 220b (currently the backup controller) can notify the control device 200a (currently the primary controller) that it has become the primary controller by sending a message directly to the control device 200a. Alternatively, the control device 220b can store the notification in the NPR, or the change can be implicit because the NPR now has a lease with the control device 220b instead of the control device 220a.
[0083] The next action of control device 220b (now used as the main controller) is to obtain a lease with NRP, which is done by sending a lease request 116-L to NRP and waiting for confirmation 118-A.
[0084] refer to Figure 1F The sequence diagram will describe the scenario where the main controller 220a replaces the NRP during operation. Initially, network node NRP1 is designated as the NRP.
[0085] The main controller 220a acquires a lease with NRP1 by sending a lease request 116-L to NRP and waiting for confirmation 118-A. The first lease Ls1.1 begins operation from the time of confirmation 118-A. As long as industrial system 210 has a valid lease, the main controller 220a is authorized to feed control signals 110 to it. The backup controller 220b sends a lease query 116-Q (call), which NRP responds to during the lease period with a positive lease response 118-R(P) (response). (It should be noted that the lease request 116-L and the lease query 116-Q can have a uniform message format.)
[0086] In this scenario, when the primary controller 220a requests the renewal of lease Ls1.1 by sending a second lease request 116-L, the NRP does not acknowledge the request. The primary controller 220a is configured to verify whether such an acknowledgment has arrived and to detect its absence. This causes the primary controller 220a to propose one or more replacement NRPs, including network node NRP2, to the backup controller 220b. The backup controller 220b responds by sending an acceptance 126 for NRP2 to the primary controller 220a.
[0087] Therefore, NRP2 is designated as NRP, and the main controller 220a can continue to function as the main controller. Optionally, the main controller 220a stores (not shown) a reference to NRP2 in NRP1, for example, by writing the identifier of NRP2 into the memory of NRP1. Thus, if the data network includes another backup controller that has not noticed the NRP change and sends a query to NRP1, the aforementioned backup controller will become aware of the NRP change and redirect accordingly to NRP2.
[0088] Next (regardless of whether the optional storage of references to NRP2 is performed), the main controller 220a continues to send lease requests 116-L to the newly appointed NRP2. When the new lease LS2.1 is confirmed, the main controller 220a may feed control signals 110 to the industrial system. During the duration of lease LS2.1, the NRP will respond to lease queries 116-Q from the backup controller 220b with a positive lease response 118-R(P).
[0089] Go to Figure 1G If the main controller 220a neither acquires a lease with the NRP nor identifies a successor to the NRP, the main controller 220a is configured to exit as the main controller and transform into a backup controller. Initially, the first control device 220a is used as the main controller, the second control device 220b is used as the backup controller, and NRP1 has been designated as the NRP.
[0090] The main controller 220a acquires a lease with NRP1 by sending a lease request 116-L to NRP and waiting for confirmation 118-A. The first lease Ls1.1 begins operation from the time of confirmation 118-A. As long as the industrial system 210 has a valid lease, the main controller 220a is authorized to feed control signals 110 to it. The backup controller 220b sends a lease query 116-Q (call), and the NRP responds to it during the lease period with a positive lease response 118-R(P) (response).
[0091] When the primary controller 220a later requests the renewal of lease Ls1.1 by sending a second lease request 116-L, the NRP does not acknowledge the request within the expected delay. This is either because the NRP has failed or because there is a network failure on the path between the primary controller 220a and the NRP. The primary controller 220a, configured to verify the arrival of such acknowledgment, notices the lack of acknowledgment and proposes at least one replacement NRP to the backup controller 220b. The backup controller 220b does not accept the proposed replacement NRP. Figure 1GIn one embodiment, the non-acceptance behavior is represented by not accepting a message within a expected time frame, i.e., implicitly; in other embodiments, the acceptance behavior can be implicitly signaled, while non-acceptance (rejection) can be explicit. Control device 220a, acting as the master controller, is configured to switch to a backup controller in response to this result. Related to this switch, control device 220b, which has so far been operating as a backup controller, assumes the role of the master controller.
[0092] The next action of control device 220b is to acquire a lease Ls1.2 with the NRP by sending a lease request 116-L to the NRP and waiting for confirmation 118-a. From the point when control device 220b receives the lease confirmation 118-A, it acts as the master controller and is therefore authorized to feed control signals 110 to industrial system 210. If lease Ls1.2 expires, control device 220b will cease to act as the master controller, or in some implementations, it will cease to act as the master controller later to allow one of the backup controllers sufficient time to become the master controller. Control device 220a, now acting as the backup controller, sends a lease query 116-Q (call), which the NRP responds to during lease Ls1.2 with a positive lease response 118-R(P) (response).
[0093] Figures 2 to 15 An example network topology and various fault locations that can occur are shown, using the same reference numerals as in Figure 16.
[0094] Starting with the current state of technology, Figure 2 An arrangement of two redundant control devices 220a and 220b is shown, which are not configured according to the full teachings of this disclosure. When one of the control devices 220a and 220b acts as the master controller, it feeds control signals to the industrial system 210 via a data network 240 and one or more of the switches 230a, 230b, 230c, and 230d arranged therein. The first control device 220a may send control signals to the industrial system 210 via a first path including the upper left switch 230a or via a second path including the lower left switch 230c. The second control device 220b, in its own right, may send control signals to the industrial system 210 via a first path including the two upper switches 230a and 230b or via a second path including the two lower switches 230d and 230c. The sending control device may be able to influence the routing of the control signals, or this may be determined by the network infrastructure, including switch 230.
[0095] It should also be noted that Figure 2The first control device 220a and the second control device 220b are connected via an upper path through upper network switches 230a and 230b, and a lower path through lower network switches 230c and 230d. Each of the first control device 220a and the second control device 220b has one port terminating the upper path and one port terminating the lower path. In this sense, the port acts as a route between the first control device 220a and the second control device 220b. Ports can have different IP addresses, or the network protocol may allow for individual addressing of ports in other ways.
[0096] Figure 3 A scenario depicting two failures in data network 240 is illustrated, as shown in red. The failure disrupts the first path from the first control device 220a to the industrial system 210, but not the second path. These failures also disrupt all paths between the first control device 220a (the master controller) and the second control device 220a (the backup controller), potentially leading the backup controller to incorrectly conclude that it should become the master controller. If control devices 220a and 220b are configured according to existing technology, this would result in an undesirable dual-master condition.
[0097] Go to Figure 4 If control devices 220a and 220b are configured according to this disclosure and jointly specify an NRP, the risk of a dual-master condition is significantly reduced. The first control device 220a starts as the master controller. When a network failure causes the NRP to become unreachable, it becomes the backup controller. Simultaneously, the second control device 220b starts as the backup controller and becomes the master controller when the first control device 220a appears to be quiescent and the NRP is reachable. The specified NRP candidate does not need to participate in this scenario. For detailed information on signaling, please refer to [link to relevant documentation]. Figure 1A The sequence diagram in [the text]. According to... Figure 4 The examples are applicable to embodiments with NRP boot failure detection and embodiments with NRP-mediated failure detection.
[0098] exist Figure 5 The diagram illustrates the link between the first control device 220a and the top-left switch 230a, which serves as the NRP endpoint. NRP reachability can be verified by sending a ping (NRP ping) on this link. The ping can conform to ICMP, which is not implemented in hard real-time devices and therefore cannot guarantee a bounded response time. As described above, this is achieved by monitoring heartbeat signals on two different network paths, such as reference... Figure 2The description of the upper and lower paths between the first control device 220a and the second control device 220b, and the determination of whether timeouts occur simultaneously, can mitigate this deficiency to some extent. It should be understood that the switch 230 is configured to forward (transmit) heartbeat signals, while the control device 220 is typically not. According to... Figure 5 The example applies to embodiments with NRP boot failure detection.
[0099] In some embodiments, such as Figure 6 As shown, each device connected to data network 240 can learn the IP address of its nearest switch 230 via the LLDP protocol upon startup. In fact, under LLDP, each switch 230 periodically announces itself using a standardized protocol supported by most commercial network devices (signal 1). Furthermore, LLDP can provide information for determining whether switch 230 supports real-time ping (i.e., not just ICMP ping). Figure 6 The examples apply to embodiments with NRP-guided fault detection and embodiments with NRP-mediated fault detection. Control device 220a, acting as the master controller, notifies the backup controller of its NRP (signal 2); this notification may optionally be included in a heartbeat signal. If the NRP is not successfully specified, the system owner can be notified accordingly via any suitable diagnostic function associated with the data network 240.
[0100] exist Figure 7 In this scenario, the first control device 220a, acting as the primary controller, starts up and malfunctions, failing to send control signals to the industrial system 210. The data network 240 operates normally. The second control device 220b starts up as a backup controller and, after noticing that the heartbeat timeout period has expired and the NRP can be successfully reached, it becomes the primary controller. The end result of these events is that the former primary controller remains in a silent state, such as in an error state or a power failure state, while the former backup controller assumes the role of primary controller. Figure 7 The example applies to embodiments with NRP boot failure detection.
[0101] Alternatively, if the second control device 220b detects heartbeat messages from the first control device 220a on both the upper and lower paths, and these messages do not share a common cause of failure, then once it determines (e.g., based on timestamps) that the heartbeat timeouts are simultaneous on both paths, the second control device 220b can switch to becoming the primary controller. The absence of heartbeat messages is unlikely to be due to two network failures (one on each path) occurring almost simultaneously. Therefore, NRP checks can be omitted without causing harm.
[0102] exist Figure 8In this configuration, the first control device 220a starts up as the primary controller, and the second control device 220b starts up as the backup controller. Subsequently, a network failure 1 occurs at a designated location, which interrupts the downstream path but does not affect the primary controller's ability to feed control signals to the industrial system 210. Due to redundancy, the backup controller continues to receive heartbeat messages via the upstream path; the network failure may not be noticed by the backup controller, and it may not necessarily trigger any action in the backup controller.
[0103] Figure 9 yes Figure 8 Continuing the scenario, if the primary controller subsequently fails 2, preventing it from sending control signals to industrial system 210, the backup controller will notice this after the timeout period. Because the backup controller will receive a response from the NRP via the upper path, it will decide to switch to primary controller. Figure 8 and Figure 9 The example applies to an embodiment of fault detection with NRP guidance.
[0104] exist Figure 10 In this configuration, the first control device 220a starts up as the primary controller, and the second control device 220b starts up as the backup controller. A network failure 1 occurs at the indicated location, disrupting the upper path and also the link between the first control device 220a and the upper left switch 230a, which acts as the NRP. The first control device 220a (the primary controller) will react by proposing one or more NRP candidates, which the second control device 220b can accept or (implicitly) reject. If the lower left switch 230c proves acceptable, it is designated as the NRP, effective for both the first and second control devices 220a and 220b. Figure 10 The examples are applicable to embodiments with NRP boot failure detection and embodiments with NRP-mediated failure detection.
[0105] exist Figure 11 In this configuration, the first control device 220a starts up as the master controller, and the second control device 220b starts up as the backup controller. This time, the top left switch 230a, designated as the NRP, stops operating. In embodiments with NRP boot failure detection, the master controller will detect NRP failures during its regular verification process (see [link to documentation]). Figure 1B In embodiments with NRP-mediated fault detection, the primary controller will detect the NRP's failure when it attempts to renew the lease. In either case, the primary controller proposes designating the lower left switch 230c as the replacement NRP, and this is accepted by the backup controller. The primary / backup roles remain unchanged.
[0106] exist Figure 12In this configuration, the first control device 220a starts up as the primary controller, and the second control device 220b starts up as the backup controller. The top-left switch 230a is designated as the NRP. A network failure 1 occurs at the indicated location, disrupting the upper path between the first control device 220a and the second control device 220b, but leaving the lower path intact. However, network failure 1 interrupts the only path from the NRP to the second control device 220b (the backup controller). In this embodiment, the backup controller routinely verifies the reachability of the NRP. Since this verification results in a negative result, due to network failure 1, the backup controller requests the primary controller to replace the NRP. In other words, the top-left switch 230a is no longer an acceptable NRP. The primary controller proposes the bottom-left switch 230c (one of its NRP candidates), which is accepted by the backup switch. The primary / backup roles remain unchanged. Figure 12 The example applies to embodiments with NRP-guided fault detection. For embodiments with NRP-mediated fault detection, similar behavior is expected, except that the backup controller does not routinely verify NRP reachability, but instead verifies whether the primary controller has a valid lease with the NRP.
[0107] exist Figure 13 In this configuration, the first control device 220a starts up as the master controller, and the second control device 220b starts up as the backup controller. The top left switch 230a is designated as the NRP. Two network failures, 1 and 2, occur at the indicated locations, disrupting the upper and lower paths between the first control device 220a and the second control device 220b. This dual network failure is an example of splitting the data network 240 into separate components. Despite the network failures, the master controller can still feed control signals to the industrial system 210 through two disjoint paths in the data network 240. When the backup controller notices the absence of a heartbeat message, it attempts to contact the NRP. Because the NRP does not respond, the backup controller does not transform into the master controller. (The above "special case" is not satisfied because network failures 1 and 2 do not occur simultaneously in terms of timestamp resolution.) In some embodiments, the backup controller is configured to transition to a startup-initialization / wait state 1710 after determining that it cannot receive a heartbeat signal and cannot reach the NRP; upon receiving a heartbeat signal again, it transitions to a backup controller-supervisory state 1720 (backup controller role). In these events, the first control device 220a continues to operate in the master controller-supervised mode 1730 (master controller role). According to Figure 13 The example is immediately applicable to embodiments with NRP boot failure detection.
[0108] exist Figure 13In the context of scenarios and NRP-mediated fault detection, the backup controller will be unable to reach the NRP, thus noticing a dual network failure. It must find a node in the data network that, instead of the unreachable NRP, can act as a backup NRP if it is to become the primary controller.
[0109] Figure 14 The scenario also involves a dual network failure, where data network 240 is split into separate components (partitions). A first control device 220a starts up as the primary controller, and a second control device 220b starts up as the backup controller. The top-left switch 230a is designated as the NRP. Two network failures, 1 and 2, occur at the indicated locations, disrupting the upper and lower paths between the first control device 220a and the second control device 220b. Network failure 2 also disrupts the path from the primary controller to the NRP. In an embodiment with NRP-guided failure detection, this will be detected during the primary controller's routine verification of NRP reachability, after which it will propose a different switch 230 as the NRP. Similarly, in an embodiment with NRP-mediated failure detection, the primary controller will notice the impact of network failure 2 when attempting to renew a lease with the NRP. However, the bottom-left switch 230c is the only NRP candidate reachable by the primary controller but unreachable by the backup controller, and it will be (implicitly) rejected by the backup controller. As a result, the first control device 220a will transition to a startup-initialization / wait state 1710, and the second control device 220b will transition to a master controller-supervised mode 1730. Therefore, the second control device 220b will operate without supervision until the data network 240 is at least partially restored.
[0110] exist Figure 15 In this configuration, the first control device 220a starts up as the master controller, and the second control device 220b starts up as the backup controller. The top left switch 230a is designated as the NRP. Consider a scenario where the NRP stops working and there is a network failure 1 on the lower path between the first control device 220a and the second control device 220b. During routine verification, the master controller will find that it can no longer reach the NRP and will propose alternative NRPs from the NRP candidates. Similarly, the bottom left switch 230c is the only NRP candidate that the master controller can reach but the backup controller cannot, and it will be (implicitly) rejected by the backup controller. As a result, the first control device 220a will transition to the startup-initialization / waiting state 1710. The second control device 220b, acting as the backup controller, will find that, on the one hand, the heartbeat signal is lost on both the upper and lower network paths, and on the other hand, it cannot reach the NRP. Therefore, the second control device 220b will also transition to the startup-initialization / waiting state 1710. Industrial system 210 will operate according to the predetermined output set (OSP) until network 240 recovers. Figure 15 The example applies to embodiments with NRP boot failure detection.
[0111] The foregoing has primarily described various aspects of this disclosure with reference to several embodiments. However, as will be readily understood by those skilled in the art, other embodiments besides those disclosed above are also possible within the scope of the invention as defined by the appended claims.
Claims
1. A control device (220) for use with at least one additional control device when controlling an industrial system (210), said control device and said additional control device being connected to said industrial system via a data network (240), The control device described herein is capable of operating to - Used as a main controller (1730), wherein the control device feeds control signals (110) to the industrial system; and - Used as a backup controller (1720), wherein the control device routinely performs fault detection (114) on the main controller via the data network, and switches to the main controller in response to a positive fault detection. Its features are, The transition from the backup controller to the main controller depends on whether the verification network reference point (NRP) responds (118) to a call (116) from the backup controller, wherein the NRP is the node in the data network that connects the main controller and the backup controller to the industrial system.
2. The control device (220) according to claim 1, wherein: When used as a main controller, the control device emits a heartbeat signal (112); and When used as a backup controller, the control device performs the fault detection by detecting the heartbeat signal.
3. The control device (220) according to claim 2, wherein the heartbeat signal comprises a timestamped heartbeat message stream.
4. The control device (220) according to claim 3, wherein when used as a backup controller, in response to the simultaneous interruption of the timestamped heartbeat message stream on at least two different paths of the data network (240), the control device transforms into the master controller, regardless of whether the NRP responds (118) to the call (116).
5. The control device (220) according to any one of the preceding claims, wherein: When used as a master controller, the control device initiates (116-L) a time-limited lease with the NRP, wherein the lease grants the holder of the lease the right to act as the master controller, and the lease is renewed; and When used as a backup controller, the control device performs the fault detection by querying the NRP (116-Q) to see if the control device has a lease with the main controller, the query being included in the call (116) from the backup controller.
6. The control device (220) according to claim 5, wherein when used as a master controller, the control device verifies the NRP confirmation (118A) of the lease in response to initiation (116-L) or renewal.
7. The control device (220) according to claim 6, wherein when used as a main controller, the control device: If the NRP does not confirm the lease, the switch (130) is converted to a backup controller.
8. The control device (220) according to claim 6, wherein when used as a main controller, the control device: If the NRP fails to confirm the lease, one or more replacement NRPs are proposed to the backup controller (124); and If at least one additional control device used as a backup controller does not accept any of the replacement NRPs in the replacement NRPs described in (126), then (130) is transformed into a backup controller.
9. The control device (220) according to any one of claims 5 to 8, wherein: The control device is configured to initiate or renew (116-L) the lease with the NRP using a common message format when acting as a primary controller, and to query (116-L) the NRP when acting as a backup controller.
10. The control device (220) according to any one of the preceding claims, wherein the NRP is a node in the data network (240), and the node has been jointly designated (128) as an NRP by the control device.
11. The control device (220) according to claim 10 is further configured to propose (124) a node in the data network (240) to the additional control device for designation (128) as an NRP, wherein the control device and the node have no known common cause of failure.
12. The control device (220) according to claim 10 or 11 is further configured to propose (124) a node in the data network (240) to the other control device for designation (128) as an NRP, the node being located on a path from the control device itself to at least one other control device.
13. The control device (220) according to any one of the preceding claims, wherein when used as a master controller, the control device routinely verifies (122) whether the NRP responds (118) to a call (116) from the master controller.
14. The control device (220) according to claim 13, wherein when used as a main controller, the control device: If the NRP does not respond, one or more replacement NRPs are proposed to the backup controller (124); and If at least one additional control device used as a backup controller does not accept any of the replacement NRPs in the replacement NRPs described in (126), then (130) is transformed into a backup controller.
15. The control device (220) according to claim 13 or 14, wherein the period (T1) of fault detection (114) and the period (T3) of verification of the NRP response (122) differ by a maximum of 10 times, preferably by a maximum of 5 times, preferably by a maximum of 2 times, and most preferably approximately equal.
16. The control device (220) according to any one of the preceding claims, wherein the call (116) from the backup controller to the NRP is a ping message, the ping message being specified to have a bounded response time in a protocol executed by the data network (240).
17. The control device (220) according to any one of the preceding claims, wherein when used as a backup controller, in response to the detection that the NRP is not responsive (118) to a call (116) from the backup controller, the control device enters a waiting state (1710).
18. A method for operating redundant control devices (220) in controlling an industrial system (210), the control devices and additional control devices being connected to the industrial system via a data network (240), the method comprising: At the first control device (220a) currently used as the main controller, a control signal (110) is fed to the industrial system. At the second control device (220b) currently used as a backup controller, a routine fault detection (114) is performed on the main controller. At the second control device, in response to a positive fault detection, a call (116) is made to the node of the data network acting as a network reference point (NRP); and At the second control device, a decision (120) is made to switch to the master controller only when a response (118) from the NRP is received.
19. The method of claim 18, further comprising: At the first control device (220a) currently used as the main controller, routine verification (122) is performed to determine whether the NRP responds (118) to a call (116) from the main controller.
20. The method of claim 19, further comprising: At the first control device (220a) currently used as the main controller, in the absence of a response (118) from the NRP, one or more replacement NRPs are proposed (124) to the backup controller; and One of the replacement NRPs proposed by all control devices (128).
21. The method of claim 20, further comprising: At the first control device (220a) currently used as the main controller, in the absence of a response (118) from the NRP, one or more replacement NRPs are proposed (124) to the backup controller; as well as At the first control device, if at least one backup controller does not accept the replacement NRP proposed by (126), (130) is transformed into a backup controller.
22. A computer program comprising instructions for causing a computer to behave as a control device (220) according to claim 1.
23. A computer program comprising instructions for causing a computer to perform the method of claim 18.
Citation Information
Patent Citations
A method for failure detection and role selection in a network of redundant processes
EP3933596A1
Configuring redundancy in a supervisory process control system
US20060056285A1
Distinguishing between link and node failure to facilitate fast reroute
US7986618B2
Technique for distinguishing between link and node failure using bidirectional forwarding detection (BFD)
US8082340B2