Detection of component degradation in industrial process plants based on loop component responses
The system detects control loop component degradation in industrial process plants by monitoring round trip times using heartbeat messages, addressing the challenge of undetected performance issues and preventing catastrophic events.
Patent Information
- Application Number
- JP2021203168
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2021-01-14
- Filing Date
- 2021-12-15
- Publication Date
- 2025-12-11
- Estimated Expiration
- 2041-12-15
Smart Images

Figure 0007784284000001 
Figure 0007784284000002 
Figure 0007784284000003
Abstract
Description
[Technical Field]
[0001] This application relates generally to industrial process control systems for industrial process plants, and more particularly to industrial process control systems capable of detecting degradation of control loop components. [Background technology]
[0002] A distributed industrial process control system, such as those used to manufacture, refine, transform, create, or produce physical materials or products in chemical, petroleum, industrial, or other process plants, typically includes one or more process controllers communicatively coupled to one or more field devices via a physical layer that may be an analog bus, a digital bus, or a combined analog / digital bus, or that may include one or more wireless communication links or networks. The field devices may be, for example, valves, valve positioners, switches, and transmitters (e.g., temperature, pressure, water level, and flow sensors) located within the process environment of the industrial process plant (interchangeably referred to herein as the “field environment” or “plant environment” of the industrial process plant) and generally perform physical process control functions, such as opening and closing valves, measuring process and / or environmental parameters such as flow rate, temperature, or pressure, and controlling one or more processes running within the process plant or system. Smart field devices, such as field devices conforming to the well-known FOUNDATION® Fieldbus protocol, may also perform control calculations, alarm functions, and other control functions typically implemented within process controllers.
[0003] A process controller, which may or may not be physically located within a plant environment, receives signals indicative of process measurements made by field devices and / or other information related to the field devices, executes control routines, applications, or logic that operate various control modules that utilize various control algorithms to make process control decisions, and generates process control signals based on the received information, interfacing with control modules or blocks implemented within field devices such as HART® field devices, WirelessHART® field devices, and FOUNDATION® Fieldbus field devices. To accomplish this communication, control modules within the process controller send control signals to a variety of different input / output (I / O) devices, which then transmit these control signals over dedicated communication lines or links (communication physical layers) to the actual field devices, thereby controlling the operation of at least a portion of a process plant or system, e.g., at least a portion of one or more industrial processes operating or executing within the plant or system. Thus, a process control loop refers to the process controller, one or more I / O devices, and one or more field devices controlled by the process controller via signals and / or data delivered to and from the field devices via the I / O devices. As used herein, the term "process control loop" may be interchangeably referred to as a "control loop" or a "loop," and as used herein, the term "process controller" may be interchangeably referred to as a "controller" or a "control device." A process controller or control device may be a physical device or a virtual device. For example, one or more logic or virtual process controllers may be allocated to run or operate on a physical host server or computing platform. A process control system may include both physical and virtual process controllers.
[0004] Additionally, I / O devices, typically located within a plant environment, are generally disposed between a process controller and one or more field devices to enable communication therebetween, for example, by converting electrical signals to digital values and vice versa. Different I / O devices are provided to support field devices that use different proprietary communication protocols. More specifically, in some configurations, different physical I / O devices are provided between the process controller and each of the field devices that use different communication protocols, such that a first I / O device is used to support HART field devices, a second I / O device is used to support Fieldbus field devices, a third I / O device is used to support Profibus field devices, and so on.
[0005] In some configurations, instead of using individual physical I / O devices to distribute data between process controllers and their respective field devices, an I / O gateway can be disposed between multiple process controllers and their corresponding field devices, where the I / O gateway switches or routes I / O data between each of the process controllers and their corresponding field devices. Thus, in these configurations, a control loop can include a process controller or control device, an I / O gateway, and one or more field devices. The I / O gateway can be communicatively connected to the field devices via various physical ports and connections supporting the respective industrial communication protocols utilized by the field devices (e.g., HART, Fieldbus, Profibus, etc.), and the I / O gateway can be communicatively connected to the process controllers via one or more communication networks or data highways (e.g., supporting wired and / or wireless Ethernet, Internet Protocol (or IP), and / or other types of packet protocols, etc.). An I / O gateway may be implemented at least partially on one or more computing platforms and configured to distribute, route, or switch I / O data between multiple process controllers and their respective field devices, thereby accomplishing I / O data distribution for process control. That is, the I / O gateway may enable a control device to coordinate one or more corresponding field devices by using network functionality provided by the I / O gateway. For example, within an I / O gateway, various I / O data distribution functions, routines, and / or mechanisms may be hosted on one or more servers and utilized to switch I / O data between ports communicatively connecting control devices to the I / O gateway and ports reliably connecting field devices to the I / O gateway.
[0006] As used herein, field devices, controllers, and I / O devices or gateways are generally referred to as "process control devices." Field devices and I / O devices are generally located, located, or installed in the field environment of a process plant. Control devices or control devices can be located, located, or installed in the field environment and / or in the back-end environment of a process plant. I / O gateways can be at least partially installed in the field environment and at least partially installed in the back-end environment of a process plant.
[0007] Additionally, information from the field devices and process controllers is typically made available via a data highway or communication network through the process controller to one or more other hardware devices, such as operator workstations, personal computers or computing devices, data historians, report generators, centralized databases, or other centralized computing devices typically located in a control room or other location away from the plant's harsher field environment, e.g., the back-end environment of a process plant. Each of these hardware devices is typically centralized throughout the process plant or throughout portions of the process plant. These hardware devices run applications that may enable operators to perform functions related to controlling the process and / or operating the process plant, such as, for example, changing settings in process control routines, modifying the operation of control modules in controllers or field devices, displaying the current state of the process, displaying alarms generated by field devices and controllers, simulating the operation of the process for purposes of training personnel or testing process control software, maintaining and updating configuration databases, etc. The data highways utilized by the hardware devices and process controllers may include wired communication paths, wireless communication paths, or a combination of wired and wireless communication paths, and typically use packet-based communication protocols and non-time-sensitive communication protocols such as Ethernet or IP protocols.
[0008] As an example, the DeltaV™ control system sold by Emerson Process Management includes multiple applications stored in and executed by different devices located at various locations within a process plant. Configuration applications resident in one or more workstations or computing devices allow users to create or modify process control modules and download them to dedicated distributed controllers via a data highway. Typically, these control modules are composed of communicatively interconnected function blocks, which may be objects in an object-oriented programming protocol that perform functions within the control scheme based on inputs to the control scheme and provide outputs to other function blocks within the control scheme. The configuration application may also allow configuration engineers to create or modify operator interfaces that are used by viewing applications to display data to an operator and allow the operator to change settings, such as setpoints, within process control routines. Each dedicated controller, and in some cases, one or more field devices, stores and executes a respective controller application that runs the control modules assigned to it and downloaded (or otherwise acquired by the controller) to implement the actual process control functions. The viewing application may run on one or more operator workstations (or one or more remote computing devices in communicative connection with the operator workstations and the data highway) and may receive data from the controller application via the data highway and display this data to a process control system designer, operator, or user using a user interface to provide any of several different views, such as an operator's view, an engineer's view, a technician's view, etc.While the data historian application is typically stored on and executed by a data historian device that collects and stores some or all of the data provided over the data highway, a configuration database application may be run on an even further remote computer attached to the data highway to store the current process control routine configuration and its associated data. Alternatively, the configuration database may be located on the same workstation as the configuration application.
[0009] Process control devices and process control loops require a great deal of planning and configuration to ensure not only the performance and safety of the device, but also the performance and safety of the entire plant. Such planning and configuration includes taking into account the response performance or responsiveness of the process controller or control device. Consider the example of a process plant that heats a volatile chemical (e.g., crude oil or gasoline) to a specific temperature. The control device (e.g., process controller, safety system controller, etc.) may be configured to monitor the temperature as quickly as once every 50 milliseconds to ensure that the gasoline does not heat up to a self-flammable temperature. If at any time the control device detects that the self-flammable temperature is being reached, the control device responds very quickly (e.g., within 50 milliseconds) to reduce the intensity of (or shut off) the heating element, preventing the gasoline from exploding while attempting to maintain the heating at a safe temperature. The heating element controlled by the control device is an example of a “final control element.” A final control element is typically a device or component that changes its behavior in response to a control command or instruction, thereby moving the value of a controlled variable toward a desired setpoint and thereby controlling at least a portion of an industrial process. The final control element may be, for example, an I / O device or a field device communicatively connected to the control device via an I / O gateway.
[0010] The ability of a control device to respond within required parameters can be verified during its initial installation. However, over time, the control device may experience degradation in response performance or responsiveness due to, for example, software design, insufficient computing resources, obsolete and / or deteriorating hardware, communication interruptions (e.g., below thresholds detectable by diagnostic processes), environmental conditions, etc. Other components of the control loop may similarly degrade. The consequences of such degradation of control loop components can be disastrous. For example, in the gasoline heating example discussed above, if the control device suffers from degraded responsiveness and / or the I / O gateway is overloaded, the control loop including the control device, I / O gateway, and heating element may not be able to shut down the heating element within the required time (e.g., within 50 milliseconds), and the heating element may continue to heat the gasoline above the temperature threshold for self-combustion.
[0011] In computerized process control systems (e.g., process control systems in which at least some control devices, I / O devices or gateways, and / or other control loop components are implemented on one or more computing platforms and shared resources provided by the one or more computing platforms), degradation of response performance of control devices, I / O devices or gateways, and / or other control loop components can be particularly affected by scarcity or utilization of hardware and / or software computing resources of the supporting platforms. For example, scarcity of central processing unit (CPU) resources, scarcity of memory, scarcity of persistent storage space, contention for networking resources, contention for logic resources, and / or scarcity of other computing resources in the various computing platforms can degrade performance of the control loop and / or various components of the control loop.
[0012] For example, a virtualized or containerized control device may be assigned to operate or run on a host server or computing platform along with other virtualized / containerized control devices. If too many virtualized control devices are operated on a single host, the hosted virtualized control devices must compete for CPU, memory, and disk resources provided by the host. Even in situations where overall host CPU availability appears sufficient (e.g., 30% availability), contention between virtualized control devices for other types of host resources (e.g., scheduling of control device instances, control device algorithm logic, etc.) may cause degradation of control performance of one or more of the virtualized control devices. Furthermore, in configurations where virtualized control devices are implemented using virtual machines, issues such as hypervisor load and / or CPU utilization may also affect the responsiveness performance of the virtualized control devices.
[0013] In a similar manner, because an I / O gateway may be disposed between a control device and a corresponding final control element, the responsiveness of a control loop may be affected by strain on the I / O gateway and / or contention for network resources (e.g., software and / or hardware network resources) provided by the I / O gateway for I / O data delivery. For example, I / O gateway strain and / or resource contention (e.g., scheduling, I / O data delivery logic, computing resources, etc.) may increase the latency and / or jitter of the control loop. Furthermore, I / O gateway strain and / or resource contention may not only adversely affect the performance of the control loop, but may also adversely affect the performance of the overall system. For example, increased latency introduced by the I / O gateway may degrade the overall operation of the process control system, and increased jitter introduced by the I / O gateway may degrade the performance of the overall control strategy.
[0014] Unfortunately, because known diagnostic procedures are typically configured to perform at a much slower rate than process control and because the diagnostic procedures must be configured for a specific situation, some types of control loop component performance degradation may remain undetected for some time. In fact, some types of control loop component degradation may only be detected after the control system or process plant experiences a catastrophic event and / or failure to fulfill business needs. Thus, users of control systems may suffer losses in productivity, profits, equipment, capital, and / or even human life due to delayed and / or undetected performance degradation of control loop components. Summary of the Invention
[0015] Disclosed are techniques, systems, and / or methods for detecting degradation of industrial process plant components based on the responsiveness of the loop components. Generally, the responsiveness of loop components (e.g., process controllers, safety system controllers, I / O gateways, I / O devices, etc.) can be monitored for degradation as the loop components operate during operation and control at least a portion of an industrial process. As a result, degradation of a loop component may be detected when it begins to occur, rather than after a hard failure of the component or when a catastrophic event occurs. The techniques, systems, and methods disclosed herein can alert operating personnel when degradation of a loop component is detected. In some embodiments, the techniques, systems, and methods provide automatic degradation mitigation, thereby automatically containing, minimizing, or even eliminating the impact of a detected degradation on system performance and safety.
[0016] In one embodiment, a system for detecting component degradation in an industrial process plant includes a first component and a second component of a process control system of the industrial process plant, where the first component and the second component are communicatively connected via a diagnostic channel and a communication channel. The first component and the second component may be included in a process control loop. For example, the first component may be an I / O gateway or one of a process controller included in a plurality of process controllers communicatively connected to the I / O gateway via respective communication channels, and the second component may be the other of the I / O gateway or the process controller. The I / O gateway communicatively connects the plurality of process controllers to respective one or more field devices, thereby controlling the industrial process within the process plant, and
[0017] The first component of the system is configured to sequentially transmit a plurality of heartbeat messages to the second component via the diagnostic channel and to receive, via the diagnostic channel, at least a subset of the plurality of heartbeat messages returned by the second component to the first component upon receipt at the second component. The first component of the system is further configured to determine an average response time of the second component based on a round trip time (RTT) of each of the at least a subset of the plurality of heartbeat messages, where the respective RTTs are determined based on respective times of transmission and reception at the first component of the at least a subset of the plurality of heartbeat messages. The first component is further configured to detect degradation of the second component when the RTT of a subsequent heartbeat message sent by the first component to the second component via the diagnostic channel exceeds a threshold corresponding to the average response time of the second component. In some configurations, the first component is additionally or alternatively configured to detect degradation of the second component when an RTT of a subsequent heartbeat message sent by the first component to the second component via the diagnostic channel exceeds a threshold corresponding to a periodicity of module scheduler execution at the second component.
[0018] In one embodiment, a method for detecting component degradation in a process control system includes sequentially transmitting multiple heartbeat messages from a first component of the process control system to a second component of the process control system via a diagnostic channel. The first component and the second component are communicatively connected via the diagnostic channel and the communication channel, and the first component and the second component may be included in the same process control loop. For example, the first component may be an I / O gateway or one of a plurality of process controllers communicatively connected to the I / O gateway via respective communication channels, and the second component may be the other of the I / O gateway or the process controller. The I / O gateway communicatively connects the plurality of process controllers to respective one or more field devices, thereby controlling an industrial process within the process plant.
[0019] The method additionally includes receiving at least a subset of the plurality of heartbeat messages at the first component from the second component via a diagnostic channel, wherein each heartbeat message of the at least the subset of the plurality of heartbeat messages is transmitted by the second component back to the first component upon respective reception at the second component. Furthermore, the method includes determining, by the first component, an average response time of the second component based on a round trip time (RTT) of each of the at least the subset of the plurality of heartbeat messages. The respective RTTs may be determined based on respective times of transmission and reception of the at least the subset of the plurality of heartbeat messages at the first component. Still further, the method includes detecting degradation of the second component when the RTT of a subsequent heartbeat message transmitted by the first component to the second component exceeds a threshold corresponding to the average response time of the second component. In some configurations, the method additionally or alternatively includes detecting degradation of the second component when an RTT of a subsequent heartbeat message sent by the first component to the second component via the diagnostic channel exceeds a threshold corresponding to a periodicity of module scheduler execution in the second component. [Brief explanation of the drawings]
[0020] [Figure 1] FIG. 1 is a simplified block diagram of an example portion of a process control system of an industrial process plant configured to detect degradation of a control loop component. [Figure 2A] FIG. 1 illustrates an exemplary message flow for detecting degradation of a control loop component. [Figure 2B] FIG. 1 illustrates an exemplary message flow for detecting degradation of a control loop component. [Figure 3]FIG. 2 is a simplified block diagram of an example control loop component configured to detect degradation of another control loop component. [Figure 4] FIG. 1 is a flow diagram of an exemplary method for detecting degradation of a control loop component. DETAILED DESCRIPTION OF THE INVENTION
[0021] FIG. 1 is a simplified block diagram of an example portion 100 of a process control system for an industrial process plant. The portion 100 of the process control system shown in FIG. 1 includes multiple controllers 102a-102n, 105a-105m, communicatively connected to multiple field devices 108a-108p via an I / O gateway 110 (also interchangeably referred to herein as an “I / O server 110”). The controllers 102a-102n, 105a-105m, the I / O gateway 110, and the field devices 108a-108p work cooperatively during operation of the industrial process plant to control the industrial process of the industrial process plant. Other components of the process control system, such as configuration and other centralized databases, communication network architecture components, diagnostic and other tools, operator and engineer user interfaces, management computing devices, etc., are not shown in FIG. 1 for ease of explanation (and not by way of limitation).
[0022] At least some of the field devices 108a-108p may be final control elements, such as heaters, pumps, actuators, sensors, transmitters, switches, etc., and each field device is communicatively connected to an I / O gateway 110 via one or more respective wired and / or wireless links 112a-112p. The links 112a-112p are configured to operate safely within the harsh field environment of a process plant. The I / O gateway 110 is communicatively connected to each of the controllers 102a-102n, 105a-105m via a data highway 115, which may be implemented by utilizing one or more suitable high-capacity links, such as Ethernet, high-speed Ethernet (e.g., 100M Ethernet), optical links, etc. The data highway 115 may include, for example, one or more wired and / or wireless links. In some configurations, the field devices 108a-108p and / or at least one of the links 112a-112p to the data highway 115 support advanced physical layer (APL) transmission technology, which in turn supports one or more protocols to enable intrinsically safe connectivity of field devices, other devices, various other equipment, and / or other apparatus located in remote and hazardous locations, such as a process plant field environment.
[0023] The plurality of controllers 102a-102n, 105a-105m (also referred to interchangeably herein as "control devices" or "controllers") may include one or more physical controllers 102a-102n and / or one or more logic or virtual controllers 105a-105m, where each controller 102a-102n, 105a-105m controls a respective portion of an industrial process during plant operation by executing one or more respective control routines, control modules, or control logic. For example, some of the controllers 102a-102n, 105a-105m may be process controllers that execute a respective portion of an industrial process plant's operational control strategy. Some of the controllers 102a-102n, 105a-105m may be safety controllers operating as part of a safety instrumented system (SIS) that supports the process plant.
[0024] Each logic or virtual controller 105a-105m may be a respective virtualized, containerized, or other type of logic control device running or operating on a respective host server or computing platform 118a-118b. The set of host servers 118a-118b may be implemented using any suitable host server platform, such as multiple networked computing devices, a server bank, a cloud computing system, etc. Each server 118a-118b may host one or more respective logic control devices 105a-105m. For example, the logic control devices 105a-105m may be implemented via containers, which may be assigned to run on a particular host server, and / or the logic control devices 105a-105m may be implemented by virtual machines running within a hypervisor on a particular host server.
[0025] As previously mentioned, the I / O gateway 110 is communicatively coupled to the field devices via various physical ports and connections 112a-112p that support the respective industrial communication protocols utilized by the field devices (e.g., HART, Fieldbus, Profibus, etc.), and the I / O gateway 110 is communicatively coupled to the process controllers 102a-102n, 105a-105m via one or more communication networks or data highways 115, which may support wired and / or wireless Ethernet, Internet Protocol (i.e., IP), and / or other types of packet protocols, etc. The I / O gateway 110 may be at least partially implemented on one or more computing platforms and configured to distribute, route, or switch I / O data between the multiple process controllers 102a-102n, 105a-105m and their respective field devices 108a-108p, thereby executing respective control logic to accomplish process control. For example, within I / O gateway 110, various I / O data distribution functions, routines, logic, and / or mechanisms may be hosted on one or more servers and utilized to switch I / O data between ports communicatively connecting control devices 102a-102n, 105a-105m to I / O gateway 110 and ports communicatively connecting I / O gateway 110 to field devices 108a-108p. In one embodiment, each control device 102a-102n, 105a-105m may be a respective client of I / O gateway 110, with I / O gateway 110 servicing their respective I / O data distribution requests. Various hardware and / or software resources of I / O gateway 110 (e.g., CPU resources, memory resources, persistent storage space, disk space, network resources, logic resources, computing resources, etc., of the computing platform supporting I / O gateway 110) may be shared to service requests of multiple clients.
[0026] 1 includes one or more process control loops (also referred to interchangeably herein as "control loops" or "loops"), each of which includes a physical 102x or virtual 105y controller, an I / O gateway 110, and at least one field device 108z, which are generally referred to herein as "components" of the control loop. For example, the components of a first control loop may include a physical control device 102a, an I / O gateway 110, and a field device 108a; the components of a second control loop may include a physical control device 102b, an I / O gateway 110, and a field device 108b; the components of a third control loop may include a virtual control device 105g, an I / O server 110, and a field device 108c; and the components of a fourth control loop may include a virtual control device 105m, an I / O gateway 110, and a field device 108p. The components and control logic of each control loop may be defined or configured within the process control system, and the control devices for each loop may be configured with their respective control routines or control logic that they execute during operational operation. Typically, a particular final control element or field device 108a-108p may be configured or assigned to be controlled exclusively by only one control device 102a-102n, 105a-105m. Also, typically, but not necessarily, a particular control device 102a-102n, 105a-105m may be configured or assigned to control multiple final control elements or field devices 108a-108p.In a general sense, within a control loop, a controller 102a-102n, 105a-105m receives one or more input signals from one or more field devices 108a-108p, one or more other controllers 102-102n, 105a-105m, and / or one or more other devices within the process plant (e.g., via the data highway 115 and the I / O gateway 110), applies one or more control routines or control logic to the input signals to generate one or more output signals, and transmits the output signals to one or more field devices 108 in the control loop (e.g., via the data highway 115 and the I / O gateway 110), thereby modifying the behavior of the field devices 108 and thereby controlling at least a portion of an industrial process. In some configurations, the controller 102a-102n, 105a-105m also transmits one or more output signals to one or more other controllers 102a-102n, 105a-105m for the purpose of further controlling the industrial process.
[0027] Some of the control loop components shown in Figure 1 are specifically configured to detect degradation of one or more other components included in the loop, indicated as such in Figure 1 by a circled "DD." For example, as shown in Figure 1, control devices 102a, 105g, and 105h are specifically configured to detect degradation of I / O gateway 110, which in turn is specifically configured to detect degradation of any number of control devices 102a-102n, 105a-105m.
[0028] For purposes of explanation, FIG. 2A shows an example message flow 200 for detecting degradation of a control loop component. For ease of illustration, and not by way of limitation, FIG. 2A is discussed herein with simultaneous reference to FIG. 1. In FIG. 2A, a first loop component 202 and a second loop component 205 are configured to be included in the same control loop and are communicatively coupled via a communication channel, such as a communication channel of data highway 115. For example, first component 202 can be one of physical or logic controllers 102a-102n, 105a-105m, and second component 205 can be an I / O gateway or server 110, or first component 202 can be an I / O gateway or server 110, and second component 205 can be one of physical or logic controllers 102a-102n, 105a-105m.
[0029] The first component 202 and the second component 205 are also communicatively connected via a diagnostic channel, which may be another channel of the data highway 115 different from the communication channel. Generally, the first and second components 202, 205 send and receive control and communication messages or signals to each other to execute control strategies for the control loops via the communication channel and not via the diagnostic channel, and the first and second components 202, 205 send and receive diagnostic messages or signals (including messages / signals related to detecting component impairments) to each other via the diagnostic channel and not via the communication channel. In one embodiment, the diagnostic channel communicatively connecting the first and second components 202, 205 is utilized only by the first and second components 202, 205 of the control loops included within the process control system and is not otherwise utilized by any other components or devices of the process control system. That is, in this embodiment, the diagnostic channel is a channel dedicated for use only by the first and second components 202, 205 and is not shared by any other components or devices. The communication channel communicatively connecting the first and second control loop components 202, 205 may be a dedicated channel or a shared channel.
[0030] Message flow 200 illustrates messages delivered between a first component 202 and a second component 205 of a process control loop via a diagnostic channel to detect degradation of the second component 205. Thus, the first component 202 can be considered a monitoring or degradation-detecting component, and the second component 205 can be considered a monitoring or target component that the first component 202 monitors for degradation. As shown in FIG. 2A , the first component 202 sequentially sends multiple heartbeat messages HBn 208 (interchangeably referred to herein as “requests 208”) to the second component 205 over time. For example, the first component 202 may sequentially send multiple heartbeat messages HBn or requests 208 periodically, aperiodically, randomly, during times of relative hardware and / or software resource availability of the first component 202, on-demand or per user command, and / or at other suitable times. Each heartbeat message HBn 208 may be distinguished from other heartbeat messages HBn 208 via a respective identifier, such as a number, count, alphanumeric character, etc. In an embodiment, the heartbeat message identifier may increment and repeat periodically, e.g., n=1, 2, 3, . . . , x, 1, 2, 3, . . . , x, etc. Upon receiving a heartbeat message HBn 208 at the second component 205, the second component 205 forwards or otherwise returns the received heartbeat message HBn to the first component 202, as indicated by the numeral 210 (interchangeably referred to herein as a “response 210”).For example, the second component 205 may receive an incoming heartbeat message HBn or request 208 over a diagnostic channel, process the heartbeat message HBn or request 208 through a module or subcomponent of the second component 205 that processes incoming and outgoing process control and communication messages, and forward or return the heartbeat message HBn or response 210 to the first component 202 over the diagnostic channel at the fastest update rate supported by the second component 205. For example, if the second component 205 is a control device 102a-102n, 105a-105m, upon receipt of the request 208 at the control device 102a-102n, 105a-105m, a response module executing within the control execution scheduler of the control device 102a-102n, 105a-105m sends the return heartbeat message HBn or response 210 to the first component 202. In another example, if the second component 205 is an I / O gateway 110, upon receiving the request 208 at the I / O gateway 110, the RTT test initiator of the I / O server 110 sends a return heartbeat message HBn or response 210 to the first component 202. That is, the I / O gateway 110 treats the incoming heartbeat message HBn 208 as if the I / O gateway 110 were forwarding an I / O command, but the "forwarding" includes looping the heartbeat message HBn or request 208 back to the first component 202 as response 210.
[0031] The first component 202 tracks the respective times that it sends each heartbeat message HBn or request 208 and receives the corresponding return heartbeat message HBn or response 210, e.g., via timestamps TS1, TS2 as shown in FIG. 2A or via any other suitable mechanism. Based on the timestamps TS1, TS2, the first component 202 determines the round trip time (RTTn) 212 of each heartbeat message HBn delivered between the first component 202 and the second component 205. The RTTn 212 may be analogous to the time interval in a process control loop that includes a physical I / O device or I / O card in place of the I / O gateway 110, from when the control device sends an output command for I / O to the I / O device until the control device receives the corresponding confirmation return from the I / O device.
[0032] In an embodiment where the first component 202 is a physical control device 102a-102n or an I / O gateway 110, the hardware clock of the processor of the first component 202 may be utilized to determine the timestamps TS1, TS2. In an embodiment where the first component 202 is a logic control device 105a-105m and the second component 205 is an I / O gateway 110, however, if the logic control device 105a-105m is used to determine the timestamp TS2, TS2 may be inaccurate and inconsistent due to at least the nature of the virtual machine running on the hypervisor and / or the hosted architecture. In these embodiments, the timestamps TS1, TS2 may be determined in the manner shown in FIG. 2B.
[0033] FIG. 2B illustrates an example message flow 220 for detecting degradation of a control loop component. In FIG. 2B, the first loop component 202 is a logic control device 105a-105m, and the second loop component 205 is an I / O gateway or server 110. Similar to FIG. 2A, messages in the message flow 220 are delivered between the first and second components 202, 205 via a diagnostic channel. Also similar to FIG. 2A, the logic control device 202 sequentially sends multiple heartbeat messages HBn 208 to the I / O gateway 110. Because the I / O gateway 205 includes a processor with a hardware clock, the I / O gateway 205 can act as a proxy for the logic control device 202 for purposes of determining the respective timestamps TS1, TS2 corresponding to each received heartbeat message HBn. For example, the I / O gateway 205 can determine TS1 as the time of receipt of heartbeat message HBn at the I / O gateway 205, and the I / O gateway 205 can determine TS2 to be the time immediately before the I / O gateway 110 returned the received heartbeat message HBn 225 to the logic control device 202. The I / O gateway 205 can send TS1 and TS2 to the logic control device 202 along with the return heartbeat message HBn 225 (e.g., by inserting TS1 and TS2 into the return heartbeat message HBn 225 or by associating the heartbeat message HBn 225 with another transmission that includes TS1 and TS2), and the logic control device 202 can use the received timestamps TS1, TS2 to determine the corresponding round trip time RTTn 212 of the heartbeat message HBn.
[0034] In any event, regardless of whether message flow 200 or message flow 220 is utilized, in one embodiment, after a threshold or minimum number of sample RTTs are obtained or determined by the first component 202 during a quiescent or normal operating or operational state of the process control system 100, the first component 202 may determine an average or baseline RTT for heartbeat messages sent between the first and second components 202, 205, and optionally, a corresponding standard deviation. The threshold or minimum number of sample RTTs utilized to determine the average or baseline RTT may be predefined or preconfigured and may be dynamically adjustable, e.g., automatically and / or manually. The average or baseline RTT between the first and second components 202, 205 and the standard deviation may be stored in the first component 202. Generally speaking, the average or baseline RTT may provide a measure, level, or indicator of an expected or steady-state response or reaction time of the second component 205 during a quiescent or normal operating or operational state of the control system 100. The corresponding standard deviation may indicate a range of RTTs within which the response speed or reaction time of the second component 205 is considered deterministic or operating within a suitable performance range (e.g., the mean RTT minus one standard deviation to the mean RTT plus one standard deviation), e.g., a range of acceptable RTTs. Thus, a comparison of a subsequently measured or calculated RTTn 208 with the mean or baseline RTT and corresponding standard deviation may indicate a measure of the response performance level of the second component 205. For example, a measured RTT that is within plus or minus one standard deviation of the mean or baseline RTT may be considered an acceptable RTT, while a measured RTT that is outside plus or minus one standard deviation of the mean or baseline RTT may be considered an unacceptable RTT, which may indicate degradation in the second component 205.
[0035] Alternatively, in another embodiment where the target or second component 205 is a control device 102a-102n, 105a-105n, an acceptable RTT may be defined as an RTT that is less than or equal to one quantum time period (e.g., less than or equal to the periodicity length) of the control device's module scheduler execution. The threshold limit (e.g., the "periodicity-based threshold") for the acceptable RTT may be predefined or preconfigured and may be adjustable. For example, a user may set the periodicity-based threshold to a percentage from 90% of the periodicity length of the module scheduler execution up to and including 100% of the periodicity length. Thus, in this embodiment, any measured RTT that is less than or equal to the periodicity-based threshold may be an acceptable RTT for the target control device 205, and any measured RTT that is greater than the periodicity-based threshold may be an unacceptable RTT for the target control device 205.
[0036] In any event, after determining and storing the mean or baseline RTT (and corresponding standard deviation) for the target component 205 and / or saving the periodicity-based threshold for the target component 205, the first component 202 continues to send heartbeat messages HBn or requests 208 to the second component 205, continue to receive corresponding return heartbeat messages HBn or responses 210, continue to determine or calculate corresponding round trip times RTTn 212, and continue to compare the determined or calculated round trip times RTTn 212 with the stored mean RTT and standard deviation and / or, as the case may be, the periodicity-based threshold. For example, an RTTn 212 that is outside of a range of acceptable RTTs around the baseline RTT for the second component 205 (e.g., an unacceptable RTT) indicates that the second component 205 is exhibiting non-deterministic behavior and, therefore, may be experiencing performance degradation. In another example, for a second component 205 that is a control device 102a-102n, 105a-105m, an RTTn 212 that exceeds a periodicity-based threshold (e.g., an unacceptable RTT) for the second component 205 indicates that the second component 205 is exhibiting non-deterministic behavior and may be experiencing performance degradation. If the first component 202 observes a statistically significant number of unacceptable RTT measurements for the second component 205, the observation may indicate that the execution health of the second component is degraded or may have degraded. For example, if the second component is a control device 102a-102n, 105a-105m, the observation may indicate that the module scheduler of the control device 102a-102n, 105a-105m is unable to meet the control determinism of the periodicity of its corresponding control module.
[0037] 1, performance degradation of the physical control devices 102a-102n can be caused by the deterioration of hardware and / or software components. For example, deterioration of the hardware and / or software components of the physical control devices 102a-102n can be caused by improper automatic or manual loading of control logic onto the control devices 102a-102n, excessive interrupt loading, cache memory failures (leading to full memory accesses that are slower than cache memory accesses), denial of service attacks, extreme environmental conditions (e.g., excessive heat or cold) that can alter the execution quality of the control devices 102a-102n, and excessive radiation exposure that can flip or alter bits in instruction data, causing improper software operation, to name a few.
[0038] Degraded performance of the more computerized components of the control loop, such as the logic control devices 105a-105m and the I / O gateway 110, may be affected by a lack of hardware and / or software computing resources supporting the computing platform, and thus the resources may be shared resources. For example, the hardware and software resources of each host server 118a, 118b may be shared by a respective set of virtual control devices 105a-105g, 105h-105m executing thereon. Thus, each virtual control device 105a-105g, 105h-105m executing on the respective host server 118a, 118b must compete for host server resources with the other virtual control devices 105a-105g, 105h-105m executing on the host server 118a, 118b. Thus, a lack of hardware resources such as CPU resources, memory, and / or persistent storage space, as well as contention for software resources such as networking resources, logic resources, and / or other computing resources in each host server 118a, 118b, may degrade the performance of one or more virtual control devices 105a-105g, 105h-105m, respectively, running thereon.
[0039] Performance degradation of the I / O gateway 110 can be affected by the load on the I / O gateway 110 and / or contention for the network resources (e.g., software and / or hardware network resources) provided by the I / O gateway 110 for the delivery of I / O data to and from multiple control devices 102a-102n, 105a-105m. For example, I / O gateway load and / or resource contention for servicing the I / O gateway 110's clients (e.g., for scheduling, I / O data delivery logic, computing, etc.) can result in increased latency and / or jitter. Furthermore, I / O gateway load and / or resource contention can not only adversely affect the performance of various control loops, but can also adversely affect the overall performance of the process control system. For example, increased latency introduced by the I / O gateway 110 can degrade the overall operation of the process control system, and increased jitter introduced by the I / O gateway 110 can degrade the performance of the entire control strategy.
[0040] In either case, when the first component 202 determines that a particular RTTn 212 is outside a range of acceptable RTTs around the baseline RTT for the second component 205 and / or above a periodicity-based threshold of acceptable RTTs for the second component 205, the first component 202 may notify operations personnel that the second component 205 is experiencing non-deterministic behavior indicative of performance degradation that may lead to unpredictable results within the process control system. For example, the first component 202 may generate an alert or warning that may be displayed in one or more operator interfaces of the process control system.
[0041] Additionally or alternatively, the first component 202 may determine one or more mitigation actions in response to the detected degradation of the second component 205. For example, if the second component 205 is a control device 102a-102n, 105a-105m, the first component 202 (in this example, the I / O gateway 110) may determine one or more mitigation actions including adjusting the load of the control device 102a-102n, 105a-105m, slowing down the execution rate of control logic within the control device 102a-102n, 105a-105m, or determining and initiating changes to the control logic of the control device 102a-102n, 105a-105m. If the second component 205 is a logic control device 105a-105m, the one or more mitigation actions may include migrating the logic control device 105a-105m to another, less loaded host server, rebalancing the load of the host server among virtual control devices supported by the host server, rebalancing the load distribution among multiple host servers, etc. For example, actions to mitigate the detected degradation of the logic control device 105a-105m may include changing the load of one or more CPUs of one or more of the host servers, changing the memory usage of one or more of the host servers, and / or changing the disk space usage of one or more of the host servers.
[0042] In another example, if second component 205 is I / O gateway 110, first component 202 (in this example, physical control devices 102a-102n or logic control devices 105a-105m), or some other device in system 100, may determine one or more actions to mitigate the degradation of I / O gateway 110, including, for example, reducing the report rate of I / O gateway 110, reducing the number of clients served by I / O gateway 110, slowing down the I / O update rate for one or more clients served by I / O gateway 110, or otherwise modifying the load of I / O gateway 110. For example, actions to mitigate the detected degradation of I / O gateway 110 may include modifying the load of one or more CPUs of I / O gateway 110, modifying memory usage of I / O gateway 110, and / or modifying disk space usage of I / O gateway 110.
[0043] In one embodiment, the first component 202 may cause the determined mitigation actions to be presented to operations personnel in one or more operator interfaces as recommended and / or selectable options. When the operator selects one or more of the presented options, the process control system may execute the selected options. In another embodiment, upon determining one or more mitigation actions, the first component 202 may cause the process control system to automatically initiate and execute at least one of the determined mitigation actions, for example, without requiring any user input, and may optionally notify operations personnel about the automatically executed mitigation actions. Whether a particular mitigation action is executed automatically or manually may be pre-configured as needed. Thus, when performance degradation is detected, the process control system may be able to mitigate and correct (either manually or automatically) at least some of the issues causing the performance degradation of the component 205, rather than having to wait until an undesirable or disruptive event occurs. Thus, the techniques described within this disclosure may provide earlier detection, warning, and even automatic mitigation of degradation in process control loop components compared to currently known techniques.
[0044] In some implementations, the occurrence of a single RTT may not, by itself, trigger a warning, alert, or mitigation action. For example, upon receiving an unacceptable RTT, the first component 202 waits for a given time interval and / or a given number of additional heartbeat messages HBn to be sent or received to determine whether the unacceptable RTT is an anomaly or a trend, and only after the trend is confirmed (e.g., after receiving a statistically significant number of unacceptable RTTs), the first component 202 generates an alert, warning, and / or mitigation action. The given time interval, the given number of additional heartbeat messages HBn, and / or other information utilized to determine or confirm a statistically significant trend may be pre-configured and adjustable.
[0045] When the target or monitored component 205 is an I / O gateway 110, the group of control devices or groups of RTTs observed by each of the first components 202 (physical 102a-102n, logic 105a-105m, or both physical and logic) may be utilized to monitor and detect degradation of the I / O server 110. For example, one of the control devices in the group of first components 202 (or another node connected to the data highway 115) may collect records of abnormal RTTs observed among the group of first components 202. When one or more predetermined thresholds (e.g., corresponding to the rate of abnormal RTTs observed among the group of first components 202 over a given time interval, the percentage of the group of first components 202 experiencing abnormal RTTs, the variance in abnormal RTTs observed by the group of first components 202 over a particular time interval, the rate of occurrence of the variance, and / or other suitable thresholds) are exceeded, the process control system may generate an alert or warning indicating performance degradation of the I / O server 110 and may automatically determine, recommend, and / or initiate one or more mitigation actions. For example, a high variance rate in RTTs observed by the group of first components 202 may indicate increased jitter introduced by the I / O server 110, which may be caused by increased contention for the computing resources of the I / O server 110 and may adversely affect the performance of control strategies being executed by process control loops via the I / O server 110. A significant increase in the RTT duration between the first group of components 202 may indicate an increase in latency in the I / O server 110, which may be caused by an increase in the load on the I / O server 110 and may have a negative impact on the performance of the entire process control system.Accordingly, the process control system may determine, recommend, and / or initiate one or more suitable mitigating actions, such as reducing the reporting rate of the I / O gateway 110, reducing the number of clients served by the I / O gateway 110, slowing down the I / O update rate for one or more clients served by the I / O gateway 110, changing the load on one or more CPUs of the I / O gateway 110, changing the memory usage of the I / O gateway 110, changing the memory usage of the I / O gateway 110, and / or otherwise changing the allocation of resources of the I / O gateway 110.
[0046] In some embodiments, the process control system may aggregate RTTs observed by multiple controllers 102a-102n, 105a-105m, and / or the I / O gateway 110 to determine a score indicative of the overall health of the process control system or portions thereof. For example, the RTTs of the physical control devices 102a-102n determined by the I / O gateway 110 may be aggregated and used to determine corresponding latencies and jitters for the physical control devices 102a-102n as a group, thereby determining a score or indicator of the overall health of the set of physical control devices 102a-102n as a whole. The RTTs of the logic control devices 105a-105m determined by the I / O gateway 110 may be aggregated and used to determine corresponding latencies and jitters for the logic control devices 105a-105m as a group, thereby determining a score or indicator of the overall health of the logic control devices 105a-105m as a whole. Additionally, as discussed above, the RTTs of the I / O gateway 110 determined by multiple control devices 102a-102n, 105a-105m may be aggregated and utilized to determine the corresponding latency and jitter of the I / O gateway 110, which may in turn be utilized to determine a score or indicator of the overall health of the I / O gateway 110.
[0047] In particular, when the I / O gateway 110 is the target or monitored component 205, the average or overall RTT may indicate the communication delay time introduced by the I / O server 110 while it is forwarding messages from or vice versa to the control devices 102a-102n, 105a-105m, and final control elements 108a-108p. Thus, the average or overall RTT of the I / O server 110 (e.g., the control delay introduced by the I / O server 110 as a whole) may be determined from multiple RTTs measured by multiple control devices 102a-102n, 105a-105m. The minimum total number of control devices 102a-102n, 105a-105m measuring the respective RTTs of the I / O server 110 from which the average or overall RTT of the I / O server 110 is determined may be predefined and / or adjustable. However, for the most accurate estimation of the overall RTT of the I / O server 110, the respective RTTs measured by a majority or even all of the control devices 102a-102n, 105a-105m can be averaged to determine the overall RTT of the I / O server 110. For example, one of the control devices 102a-102n, 105a-105m or another device in the process control system 100 can determine the overall RTT of the I / O server 110 from the RTTs measured by multiple control devices 102a-102n, 105a-105m.
[0048] Furthermore, because the overall RTT of an I / O server 110 depends on the engineering of the control loops that utilize the I / O server 110 for I / O delivery, the average health of the I / O server 110 may be determined by comparing its average or overall RTT during operation to a baseline average or overall RTT obtained while the control system 100 was operating under quiescent or normal operating conditions. A threshold (e.g., a “difference threshold”) corresponding to the maximum allowable difference between the measured average RTT during operation and the baseline average RTT may be utilized to identify acceptable and unacceptable average or overall RTTs for the I / O server 110. Thus, a difference between the measured average or overall RTT of the I / O server 110 and the baseline average or overall RTT of the I / O server 110 that is greater than the difference threshold may indicate an unacceptable degradation in the performance of the I / O server 110. The difference threshold may be defined with respect to the execution period of the control module, for example, for an X% execution rate of the control module, and / or based on other criteria. The difference threshold may be predefined or preconfigured and may be adjustable.
[0049] In some configurations, the control system 100 may include multiple chained I / O servers 110 (not shown) that collectively act as a single logical I / O server to distribute messages between the control devices 102a-102n, 105a-105m and the final control elements 108a-108p. In these configurations, the overall RTT or control delay introduced by the chain of I / O servers 110 may be determined by aggregating or cumulatively adding the respective RTTs of each of the chained I / O servers 110. Differences between the measured RTTs of the chain of I / O servers 110 may be compared to the average or baseline RTT of the chain of I / O servers 110 to detect any degradation within the chain, e.g., in a manner similar to that of a single I / O server 110 discussed above.
[0050] Additionally, the respective RTT of each of the chained I / O servers 110 may be compared to the respective baseline RTT of each of the chained I / O servers 110 to identify or narrow the cause of control delay to a particular I / O server 110 in the chain. For example, if the difference between the operational RTT and the baseline RTT of a first I / O server 110 in the chain exceeds a respective difference threshold, while the difference between the operational RTT and the baseline RTT of a second I / O server 110 in the chain does not exceed a respective difference threshold, then the first I / O server 110 may be identified as a potential cause of control delay in the chain of I / O servers 110, and appropriate mitigation action may be taken against the first I / O server 110.
[0051] As discussed above, the measured RTT of an I / O server 110 (or a chain of I / O servers 110) may indicate the communication delay introduced by the I / O server 110 in the control loop. To illustrate, in one example, the monitoring device 202 may be a control device 102a that drives the action of the valve 108a by sending a message to the valve 108a every 500 milliseconds (ms) via the I / O server 110. The control device 102a (e.g., the monitoring device 202) may perform an RTT test on the I / O server 110 (e.g., the target or monitored device 205), and the measured RTT of the RTT test may be 100 ms. Thus, the communication delay introduced by the I / O server 110 in the control loop (e.g., the control loop including the control device 102a, the I / O server 110, and the valve 108a) may be 100 ms. Therefore, the overall time for control device 102a to receive input from I / O server 110, calculate a new valve position, and drive the new valve position via a corresponding signal to valve 108a may be delayed by an additional 100 ms due to the communication delay introduced by I / O server 110.
[0052] The amount of communication delay introduced by the I / O server 110 (e.g., as described above, e.g., overall or average measured RTT) may be stored in a parameter (e.g., an “I / O server communication delay parameter”) within the control system 100 and utilized to refine or improve the operation of a control loop that utilizes the I / O server to account for the communication delay. For example, in an exemplary control loop including the control device 102a, the I / O server 110, and the valve 108a, an indication of the value of the I / O server communication delay parameter may be included in an application time field included in a control signal (e.g., an output of the control device 102a) that is communicated to the valve 108a and controls the behavior of the valve 108a. Thus, in this example, the control or output signal sent to the valve 108a includes both an indication of the new / updated target valve position (e.g., as determined by the control device 102a) and an application time field that includes an indication of the value of the I / O server communication delay parameter. Thus, the content of the Apply Time field indicates the time at which the valve 108a will act on the indicated new / updated target valve position, e.g., the time at which the new / updated target valve position becomes effective for the valve 108a. Thus, the timing of the position change of the valve 108a takes into account communication delays introduced by the I / O server 110.
[0053] In an exemplary implementation, the valve 108a is a wireless valve 108a, and the control or output signal generated by the control device 102a to drive the valve 108a may be a WirelessHART command (or another type of wireless signal) that includes an indication of a new / updated target valve position and includes an application time field populated with an indication of the value of the I / O server communication delay parameter. Thus, upon receiving the command generated by the control device 102a, the valve 108a can delay acting on the new / updated target valve position according to the value of the application time field. Furthermore, the valve 108a can populate its READBACK parameter with the new / updated target valve position delayed by the value of the application time field. Thus, the READBACK parameter value reflects the target valve position regardless of the communication delay. As a result, if the READBACK parameter value begins to fluctuate more widely over time, the wider change may indicate that the valve 108a is performing differently and may need to be evaluated.
[0054] In some circumstances, a wireless gateway (through which WirelessHART or other types of wireless commands are sent to the wireless valve 108a) may utilize the value of the application time field to maintain a common sense of time among the devices of the wireless network, for example, by distributing or redistributing time slots among the devices of the wireless network based on the value of the application time field. Of particular note, because the I / O server communication delay parameter value is determined based on a statistically significant number of RTT measurements, changes in the value may indicate changes in the load on the I / O server 110 and / or changes in resource contention at the I / O server 110. Thus, because the delay time field value is determined based on the I / O server communication delay parameter value, providing the delay time field value to an ultimate control element (e.g., of the valve 108a) may enable the ultimate control element to respond to changing conditions at the I / O server 110. That is, the behavior of the ultimate control element (e.g., of the valve 108a) may automatically adjust or adapt to accommodate changes in load and / or resource usage at the I / O server 110. Additionally, and advantageously, since changes in the I / O server communication delay value indicate changes in the performance of the I / O server 110, the value of the I / O server communication delay value and its fluctuations can be monitored to easily detect degradation or performance problems of the I / O server 110.
[0055] Additionally, in some implementations of message flow 200 and / or message flow 220, the RTTs observed by the various components 102a-102n, 105a-105m, and 110 of the process control loop may be utilized to monitor and determine utilization of computing resources within the process control loop and / or the process control system. Computing resource utilization may be measured, for example, via one or more standard API calls to an underlying operating system of the target or monitored device 205, and computing resource utilization information obtained via the one or more standard API calls may be included in a return heartbeat message HBn or response 210 sent by the target or monitored device 205 to the monitoring device 202. If the target or monitored device 205 is an I / O server 110 or a logic control device 105a-105m, other computing resource utilization information of the I / O server 110 (e.g., network bandwidth, CPU availability or usage, memory availability or usage, etc.) may additionally or alternatively be sent by the target or monitored device 205 to the monitoring device 202 in the return heartbeat message HBn 210. The computing resource utilization may indicate the total capacity of the system computing resources currently being consumed. An increase in utilization may indicate a degradation in the overall performance of the system.
[0056] 3 shows a simplified block diagram of an example component 300 of a process control loop. For example, component 300 may be one of control devices 102a-102n, 105a-105m, or an I / O gateway, or server 110, or component 300 may be component 202 or component 205 of FIGS. 2A and 2B. For ease of illustration, and not by way of limitation, component 300 is described with simultaneous reference to FIGS. 1, 2A, and 2B.
[0057] 3, the component 300 includes or utilizes one or more processors 302, one or more memories 305, and one or more network interfaces 308 that communicatively connect the component 300 to a data highway or communication link of a process control system, such as the data highway 115. In embodiments in which the component 300 is a logic control device 105a-105m, the processor 302, memory 305, and network interface 308 utilized by the component 300 may be resources shared among multiple logic control devices. For example, the processor 302, memory 305, and network interface 308 of the logic control device component 300 may be provided by the host server 118a, 118b on which the logic control device component 300 and other logic control devices execute.
[0058] The network interface 308 enables the component 300 to communicate with the target component through two separate channels of the data highway. One of the channels 310 is a communication channel 310 through which the component 300 sends and receives process control and signaling messages to and from other loop components during control loop operation, thereby controlling at least a portion of an industrial process. The other channel 312 is a diagnostic channel through which the component 300 sends and receives heartbeat messages (e.g., heartbeat messages HBn 208, 210, 225) to and from target components (e.g., component 205) of the control loop to monitor and detect performance degradation of the target component. The communication channel 310 may be a dedicated channel or may be shared, for example, among multiple components and devices. In one embodiment, the diagnostic channel 312 may be a dedicated channel utilized exclusively by the component 300 and its corresponding target component to exclusively distribute heartbeat messages HBn 208, 210, 225 and, optionally, other types of diagnostic messages between them. For example, the component 300 may prevent communication and control messages utilized for operational process control from being sent or received over the diagnostic channel 312 .
[0059] The component 300 also includes or utilizes a process control message interpreter 315 and one or more process control loop modules 318. The process control message interpreter 315 and the process control loop modules 318, in embodiments, may include respective sets of computer-executable instructions stored on the memory 305 and executable by the one or more processors 302. In some embodiments, at least a portion of the process control message interpreter 315 may be implemented using the firmware and / or hardware of the component 300. Generally speaking, the process control message interpreter 315 and the process control loop modules 318 operate in conjunction to process incoming and outgoing process control messages (e.g., both control and signal messages) sent and received by the component 300 via the communication channel 312. While FIG. 3 depicts the process control message interpreter 315 and the process control loop modules 318 as separate modules or entities, in some embodiments of the component 300, the process control message interpreter 315 and the process control loop modules 318 may be implemented as an integrated module or entity.
[0060] In an exemplary configuration in which the component 300 is a control device 102a-102n, 105a-105m, the component 300 receives control messages via the communication channel 312 and the network interface 308, and the process control message interpreter 315 processes the control messages to obtain the message payload or content for the process control loop module 318. The process control loop module 318 includes one or more control routines or control logic for which the component 300 is specifically configured. The control routines or logic operate on the message content as input (possibly in combination with other inputs) to generate control signals that are packaged by the message interpreter 315 and transmitted from the component 300 to a receiving component or device via the network interface 308 and the communication channel 310. In embodiments in which the component 300 is a logic control device 105a-105m, the process control message interpreter 315 and the process control loop module 318 utilized by the component 300 may be resources shared among multiple logic control devices. For example, the process control message interpreter 315 and process control loop module 318 of the logic control device component 300 may be provided by the host server 118a, 118b on which the logic control device component 300 executes. The host server 118a, 118b may, for example, activate / deactivate more or fewer instances of the process control message interpreter 315 and / or process control loop module 318 to service its hosted logic control device as needed.
[0061] In another exemplary configuration in which the component 300 is the I / O gateway 110, the component 300 receives a control message or signal that is routed to a process control loop component or device, and the process control message interpreter 315 processes the message or signal and determines the recipient device of the message / signal. The recipient device may be, for example, a field device 108a-108p or a control device 102a-102b, 105a-10m. The process control loop module 318 includes switching or routing logic or routines that optionally convert or transform the process control message / signal into a format suitable for transmission to the recipient device and transmits the message / signal to the recipient device. For example, if the component 300 receives a control message from one of the physical or logic controllers 102a-102n, 105a-105m via a communication channel 310 of the data highway 115 and the control message is intended for distribution to another one of the physical or logic controllers 102a-102n, 105a-105m, the process control message interpreter 315 and / or the process control loop module 318 may simply forward the control message to that receiving controller 102a-102n, 105a-105m via the communication channel 310, for example. In another example, in which the component 300 receives a control message from one of the physical or logic controllers 102a-102n, 105a-105m via the communication channel 310 and the control message is intended to be distributed to the field devices 108a-108p, the process control loop module 318 may convert the message into a signal distributable to the receiving field device 108a-108p via the respective link 112a-112p and route the signal to the receiving field device 108a-108p via the respective link 112a-112p.In yet another example, where the component 300 receives a signal from one of the field devices 108a-108p via a respective link 112a-112p and the signal content is distributed to the control devices 102a-102n, 105a-105m, the process control loop module 318 may convert the signal content into a control message and transmit the control message to the receiving control device 102a-102n, 105a-105m via the communication channel 310. Note that because the I / O gateway 110 is typically implemented on a computing platform, the I / O gateway 110 may support multiple instances of the process control message interpreter 315 and / or the process control loop module 318. For example, the I / O gateway 110 may activate / deactivate more or fewer instances of the process control message interpreter 315 and / or the process control loop module 318 as needed.
[0062] In some embodiments, component 300 is an impairment detection component that monitors a target component of a process control loop for degradation in response performance and detects degradation in the target component's response performance. The impairment detection component 300 and the target component are included within the same process control loop and, therefore, are both components of the process control loop. For example, degradation detection component 300 may be component 202 of FIGS. 2A and 2B. In such embodiments, component 300 includes a degradation detector 320 and a degradation store 322 stored on one or more memories 305. The degradation detector 320 may include computer-executable instructions executable by one or more processors 302 to perform corresponding actions described above for message flows 200, 220 and component 202 of FIGS. 2A and 2B. Additionally or alternatively, degradation detector 320 may be executable by one or more processors 302 to perform at least a portion of method 400 for detecting degradation of a loop component, which will be discussed in more detail below. Generally speaking, degradation detector 320 is configured to send and receive heartbeat messages to and from a target component (e.g., component 205) via diagnostic channel 312 to monitor, detect, and diagnose reasons for a degradation in the target component's response performance outside of its normal, acceptable operating range, e.g., degradation of the target component. In some embodiments, degradation detector 320 is configured to notify operations personnel of a detected degradation of the target component and to determine one or more mitigation actions and / or initiate at least one of the mitigation actions as described elsewhere in this disclosure.
[0063] Further, in embodiments in which component 300 is a degradation detection component, degradation detector 320 may be configured to determine a normal, typical, or acceptable operating range for the target component, e.g., by determining the average round-trip time (RTT) for a predetermined number of heartbeat messages sent to and received from the target component and the corresponding standard deviation, such as in the manner described above with respect to FIGS. 2A and 2B. Degradation detector 320 may, for example, store the determined average RTT and corresponding standard deviation in degradation detection data store 322 and utilize the stored data to monitor and detect degradation of the target component. Calculation or determination of the average RTT and corresponding standard deviation may be performed automatically by component 300 (e.g., periodically, upon occurrence of a particular event, such as a reconfiguration of the target component, a software upgrade, etc.) and / or manually or upon user command. For example, an operator may instruct component 300 to determine the average RTT and standard deviation associated with the target component at different points in the target component's lifecycle, such as upon completion of a software upgrade on the target component, upon completion of a reconfiguration of the target component, upon completion of maintenance on the component, under various system loads and system configurations, etc. The degradation detector 320 may determine and store (eg, in the degradation data store 322) multiple average RTTs and standard deviations for different loads, configurations, and scenarios, as desired.
[0064] Of course, component 300 may further include other instructions 325 and other data 328 for utilization in its process control and / or impairment detection operations, and / or other operations.
[0065] Additionally, it should be noted that in some cases, component 300 may detect component degradation for multiple target components. For example, I / O gateway 110 may be configured to monitor and detect degradation for multiple controllers 102a-102n, 105a-105m.
[0066] Further, it should be noted that not all components of a process control loop need be configured to perform impairment detection. For example, in Figure 1, components 102a, 105g, 105h, and 110 are shown by the circled DDs as being configured for impairment detection, and therefore are each configured to include a respective instance of the impairment detector 320 and the impairment detection store 322. On the other hand, components 102n, 105a, and 105h are shown as not configured for impairment detection, and therefore each of components 102n, 105a, and 105h may omit or deactivate a respective instance of the impairment detector 320 and the impairment detection store 322.
[0067] In some embodiments, component 300 may additionally or alternatively be a target component of a process control loop being monitored for performance degradation by another impairment-detection component of the process control loop. For example, component 300 may be component 205 of FIGS. 2A and 2B. In these embodiments, component 300 may additionally be an impairment-detection component, i.e., component 300 may or may not include or utilize an impairment detector 320 and impairment-detection data 322. In either case, in embodiments in which component 300 is a target component, component 300 receives a heartbeat message HBn 208 from the impairment-detection component via a diagnostic channel 312. Upon receiving the heartbeat message HBn 208, component 300 processes the received heartbeat message HBn 208 via a process control message interpreter 315 and forwards or otherwise transmits the heartbeat message HBn 210, 225 to a sending component via the diagnostic channel 312. In particular, the component 300 sends heartbeat messages HBn 210, 225 back to the sending component via the process control message interpreter 315 at the fastest rate at which the component 300 is configured to report or transmit process control values. For example, if the component 300 is an I / O gateway 110 and the I / O gateway is configured to report or transmit process control values at a maximum rate of 50 ms, the component 300 sends heartbeat messages HBn 210, 225 back to the sending component at a rate of 50 ms.
[0068] 4 illustrates a block diagram of an example method 400 for detecting an impairment in a component of a process control loop included within a distributed process control system (DCS) of a physical industrial process plant, such as portion 100 of the process control system illustrated in FIG. 1. In embodiments, different instances of at least a portion of method 400 may be performed by one or more control devices 102a-102n, 105a-105m, respectively, and / or by an I / O gateway or server 110. Additionally or alternatively, at least a portion of method 400 may be performed by component 202 of FIGS. 2A and 2B or component 300 of FIG. 3. For example, at least a portion of method 400 may be performed by impairment detector 320 of component 300. In embodiments, method 400 may include additional or alternative blocks other than those discussed within this disclosure.
[0069] At block 402, a method 400 for detecting degradation of a component in a process control loop includes sequentially transmitting a plurality of heartbeat messages at a first component of the process control loop to a second component of the process control loop via a diagnostic channel. The first and second components of the process control loop are communicatively coupled via both a diagnostic channel and a communication channel, such as diagnostic channel 312 and communication channel 310 of FIG. 3. For example, the first component may be component 202 of FIGS. 2A and 2B, and thus may be I / O gateway 110 or one of process controllers 102a-102n, 105a-105m. Thus, if the first component is an I / O gateway or server 110, the second component may be one of the process controllers 102a-102n, 105a-105m, and if the first component is one of the process controllers 102a-102n, 105a-105m, the second component may be an I / O gateway or server 110. For example, the second component may be component 205 of FIGS. 2A and 2B.
[0070] As the second component receives each heartbeat message sent by the first component, the second component forwards or otherwise returns the heartbeat message to the first component via the diagnostic channel. Thus, at block 405, method 400 includes receiving at least a subset of the plurality of heartbeat messages at the first component from the second component via the diagnostic channel, where at least a subset of the plurality of heartbeat messages have been returned by the second component to the first component upon respective receipt at the second component.
[0071] At block 408, method 400 includes determining, by the first component, an average response time or round trip time (RTT) of the second component based on the RTTs of at least a subset of the plurality of heartbeat messages. The respective RTTs may be determined based on the respective transmission and reception times (e.g., respective TS1 and TS2 discussed above with respect to FIGS. 2A and 2B) of at least a subset of the plurality of heartbeat messages received by the first component. A minimum number of messages included in at least a subset of the received plurality of heartbeat messages and used to determine the average response time of the second component may be preconfigured and optionally adjustable. Further, at block 408, method 400 may include determining a standard deviation corresponding to the average RTT of the second component based on the RTTs of at least a subset of the plurality of heartbeat messages.
[0072] In embodiments, method 400 may include storing, in the first component, for example, the average response time or average RTT and standard deviation of the second component. Additionally or alternatively, method 400 may include determining and storing an acceptable range of RTTs for the second component. For example, a lower limit of the acceptable range of RTTs for the second component may be the average RTT minus the standard deviation, and an upper limit of the acceptable range of RTTs may be the average RTT plus the standard deviation. In some embodiments, method 400 may additionally or alternatively include storing, in the first component, a threshold value corresponding to the periodicity of module scheduler execution (e.g., a periodicity-based threshold). The periodicity-based threshold may be determined, for example, based on the configuration of the second component, and a measured RTT that exceeds the periodicity-based threshold may be an unacceptable RTT for the second component.
[0073] At some point after the average RTT and corresponding standard deviation and / or periodicity-based threshold have been determined in block 410 (block 408), method 400 may include determining the RTT of another heartbeat message subsequently sent by the first component, replied to by the second component, and received by the first component, e.g., in a manner such as discussed with respect to Figures 2A and 2B. In block 412, method 400 determines whether the RTT of the subsequent heartbeat message is between the upper and lower limits of a range of acceptable RTTs for the second component, e.g., within plus or minus a standard deviation of the average RTT for the second component, and / or whether the RTT of the subsequent heartbeat message exceeds the periodicity-based threshold for the second component. If the RTT of the subsequent heartbeat message is determined to be an acceptable RTT by either or both of the acceptance criteria (e.g., as indicated by the NO leg of block 412), method 400 continues by sending the next heartbeat message (block 415) and determining its respective RTT (block 410).
[0074] On the other hand, if method 400 determines that the RTT of the subsequent heartbeat message is an unacceptable RTT for the second component (e.g., not within plus or minus a standard deviation of the average RTT and / or exceeds a periodicity-based threshold as indicated by the YES leg of block 412), method 400 includes detecting a degradation of the second loop component based on the out-of-range RTT (block 418), such as in the manner described above, and correspondingly alerting a user of the degradation, determining one or more actions to mitigate the degradation, and / or initiating at least one of the mitigation actions (block 420).
[0075] Thus, by utilizing the techniques for detecting loop component degradation described herein, degradation within control devices and I / O gateways can be determined from changes in their respective performance responsiveness and / or trends in such changes. Thus, the process control system can notify operations personnel of component degradation before the degradation results in a failure or catastrophic event. Indeed, because known diagnostic processes are typically scheduled to occur, operate, and respond at a slower rate and / or lower priority than real-time process control and communication messages, the process control system can notify operations personnel of component degradation earlier than it can be detected by known diagnostic processes. Furthermore, in some embodiments, the process control system can recommend or suggest one or more mitigation actions to operations personnel to address the detected degradation, and in some cases, automatically initiate one or more mitigation actions to address the detected degradation. Thus, the techniques described herein advantageously provide an early degradation detection system with an optional corresponding automatic degradation mitigation system.
[0076] Further advantageously, the techniques described herein can be utilized to standardize load balancing across various control loop components of a computerized process control system. For example, the load of control logic within a control device, the report rate of an I / O gateway, the number of clients served by an I / O gateway, the load of the physical computing platform supporting the logic control device, and the load of the physical computing platform supporting the I / O gateway can be adjusted (e.g., automatically adjusted) based on a comparison of the measured RTT to the average RTT. Further advantageously, the deterministic / non-deterministic measurements of loop component RTT can be utilized as a performance or health metric of various loop components, process control loops, and even the process control system itself, thus advantageously providing a mechanism for monitoring and evaluating the overall performance, health, and utilization of various loop components, process control loops, and even the process control system itself as a whole.
[0077] If implemented in software, any of the applications, modules, etc. described herein may be stored in any tangible, non-transitory computer-readable memory, such as a magnetic disk, laser disk, solid-state storage device, molecular memory storage device, or other storage medium, RAM, or ROM of a computer or processor. It should be noted that while the exemplary systems disclosed herein are disclosed as including software and / or firmware running on hardware, among other components, such systems are merely exemplary and should not be considered limiting. For example, it is contemplated that any or all of these hardware, software, and firmware components may be embodied exclusively in hardware, exclusively in software, or in any combination of hardware and software. Thus, while the exemplary systems described herein are described as implemented in software running on processors of one or more computing devices, those skilled in the art will readily recognize that the provided examples are not the only way to implement such systems.
[0078] Thus, while the present invention has been described with reference to specific examples, it will be apparent to those skilled in the art that these are illustrative only and are not intended to be limitations of the invention, and that modifications, additions, or deletions may be made to the disclosed embodiments without departing from the spirit and scope of the invention.
[0079] The particular features, structures, and / or characteristics of any particular embodiment may be combined with one and / or more other embodiments in any suitable manner and / or in any suitable combination, including using selected features with or without the corresponding use of other features. Furthermore, many modifications may be made to adapt a particular application, situation, and / or material to the essential scope or spirit of the invention. It should be understood that other variations and / or modifications of the embodiments of the invention described and / or illustrated herein are possible in light of the teachings herein and should be considered part of the spirit or scope of the invention. Certain aspects of the invention are described herein as exemplary aspects.
Claims
1. 1. A system for detecting component degradation in an industrial process plant, the system comprising: a first component of a process control system of the industrial process plant, the first component communicatively connected to a second component of the process control system via a diagnostic channel and via a communication channel, the first component being an I / O gateway or one of a process controller included in a plurality of process controllers communicatively connected to the I / O gateway via respective communication channels, the I / O gateway communicatively connecting the plurality of process controllers to respective one or more field devices thereby controlling an industrial process within the process plant, the second component being the other of the I / O gateway or the process controller, the first component comprising: sending a plurality of heartbeat messages sequentially to the second component via the diagnostic channel; receiving, via the diagnostic channel, at least a subset of the plurality of heartbeat messages sent by the second component back to the first component upon respective receipt at the second component; determining an average response time of the second component based on at least one of a periodicity of module scheduler execution at the second component or a round trip time (RTT) of each of the at least the subset of the plurality of heartbeat messages, the each RTT being determined based on respective times of transmission and reception of the at least the subset of the plurality of heartbeat messages at the first component; and detecting a degradation of the second component when a RTT of a subsequent heartbeat message sent by the first component to the second component via the diagnostic channel exceeds a threshold corresponding to the average response time of the second component.
2. 2. The system of claim 1, wherein a control loop of the process control system includes the first component and the second component, and communication and control messages delivered between the first component and the second component via the control loop to control the industrial process are delivered between the first and second components via the communication channel and are not delivered between the first and second components via the diagnostic channel.
3. and a routine executed in the second component, the routine comprising: processing communication and control messages received at the second component via the communication channel for controlling the industrial process; and transmitting any heartbeat messages received at the second component via the diagnostic channel back to the first component via the diagnostic channel.
4. The system of claim 1 , wherein the diagnostic channel is a dedicated diagnostic channel established for exclusive use by the first component and the second component.
5. 5. The system of claim 1, wherein the degradation of the second component is detected when the RTT of the subsequent heartbeat message is outside a given number of standard deviations from the average response time.
6. The system of claim 5 , wherein the given number of standard deviations is one standard deviation.
7. 7. The system of claim 1, wherein the second component is configured to return heartbeat messages received over the diagnostic channel at a fastest update rate supported by the second component.
8. the first component, upon said detection of said impairment, generating an alert or alarm indicative of said detected degradation; automatically rebalancing a load corresponding to the second component; determining mitigation actions; causing the mitigation action to be performed automatically within the process control system; or 8. The system of claim 1, further configured to perform at least one of the following: causing a user interface to present an alert or warning indicating the mitigation action.
9. 9. The system of claim 1, wherein the process control system is further configured to determine one or more metrics indicative of an overall health of the process control system based on an average response time of each of at least one of the plurality of process controllers or the I / O gateway.
10. The system of claim 9 , wherein the process control system is further configured to detect a degradation in the overall health of the process control system based on a change in the one or more metrics.
11. 11. The system of claim 10, wherein the one or more metrics indicate a level of latency throughout the process control system, and the detected degradation in the health of the overall process control system comprises a reduction in availability of computing resources.
12. 12. The system of claim 10 or claim 11, wherein the one or more metrics indicate a level of jitter across the process control system, and the detected degradation in the health of the overall process control system comprises increased contention for computing resources.
13. 13. The system of claim 1, wherein a total number of heartbeat messages included in at least the subset of the plurality of heartbeat messages is greater than or equal to a minimum number of heartbeat messages required to determine the average response time of the second component, and wherein the minimum number of heartbeat messages is configurable.
14. The system of claim 1 , wherein the first component is the process controller and the second component is the I / O gateway.
15. the process control system is further configured to determine a metric indicative of the health of the I / O gateway based on at least one of a respective average response time, a respective average latency, or a respective average jitter corresponding to the I / O gateway; 15. The system of claim 14, wherein the at least one of the respective average response times, respective average latencies, or respective average jitters corresponding to the I / O gateways is determined based on at least one of the respective response times, respective latencies, or respective jitters corresponding to the I / O gateways and is determined by the plurality of process controllers.
16. 16. The system of claim 14 or claim 15, wherein the process controller is a physical process controller, and the physical process controller is configured to determine the RTT of each of the at least the subset of the plurality of heartbeat messages and the RTT of the subsequent heartbeat message using a hardware clock included in the physical process controller.
17. the process controller is a virtual process controller; each heartbeat message in at least the subset of the plurality of heartbeat messages returned by the I / O gateway includes an indication of a respective time that the each heartbeat message was received at the I / O gateway and a respective time that the I / O gateway sent the respective return of the each heartbeat message to the virtual process controller; 17. The system of claim 14, wherein the virtual process controller is configured to determine the respective RTTs of the at least the subset of the plurality of heartbeat messages further based on the respective times indicated in each heartbeat message.
18. The system of claim 17 , wherein the virtual process controller is implemented via a container.
19. the average response time of the second component indicates a communication delay introduced by the I / O gateway; 19. The system of claim 14, wherein the control signal generated by the process controller to drive a field device includes a target value for the field device and an application time field indicating the communication delay introduced by the I / O gateway.
20. 20. The system of claim 19, wherein the field device stores an indication of the target value for the field device delayed by the communication delay introduced by the I / O gateway.
21. The system of claim 1 , wherein the first component is the I / O gateway and the second component is the process controller.
22. 22. The system of claim 21, wherein the I / O gateway is configured to determine the respective RTTs of the at least the subset of the plurality of heartbeat messages and the RTTs of the subsequent heartbeat messages by using a hardware clock included in the I / O gateway.
23. 23. The system of any of claims 1 to 22, wherein the process controller is a safety system controller.
24. 1. A method for detecting component degradation in a process control system, the method comprising: a first component communicatively coupled to a second component via a communication channel and a diagnostic channel, the first component is one of an I / O gateway or a process controller included in a plurality of process controllers communicatively connected to the I / O gateway via respective communication channels, the I / O gateway communicatively connecting the plurality of process controllers to respective one or more field devices to thereby control an industrial process within a process plant; In the first component, the second component is the other of the I / O gateway or the process controller, transmitting, by the first component, a plurality of heartbeat messages sequentially over the diagnostic channel to the second component; receiving at least a subset of the plurality of heartbeat messages at the first component from the second component via the diagnostic channel, wherein each heartbeat message of the at least the subset of the plurality of heartbeat messages is transmitted by the second component back to the first component upon respective receipt at the second component; determining, by the first component, an average response time of the second component based on at least one of a periodicity of module scheduler execution at the second component or a round trip time (RTT) of each of the at least the subset of the plurality of heartbeat messages, the each RTT being determined based on respective times of transmission and reception of the at least the subset of the plurality of heartbeat messages at the first component; detecting degradation of the second component when a RTT of a subsequent heartbeat message sent by the first component to the second component exceeds a threshold corresponding to the average response time of the second component.
25. Upon detecting the degradation of the second component, generating an alert or alarm indicative of the detected degradation; automatically rebalancing the load on the second component; determining a mitigation action for the detected impairment; causing said mitigation action to be performed automatically; or 25. The method of claim 24, further comprising at least one of: generating an alert or warning indicating the mitigation action.
26. 26. The method of claim 25, wherein the second component is the process controller, and determining the mitigation action comprises determining a change to a control routine executed by the process controller.
27. 27. The method of claim 25 or claim 26, wherein the second component is the process controller, the process controller being a virtual process controller executing on a physical computing platform, and causing the mitigation action to be performed automatically includes migrating the virtual process controller to another physical computing platform for performance.
28. 28. The method of claim 27, wherein the virtual process controller is implemented via a container, and migrating the virtual process controller to the other physical computing platform comprises assigning the container to the other physical computing platform.
29. 29. The method of claim 25, wherein the second component is the I / O gateway, and wherein determining the mitigation action includes at least one of determining a change to a reporting rate of the I / O gateway or determining a change to a number of clients serviced by the I / O gateway.
30. 30. The method of claim 25, wherein determining the mitigation action comprises determining a change to a load on the physical computing platform that supports the second component.
31. 31. The method of claim 30, wherein determining the change to the load of the physical computing platform comprises determining a respective change to one or more of a CPU (Central Processing Unit) load of the physical computing platform, a memory usage of the physical computing platform, or a disk space usage of the physical computing platform.
32. 32. The method of claim 24, wherein detecting the degradation of the second component comprises detecting the degradation of the second component when the RTT of the subsequent heartbeat message is outside one or more predetermined standard deviations from the average response time of the second component.
33. 33. The method of claim 24, further comprising determining one or more metrics indicative of the overall health of the process control system based on an average response time of each of at least one of the plurality of process controllers or the I / O gateway.
34. 34. The method of claim 33, further comprising detecting a degradation in the overall health of the process control system based on a change in the one or more metrics.
35. the one or more metrics are indicative of a level of latency throughout the process control system; 35. The method of claim 34, wherein detecting the degradation in the overall health of the process control system comprises detecting a reduction in availability of computing resources based on the change in the level of latency.
36. the one or more metrics indicate a level of jitter throughout the process control system; 36. The method of claim 34 or claim 35, wherein detecting the degradation in the overall health of the process control system comprises detecting increased contention for computing resources based on the change in the level of jitter.
37. 37. The method of claim 24, further comprising determining, by the first component, the respective round trip time (RTT) of each heartbeat message of the at least the subset of the plurality of heartbeat messages and the RTT of the subsequent heartbeat message.
38. the first component is the process controller, the process controller is a virtual process controller, and the second component is the I / O gateway; the method further comprising receiving, for each heartbeat message in the at least the subset of the plurality of heartbeat messages, an indication of a respective reception time of the each heartbeat message at the I / O gateway and a respective transmission time of the return transmission of the each heartbeat message by the I / O gateway; 38. The method of claim 37, wherein determining the average response time of the second component is further based on the received indication of the respective receive times at the I / O gateway and the respective transmit times of the returns by the I / O gateway.
39. 39. The method of claim 37 or claim 38, wherein determining the respective RTT of each heartbeat message of the at least the subset of the plurality of heartbeat messages and the RTT of the subsequent heartbeat message is based on a hardware clock included in the first component.
40. configuring a minimum number of heartbeat messages based on which the average response time of the second component is calculated; 40. The method of claim 24, wherein determining the average response time of the second component comprises determining, at the first component, the average response time of the second component when at least the minimum number of heartbeat messages have been received back from the second component.
41. the first component is the process controller, and the second component is the I / O gateway; the method further includes determining a metric indicative of the health of the I / O gateway based on at least one of a respective average response time, a respective average latency, or a respective average jitter corresponding to the I / O gateway; 41. The method of claim 24, wherein the at least one of a respective response time, a respective latency, or a respective jitter corresponding to the I / O gateway is determined by the plurality of process controllers based on the respective heartbeat messages.
42. the first component is the process controller, the second component is the I / O gateway, and the method comprises: determining a value of a delay time parameter based on the average response time of the I / O gateway; generating, by the process controller as an output of execution of a control module, a message for driving a field device, the message including an indication of a target value for the field device and the delay time parameter populated with the value determined based on the average response time of the I / O gateway; 42. The method of claim 24, further comprising transmitting, by the process controller, the generated message to the field device via the I / O gateway.
43. 43. The method of any of claims 24 to 42, further comprising establishing the diagnostic channel between the first component and the second component for dedicated use, including preventing communication and control messages utilized in controlling the industrial process from being distributed between the first component and the second component over the diagnostic channel.
44. the first component and the second component are included in a control loop that executes to control at least a portion of the industrial process; 44. The method of any of claims 24 to 43, wherein the method further comprises processing, by the first component, communication and control messages sent by and / or received at the first component via the communication channel, thereby executing the control loop to control at least a portion of the industrial process.
45. a routine executing in the second component for processing communication and control messages sent by and / or received at the second component via the communication channel, thereby executing the control loop to control at least the portion of the industrial process; 45. The method of claim 44, wherein receiving, at the first component, the at least the subset of the plurality of heartbeat messages returned by the second component comprises receiving, at a fastest update rate supported by the second component, the at least the subset of the plurality of heartbeat messages returned by the routine executed in the second component.
Citation Information
Patent Citations
Deterioration diagnostic method in equity type communication line system
JP1992207339A