Method, system, and computer program for alarm processing
The method and system address alarm overload in industrial systems by identifying and suppressing nuisance alarms and predicting future alarms, improving operator efficiency and safety through automated alarm rationalization and proactive notification.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2021-10-26
- Publication Date
- 2026-03-10
AI Technical Summary
Industrial process control systems face issues with alarm overload due to nuisance alarms, such as chattering, redundant, and consequential alarms, which overwhelm operators and can lead to accidents and increased stress, while true alarms are often buried and overlooked.
A method and system for alarm processing that identifies and suppresses nuisance alarms, groups relevant alarms, and predicts future alarms by analyzing historical data for patterns and correlations, using median time differences and probability thresholds to generate ordered sequences of operator actions.
Reduces alarm burden by accurately identifying and suppressing nuisance alarms, automates alarm rationalization, and provides proactive notification of potential future alarms, enhancing operator efficiency and safety.
Smart Images

Figure 0007826643000001 
Figure 0007826643000002 
Figure 0007826643000003
Abstract
Description
[Technical Field]
[0001] At least one exemplary embodiment relates to the field of industrial process control systems, and more particularly to methods, systems, and computer programs for alarm processing, alarm prediction, and / or alarm rationalization within industrial process control systems. [Background technology]
[0002] Industrial environments, such as manufacturing, production, mining, and construction environments, contain complex systems and devices and equally complex workflows. Processing facilities within industrial environments, such as refineries and water purification plants, are routinely managed using process control systems. Process control systems may be configured to manage the function and operation of industrial equipment, including machines, sensors, valve devices, and / or actuators within the processing facilities. [Prior art documents] [Non-patent literature]
[0003] [Non-Patent Document 1] Industrial standard ANSI / ISA-18.2 [1, 16 pages] [Non-patent document 2] Industrial standard ANSI / ISA-18.2 [1, 18 pages] Summary of the Invention [Means for solving the problem]
[0004] At least one exemplary embodiment relates to the field of industrial process control systems, and more particularly to methods, systems, and computer programs for alarm processing, alarm prediction, and alarm rationalization within industrial process control systems.
[0005] At least one exemplary embodiment relates to the field of industrial process control systems, and more particularly to methods, systems, and computer programs for alarm processing, alarm prediction, and alarm rationalization within industrial process control systems.
[0006] At least one example embodiment provides a method for alarm processing in a process control system, the method including: (i) detecting one or more alarm events based on status data received from at least one device in the process control system; and (ii) in response to determining that the detected one or more alarm events match a stored alarm event pattern, (a) obtaining an alarm event pattern response associated with the matched alarm event pattern, the alarm event pattern response identifying one or more alarm response events; and (b) generating a control signal to implement one or more of the alarm response events.
[0007] The stored alarm event pattern may be generated based on the following steps: (i) acquiring a set of historical data including alarm and event log data; (ii) correlating, based on the alarm and event log data in the acquired set of historical data, a reference alarm event with at least one of: (a) one or more concurrent candidate alarm events, where a respective probability of concurrent occurrence of each of the one or more candidate alarm events with the reference alarm event is determined to be equal to or greater than a defined first threshold; and (b) one or more operator actions, where a respective probability of concurrent occurrence of each of the one or more operator actions with the reference alarm event is determined to be equal to or greater than a defined second threshold; and (iii) including, in the stored alarm event pattern, an ordered sequence including the one or more concurrent candidate alarm events or the one or more concurrent operator actions, where the ordered sequence is generated based on a median time difference between a timestamp associated with the reference alarm event and a timestamp associated with the concurrent candidate alarm events or the concurrent operator actions.
[0008] In an embodiment of the method, an ordered sequence including one or more concurrent candidate alarm events or one or more concurrent operator actions is additionally generated based on a median absolute deviation from a median of the determined time differences for the concurrent candidate alarm events or concurrent operator actions.
[0009] In another method embodiment, the step of correlating the reference alarm event with at least one of one or more concurrent candidate alarm events and one or more operator actions is performed based on alarm and event log data from a reduced set of historical data, the reduced set of historical data being generated based on (i) identifying chattering alarm data within the retrieved set of historical data, and (ii) generating the reduced set of historical data to include alarm and event data from the retrieved set of historical data other than the identified chattering alarm data.
[0010] In certain embodiments of the method, the chattering alarm data includes (i) one or more alarm events having an alarm gap less than or equal to a predefined first duration, or (ii) one or more alarm events having an alarm lifetime less than or equal to a predefined second duration.
[0011] In another embodiment of the method, (i) the matched alarm event pattern includes one or more alarm events that do not have a corresponding concurrent operator action, and (ii) the obtained alarm event pattern response associated with the matched alarm event pattern includes an alarm suppression process flow for the one or more alarm events that do not have a corresponding concurrent operator action.
[0012] In a further method embodiment, determining that the alarm event does not have a corresponding concurrent operator action includes (i) determining one or more probabilities of coincidence of at least one detected operator action with the alarm event; and (ii) identifying the alarm event as an alarm event that does not have a corresponding concurrent operator action in response to the determined one or more probabilities of coincidence being less than a predefined value.
[0013] In an embodiment of the method, (i) the matched alarm event pattern includes one or more alarm events having a corresponding set of concurrent operator actions, and (ii) the obtained alarm event pattern response associated with the matched alarm event pattern includes initiating a control signal for presenting a standardized set of operator actions to an operator in response to detecting one or more alarm events or in response to detecting the matched alarm event pattern, wherein the standardized set of operator actions includes a corresponding set of concurrent operator actions.
[0014] According to another embodiment of the method, determining that the alarm event has a corresponding set of concurrent operator actions includes: (i) determining one or more probabilities of co-occurrence of the one or more operator actions with the alarm event; (ii) identifying the one or more operator actions as concurrent operator actions that occur simultaneously with the alarm event in response to the determined one or more probabilities of co-occurrence being greater than a predefined value; (iii) identifying a timestamp associated with each of the one or more concurrent operator actions; and (iv) ordering each of the plurality of concurrent operator actions in a sequence, wherein a position of the concurrent operator action in the sequence is determined based on a median time difference between a timestamp associated with the concurrent operator action and a timestamp associated with the alarm event.
[0015] In one method embodiment, the position of the concurrent operator actions within the sequence is additionally determined based on a median absolute deviation from a median of the determined time differences for the concurrent operator actions.
[0016] The method may include an embodiment in which (i) the matched alarm event pattern includes a cluster of redundant alarm events, and (ii) the obtained alarm event pattern response associated with the matched alarm event pattern initiates an alarm suppression process flow for one or more alarm events in the cluster of redundant alarm events.
[0017] The method may additionally include an embodiment in which the matched alarm event pattern is generated by the steps of: (i) determining one or more probabilities of co-occurrence of at least one candidate alarm event with a reference alarm event; (ii) identifying the candidate alarm event as a co-occurring alarm event that occurs co-occurring with the reference alarm event in response to the determined one or more probabilities of co-occurrence being greater than a predefined value; (iii) generating a cluster of alarm events including the reference alarm event and the identified one or more co-occurring alarm events; (iv) identifying a timestamp associated with each of a plurality of alarm events in the generated cluster of alarm events; and (v) ordering each of the plurality of alarm events in the generated cluster of alarm events in a sequence, wherein the position of the alarm event sought to be ordered in the sequence is determined based on a median time difference between a timestamp associated with the reference alarm event and a timestamp associated with the candidate alarm event in the cluster of alarm events.
[0018] In an embodiment of the method, the position of the candidate alarm event within the sequence is additionally determined based on a median absolute deviation from a median of the time differences determined for the candidate alarm event.
[0019] The method may include an embodiment in which the step of identifying redundant alarm events for grouping in a cluster of redundant alarm events includes: (i) identifying a cluster of alarm events that occur sequentially; (ii) determining a time of occurrence of each alarm event in the identified cluster; and (iii) responding to a determination that the time of occurrence of each alarm event in the cluster is separated from a previous alarm event in the cluster by less than a defined time value by identifying the cluster of alarm events as a cluster of redundant alarm events.
[0020] In another embodiment of the method, (i) the matched alarm event pattern includes a cluster of consequential alarm events, and (ii) the obtained alarm event pattern response associated with the matched alarm event pattern includes initiating an alarm prediction process flow that includes responding to the detection of an occurrence instance of one or more preceding alarm events in the cluster of consequential alarm events by presenting to an operator information that predicts a future occurrence of one or more of the instances of subsequent alarm events in the cluster of consequential alarm events before the subsequent alarm event instance is detected.
[0021] In a further embodiment of the method, the cluster of resulting alarm events is generated based on the steps of (i) identifying a timestamp associated with each of the plurality of resulting alarm events, and (ii) ordering each of the plurality of resulting alarm events in a sequence, wherein the position of a candidate alarm event to be ordered in the sequence is determined based on the median time difference between the timestamp associated with a reference alarm event and the timestamp associated with the candidate alarm event within the plurality of resulting alarm events.
[0022] In certain embodiments of the method, the position of the candidate alarm event within the sequence is additionally determined based on a median absolute deviation from the median of the time differences determined for the candidate alarm event.
[0023] In a further embodiment of the method, identifying the resulting alarm events for grouping within the cluster of resulting alarm events includes: (i) identifying a cluster of sequentially occurring alarm events; (ii) determining an occurrence time of each alarm event within the identified cluster of sequentially occurring alarm events; and (iii) responding to a determination that the occurrence time associated with one or more (or preferably each) alarm event within the identified cluster of sequentially occurring alarm events is separated from a previous alarm event within the identified cluster of sequentially occurring alarm events by more than a defined duration by identifying the cluster of sequentially occurring alarm events as a cluster of resulting alarm events.
[0024] At least one example embodiment also provides a system for alarm processing in a process control system. The system may include a processor-implemented server configured to: (i) detect one or more alarm events based on status data received from at least one device in the process control system; and (ii) in response to determining that the detected one or more alarm events match a stored alarm event pattern, (a) obtain an alarm event pattern response associated with the matched alarm event pattern, where the alarm event pattern response identifies one or more alarm response events; and (b) generate a control signal to implement one or more of the alarm response events.
[0025] The system may be configured to generate the stored alarm event pattern based on: (i) acquiring a set of historical data including alarm and event log data; (ii) correlating, based on the alarm and event log data in the acquired set of historical data, a reference alarm event with at least one of: (a) one or more concurrent candidate alarm events, where a respective probability of concurrent occurrence of each of the one or more candidate alarm events with the reference alarm event is determined to be equal to or greater than a defined first threshold; and (b) one or more operator actions, where a respective probability of concurrent occurrence of each of the one or more operator actions with the reference alarm event is determined to be equal to or greater than a defined second threshold; and (iii) including, in the stored alarm event pattern, an ordered sequence including the one or more concurrent candidate alarm events or the one or more concurrent operator actions, where the ordered sequence is generated based on a median time difference between a timestamp associated with the reference alarm event and a timestamp associated with the concurrent candidate alarm events or the concurrent operator actions.
[0026] The system may be configured such that an ordered sequence including one or more concurrent candidate alarm events or one or more concurrent operator actions is additionally generated based on a median absolute deviation from a median of the determined time differences for the concurrent candidate alarm events or concurrent operator actions.
[0027] In one embodiment, the system may be configured such that the step of correlating the reference alarm event with at least one of one or more concurrent candidate alarm events and one or more operator actions is performed based on alarm and event log data from a reduced set of historical data, and the reduced set of historical data is generated based on (i) identifying chattering alarm data in the set of acquired historical data, and (ii) generating the reduced set of historical data to include alarm and event data from the set of acquired historical data other than the identified chattering alarm data.
[0028] In another embodiment, the system may be configured such that the chattering alarm data includes (i) one or more alarm events having an alarm gap less than or equal to a predefined first duration, or (ii) one or more alarm events having an alarm lifetime less than or equal to a predefined second duration.
[0029] In certain embodiments, the system may be configured such that (i) the matched alarm event pattern includes one or more alarm events that do not have a corresponding concurrent operator action, and (ii) the obtained alarm event pattern response associated with the matched alarm event pattern includes an alarm suppression process flow for the one or more alarm events that do not have a corresponding concurrent operator action.
[0030] The system may additionally be configured such that determining that the alarm event does not have a corresponding concurrent operator action includes (i) determining one or more probabilities of coincidence of at least one detected operator action with the alarm event; and (ii) identifying the alarm event as an alarm event without a corresponding operator action in response to the determined one or more probabilities of coincidence being less than a predefined value.
[0031] The system may be configured such that (i) the matched alarm event pattern includes one or more alarm events having a corresponding set of concurrent operator actions, and (ii) the obtained alarm event pattern response associated with the matched alarm event pattern includes initiating a control signal for presenting a standardized set of operator actions to an operator in response to detecting the one or more alarm events or in response to detecting the matched alarm event pattern, wherein the standardized set of operator actions includes a corresponding set of concurrent operator actions.
[0032] In one embodiment, the system may be configured such that the operation of determining that an alarm event has a corresponding set of concurrent operator actions includes: (i) determining one or more probabilities of co-occurrence of the one or more operator actions with the alarm event; (ii) identifying the one or more operator actions as concurrent operator actions that occur concurrently with the alarm event in response to the determined one or more probabilities of co-occurrence being greater than a predefined value; (iii) identifying a timestamp associated with each of the one or more concurrent operator actions; and (iv) ordering each of the plurality of concurrent operator actions in a sequence, wherein the position of the concurrent operator action in the sequence is determined based on a median time difference between a timestamp associated with the concurrent operator action and a timestamp associated with the alarm event.
[0033] The system may be configured such that the position of the concurrent operator actions within the sequence is additionally determined based on a median absolute deviation from a median of the determined time differences for the concurrent operator actions.
[0034] In one embodiment, the system may be configured such that (i) the matched alarm event pattern includes a cluster of redundant alarm events, and (ii) the obtained alarm event pattern response associated with the matched alarm event pattern includes initiating an alarm suppression process flow for one or more alarm events in the cluster of redundant alarm events.
[0035] In a further embodiment, the system may be configured such that the matched alarm event pattern is generated by the following operations: (i) determining one or more probabilities of co-occurrence of at least one candidate alarm event with a reference alarm event; (ii) identifying the candidate alarm event as a coincident alarm event that occurs simultaneously with the reference alarm event in response to the determined one or more probabilities of co-occurrence being greater than a predetermined value; (iii) generating a cluster of alarm events including the reference alarm event and the identified one or more coincident alarm events; (iv) identifying a timestamp associated with each of a plurality of alarm events in the generated cluster of alarm events; and (v) ordering each of the plurality of alarm events in the generated cluster of alarm events in a sequence, wherein the position of the alarm event required to be ordered in the sequence is determined based on a median time difference between a timestamp associated with the reference alarm event and a timestamp associated with the candidate alarm event in the cluster of alarm events.
[0036] The system may be configured such that the position of the candidate alarm event within the sequence is additionally determined based on a median absolute deviation from a median of the time differences determined for the candidate alarm event.
[0037] In one embodiment, the system may be configured such that the act of identifying redundant alarm events for grouping within a cluster of redundant alarm events includes the acts of (i) identifying a cluster of alarm events that occur sequentially, (ii) determining the time of occurrence of each alarm event within the identified cluster, and (iii) responding to a determination that the time of occurrence of each alarm event within the cluster is separated from a previous alarm event within the cluster by less than a defined time value by identifying the cluster of alarm events as a cluster of redundant alarm events.
[0038] In another embodiment, the system may be configured to: initiate an alarm prediction process flow where (i) the matched alarm event pattern includes a cluster of consequential alarm events; and (ii) the obtained alarm event pattern response associated with the matched alarm event pattern includes responding to the detection of an occurrence instance of one or more preceding alarm events in the cluster of consequential alarm events by presenting to an operator information that predicts a future occurrence of one or more instances of subsequent alarm events in said cluster of consequential alarm events before the subsequent alarm event instance is detected.
[0039] In certain embodiments, the system may be configured such that a cluster of resultant alarm events is generated based on (i) an operation of identifying a timestamp associated with each of a plurality of resultant alarm events, and (ii) an operation of ordering each of the plurality of resultant alarm events in a sequence, wherein the position of a candidate alarm event to be ordered in the sequence is determined based on a median time difference between a timestamp associated with a reference alarm event and a timestamp associated with the candidate alarm event within the plurality of resultant alarm events.
[0040] The system may be configured such that the position of the candidate alarm event within the sequence is additionally determined based on a median absolute deviation from a median of the time differences determined for the candidate alarm event.
[0041] In certain embodiments, the system may be configured such that the act of identifying consequential alarm events for grouping within a cluster of consequential alarm events includes: (i) identifying a cluster of sequentially occurring alarm events; (ii) determining an occurrence time of each alarm event within the identified cluster of sequentially occurring alarm events; and (iii) responding to a determination that the occurrence time associated with one or more (or preferably each) alarm event within the identified cluster of sequentially occurring alarm events is separated from a previous alarm event within the identified cluster of sequentially occurring alarm events by more than a defined duration by identifying the cluster of sequentially occurring alarm events as a cluster of consequential alarm events.
[0042] At least one example embodiment additionally provides a computer program product for alarm processing in a process control system. The computer program product may comprise a non-transitory computer-usable medium having computer-readable program code embodied therein, the computer program product including instructions for performing within a processor-based computing system: (i) detecting one or more alarm events based on status data received from at least one device in the process control system; and (ii) in response to determining that the detected one or more alarm events match a stored alarm event pattern, (a) obtaining an alarm event pattern response associated with the matched alarm event pattern, where the alarm event pattern response identifies one or more alarm response events; and (b) generating a control signal to implement one or more of the alarm response events.
[0043] In one embodiment of the computer program product, the stored alarm event pattern may be generated based on the following steps: acquiring a set of historical data including alarm and event log data; (ii) correlating, based on the alarm and event log data in the acquired set of historical data, a reference alarm event with at least one of: (a) one or more concurrent candidate alarm events, where a respective probability of concurrent occurrence of each of the one or more candidate alarm events with the reference alarm event is determined to be equal to or greater than a defined first threshold; and (b) one or more operator actions, where a respective probability of concurrent occurrence of each of the one or more operator actions with the reference alarm event is determined to be equal to or greater than a defined second threshold; and (iii) including, in the stored alarm event pattern, an ordered sequence including the one or more concurrent candidate alarm events or the one or more concurrent operator actions, where the ordered sequence is generated based on a median time difference between a timestamp associated with the reference alarm event and a timestamp associated with the concurrent candidate alarm events or the concurrent operator actions. [Brief explanation of the drawings]
[0044] [Figure 1A] FIG. 1 illustrates an exemplary process control system of the type that may be used to manage a processing facility / industrial environment. [Figure 1B] Figure 1 shows a comparison of alarm system KPIs across various industries compared to benchmarks set by EEMUA. [Figure 2] FIG. 1 is a diagram of "alarm life" and "alarm gap" as understood in relation to an alarm system. [Figure 3]1 is a graph illustrating the effect of chattering alarm identification and elimination in an alarm system. [Figure 4A] 10A-10C are exemplary diagrams corresponding to alarms that do not trigger associated operator actions. [Figure 4B] 10A-10C are exemplary diagrams corresponding to alarm events that previously triggered associated operator actions. [Figure 5A] 1 is a flowchart illustrating a method for automatic alarm streamlining in accordance with the teachings of at least one exemplary embodiment. [Figure 5B] 5B is a flowchart illustrating a method of alarm processing based on alarm rationalization from the method of FIG. 5A. [Figure 6] 1 is a flowchart illustrating a method for identifying chattering alarm data within a set of alarm and event data in accordance with the teachings of at least one exemplary embodiment. [Figure 7A] FIG. 10 illustrates the use of alarm gap and alarm lifetime parameters to identify chattering alarms in accordance with the teachings of at least one exemplary embodiment. [Figure 7B] FIG. 10 illustrates the use of alarm gap and alarm lifetime parameters to identify chattering alarms in accordance with the teachings of at least one exemplary embodiment. [Figure 8] 10 is a flowchart illustrating a method for generating an alarm event pattern response for association with an alarm event pattern that includes one or more alarm events that do not require any associated operator action to return them to a normal state. [Figure 9] 1 is a flowchart illustrating a method for presenting an operator with a standardized set of operator actions to implement in response to a detected alarm event, in accordance with the teachings of at least one exemplary embodiment. [Figure 10]FIG. 10 is a diagram of a sequence of operator actions that may be seen in response to a Flow High alarm event to illustrate types of alarm events that have associated operator actions that have been triggered in the past. [Figure 11] 1 is a flowchart illustrating a method for generating an alarm event pattern response for association with an alarm event pattern that includes one or more clusters of redundant alarm events, in accordance with the teachings of at least one exemplary embodiment. [Figure 12] 12 is a flowchart illustrating a method for identifying redundant alarm events for clustering the redundant alarm events according to the teachings of the method of FIG. 11. [Figure 13] 1 is an example diagram of a group of alarms including one or more redundant alarm events of a type that may be subject to alarm suppression in accordance with the teachings of at least one example embodiment. [Figure 14] 1 is a flowchart illustrating a method of predictive alarm event detection in accordance with the teachings of at least one exemplary embodiment. [Figure 15] 1 is a flowchart illustrating a method for clustering related or consequential alarm events in accordance with the teachings of at least one exemplary embodiment. [Figure 16] 1 is an example diagram of a group of alarm events including one or related alarm events of a type that may be used for predictive alarm event detection in accordance with the teachings of at least one example embodiment. [Figure 17] 16 illustrates the creation of an exemplary time window for implementing the steps of the method of FIG. 15. [Figure 18] 16A-16C illustrate exemplary truncations of the type of time window that may be used to implement the steps of the method of FIG. 15. [Figure 19] FIG. 2 illustrates an exemplary first matrix used to store data corresponding to the occurrence of alarm events within a particular time window. [Figure 20]FIG. 10 illustrates the principle of time difference determination when a second alarm event occurs multiple times within a time window associated with a first alarm event. [Figure 21] FIG. 10 illustrates a second matrix used to store time difference data when a second alarm event occurs multiple times within a time window associated with a first alarm event. [Figure 22] FIG. 10 illustrates a third matrix used to store timestamp data in accordance with the teachings of at least one exemplary embodiment. [Figure 23] FIG. 10 illustrates a modified first matrix in accordance with the teachings of at least one exemplary embodiment. [Figure 24] FIG. 10 illustrates an updated second matrix in accordance with the teachings of at least one exemplary embodiment. [Figure 25] FIG. 25 illustrates a fourth matrix including a binary matrix generated based on the updated second matrix of FIG. 24 in accordance with the teachings of at least one exemplary embodiment. [Figure 26] FIG. 10 illustrates time window formation for a unique alarm event in accordance with the teachings of at least one exemplary embodiment. [Figure 27] FIG. 10 illustrates the calculation of the time difference between an associated action and a corresponding focused alarm event when there are two or more occurrences of the action within the same time window, in accordance with the teachings of at least one exemplary embodiment. [Figure 28] FIG. 1 illustrates a configured server in accordance with the teachings of at least one example embodiment of a type that may be implemented within a process control system or alarm system. [Figure 29] FIG. 1 illustrates an exemplary computer system according to which various embodiments of at least one exemplary embodiment may be implemented. DETAILED DESCRIPTION OF THE INVENTION
[0045] FIG. 1 illustrates an example process control system 100 of a type that may be used to manage an industrial environment. The process control system 100 includes multiple sensors, valve devices, or actuators 102a, 102b, and 102c. The sensors, valve devices, and / or actuators represent components that may perform any of a wide variety of functions. For example, sensors may measure parameters or characteristics of the industrial environment, such as temperature, pressure, and flow rate. Valve devices are used to regulate and / or direct fluid flow. Similarly, actuators may perform a wide variety of actions that change the state of the industrial environment or alter the parameters / characteristics being monitored by the sensors. For example, actuators may represent electric motors, hydraulic cylinders, and / or transducers.
[0046] One or more sensors / valve devices / actuators 102a-102c are connected to the controllers 104a, 104b via a field network (e.g., an Ethernet network, an electrical signal network such as a HART or FOUNDATION FIELDBUS network, a pneumatic control signal network, or some other or additional type of network) that facilitates interaction between the devices connected thereto. The controllers 104a, 104b may include one or more hardware controllers that use parameter data received from one or more sensors or from an operator or server to control the operation of the one or more actuators.
[0047] The process control system 100 may additionally include a server 106 configured to perform functions necessary to support the operation and control of the controllers 104a, 104b. Example functions of the server 106 may include logging information collected or generated by the controllers 104a, 104b and executing applications that control the operation of the controllers 104a, 104b, and thereby the operation of the sensors, valve devices, and / or actuators 102a-102c. The server 106 may additionally provide secure access to the controllers 104a, 104b and / or the sensors, valve devices, or actuators 102a-102c.
[0048] The process control system 100 also includes a database 108 configured to store information received from one or more of the server 106, the controllers 104a, 104b, and / or the sensors or valve devices or actuators 102a-102c. Additionally, the process control system 100 includes one or more operator terminals 110, each comprising a processor-implemented, network-communication-enabled data processing device that provides an operator access to the server 106, the controllers 104a, 104b, and / or the sensors or valve devices or actuators 102a-102c. Each operator terminal 110 may be configured to receive data input and / or control commands from an operator and to receive and display warnings, alerts, alarms, or other messages or indications generated by the server 106, the controllers 104a, 104b, and / or the sensors or valve devices or actuators 102a-102c.
[0049] Process control systems of the type shown in Figure 1A may implement alarm systems used to generate alarms in response to the detection of problems or deviations from specified process parameters. An alarm system is defined in the industrial standard ANSI / ISA-18.2 [1, p. 16] as follows: "An alarm system is a collection of hardware and software that detects alarm conditions, communicates indications of that condition to an operator, and records changes in alarm conditions." Alarm systems form an integral part of modern process control systems, such as distributed control systems (DCS) and supervisory control and data acquisition (SCADA) systems, and play a critical role in the safe and efficient operation of modern industrial plants, such as refineries, chemical plants, petrochemical plants, power plants, and water purification plants. The primary purpose of an alarm system is to quickly indicate the occurrence of any abnormal conditions so that operators can take corrective action to return the process to its normal operating range. A common problem with some industrial alarm systems is that they tend to generate far more alarms than operators can efficiently handle. This problem, known as alarm overload or alarm flooding, is typically the result of poorly configured or poorly functioning alarm systems. The scope of the problem is illustrated in Table 1, shown in Figure 1B, which presents statistics for three key performance indicators (KPIs) of alarm systems based on a survey of 39 industrial plants from the oil and gas, petrochemical, power, and other industries.
[0050] The corresponding benchmark values according to the relevant EEMUA-191 guidelines (The Engineering Equipment and Materials Users' Association) are also presented for comparison in Table 1. It is noted that the compiled statistics of KPIs from various industries significantly exceed the EEMUA benchmarks.
[0051] The occurrence of alarm overload can be understood by classifying alarms into two groups: (i) nuisance alarms and (ii) true (real) alarms. Nuisance alarms do not affect the process and therefore do not require any specific response or action from the operator. According to the industrial standard ANSI / ISA-18.2 [1, p. 18], an alarm must indicate an equipment malfunction, a process deviation, or an abnormal condition that requires a response. On the other hand, a true alarm must indicate an abnormal situation that requires the operator's attention or timely action to prevent the abnormal situation associated with the true alarm from adversely affecting the safety and / or efficiency of the process. Nuisance alarms are the primary cause of the phenomenon of alarm overload.
[0052] It is easy to understand that alarm overload is detrimental to the role played by alarm systems. Numerous alarms generated by alarm systems are difficult to address meaningfully. They provide no useful information and distract plant operators. Ineffective management of nuisance alarms can lead to accidents and can result in increased risk of fatigue and stress for operators who must make split-second decisions on how to respond when an alarm occurs. Meanwhile, true alarms are often buried among the numerous nuisance alarms and can be overlooked by operators. As a result, operators may mistakenly pay attention to less important / nuisance alarms or may not pay attention to the alarm system as a whole. As a result, true alarms that require operator action to correct abnormal process conditions may be ignored, and necessary corrective actions may be overlooked.
[0053] Alarm handling and alarm rationalization refer to processes for addressing some of these issues. Some processes for alarm rationalization involve cross-functional teams of plant stakeholders reviewing, justifying, and documenting whether each alarm configured within the alarm system or process control system meets the criteria for being an alarm (i.e., the alarm must be relevant and useful, indicate an abnormal condition, and require necessary corrective action by an operator). The primary goal of alarm rationalization is to minimize alarm burden to operators by presenting only true alarms that are relevant and require operator action. Another goal of alarm rationalization is to suppress nuisance alarms, or other alarms that do not qualify as true alarms, allowing operators to focus on alarms that actually require attention and corrective action.
[0054] Alarm rationalization may involve defining the attributes of each alarm (such as limits, priority, classification, and type) and documenting the cause and effect, response time, and operator actions. The alarm rationalization process is typically performed manually. It is tedious, time-consuming, and requires significant manual effort. With thousands of alarms in an alarm system, it is often difficult to identify appropriate candidate alarm events for review / investigation during alarm rationalization.
[0055] Therefore, a solution is needed that enables automated alarm rationalization to handle true alarms by accurately identifying nuisance alarms for suppression, grouping, and deletion. Predictive alarm handling is also needed to provide operators with advance notification of alarm conditions that are likely to generate one or more future true alarms within a defined time window so that the operator can take proactive action to correct or completely avoid the causal events that correspond to the predicted alarm condition.
[0056] At least one exemplary embodiment provides a method, system, and computer program for alarm processing, alarm prediction, and / or alarm rationalization within an industrial process control system. Certain embodiments provide a solution that accurately identifies nuisance alarms for suppression or other appropriate action and enables automated alarm rationalization to process true alarms. At least one exemplary embodiment additionally provides predictive alarm processing to provide an operator with advance notice of alarm conditions that are likely to generate one or more future true alarms within a defined time window so that the operator can take proactive action to correct or entirely avoid the causal event or condition corresponding to the predicted true alarm.
[0057] Unlike other solutions for alarm rationalization and / or alarm handling that focus only on the alarm count (i.e., number of alarm occurrences) and alarm duration (i.e., time gap between alarm activation and recovery) of individual alarms, at least one exemplary embodiment focuses on discovering correlations and patterns among a large volume of alarm events and using these discovered correlations and patterns for alarm handling activities including, but not limited to, alarm suppression, alarm grouping, alarm elimination, alarm prediction, and standardization of alarm response procedures. As a result, the exemplary embodiment may have the technical effect of reducing resources (e.g., computer processing power and / or manual power) used to detect and resolve alarms in a process environment.
[0058] For purposes of discussion, the terms "alarm" and "alarm event" are used interchangeably to describe an alarm condition, i.e., an alert, message, or communication generated in response to the detection of a deviation from normal operating conditions. The term "alarm condition" may be understood as a process, component, device, or environmental condition that falls outside a set of process, component, device, or environmental conditions defined as normal or acceptable conditions in an industrial environment.
[0059] In addition, the terms "alarm lifetime" and "alarm gap" should be understood according to the description of Figure 2. As shown in Figure 2, the term "alarm lifetime" can be understood as the time difference between an alarm notification event (i.e., when the alarm is activated / started) and an alarm recovery event (i.e., when the alarm returns to normal). The term "alarm gap" can be understood as the time difference between an alarm recovery event and the next alarm notification event for the same alarm.
[0060] The following description relies on reference to a variety of different types of alarms or alarm events, including true alarms, chattering alarms, redundant alarms, and consequential alarms, each of which is briefly described below.
[0061] As discussed above, a true alarm is an alarm event generated in response to the detection of an abnormal condition that requires timely operator attention or action to prevent the detected abnormal process condition, component condition, device condition, or environmental condition from having an adverse effect on the safety and / or efficiency of the process.
[0062] Chattering alarms are one of the most widely encountered nuisance alarms, typically found to contribute approximately 10% to 60% of alarm counts within industrial environments. According to the industrial standard ANSI / ISA-18.2, a "chattering alarm" can be defined as one that repeatedly transitions between an alarm state and a normal state within a short period of time. As a result, chattering alarms provide little or no time for operators to analyze such alarms and take remedial steps.
[0063] Chattering alarms include two closely related alarm types: short-lived alarms and repetitive alarms. Short-lived alarms include alarm events that have a short alarm life or alarm duration and do not repeat immediately, while repetitive alarms repeat almost immediately after recovery but do not necessarily have a short alarm life. Chattering alarms of either type are typically or often triggered due to random noise and / or disturbances detected in association with process variables, especially when the process variables are operating near their alarm limits.
[0064] Figure 3 is a graph illustrating the effect of identifying and removing chattering alarms in an alarm system. The data in the graph in Figure 3 was generated based on historical alarm and event (A&E) data from a water treatment plant that appeared to be suffering from alarm overload, with approximately 77 alarm occurrences per hour. The identified chattering alarms (including short-lived and recurring alarms) were found to contribute more than 81% of the total alarm count. Before alarm rationalization, the alarm and event (A&E) data contained a total alarm count of 107,324 alarm events over 58 days, with an average alarm rate of 77.10 alarms per hour and 12.85 alarms per 10 minutes. After identifying and removing chattering alarms, the alarm and event data was found to contain a total alarm count of 19,925 alarm events over 58 days, with an average alarm rate of 14.31 alarms per hour and 2.39 alarms per 10 minutes. In other words, nuisance alarms were found to comprise approximately 81% of the total number of alarm events over the 58 days studied.
[0065] Redundant alarms and consequential alarms are two other alarm types that contribute significantly to nuisance alarms.
[0066] A "redundant alarm" can be defined as a group of alarm events that often occur together within a short time period, typically within minutes. Because all alarm events in a redundant alarm group occur within such a short time period, operators cannot respond to each and every alarm event in the redundant alarm group. Furthermore, all alarm events in a redundant alarm group are triggered by the same root cause. Therefore, all of them essentially indicate the same underlying problem and do not necessarily require different corrective actions by the operator. Misconfigured alarm variables are often the primary reason for redundant alarms. Similarly, many variables are often configured to trigger alarm events without careful consideration of the need to link alarm events to such variables or the alarm scope. Previous research has found that up to 50% of configured alarm variables are redundant alarms that could have been eliminated through alarm rationalization.
[0067] On the other hand, a "consequential alarm" is a group of alarm events that occur one after the other with a significant time difference (e.g., greater than 5 minutes) between the individual alarm events. The time difference between the individual alarm events in a consequential alarm group is typically greater than 5 minutes, so that an operator can take necessary corrective action in a timely manner upon the occurrence of one or more preceding alarm events in the alarm sequence, just as corrective action can prevent the occurrence of one or more subsequent alarm events in the consequential alarm sequence.
[0068] Consequential alarms are generally proven to occur due to the propagation of anomalies resulting from physical connections. Large industrial processes typically consist of upstream and downstream devices that are physically connected. An abnormal condition in one process unit is highly likely to propagate to downstream or upstream devices through automatic control loops or recycle connections. As a result, the propagation of anomalies can result in a sequence of alarms over a period of time from process variables associated with the device on which the alarm was set.
[0069] In addition to the above, there are two other alarm types that have been found to be useful for the purposes of the present invention. The first type is an alarm event that does not require any associated operator action to return to normal, as shown in Figure 4A. Such alarms can be considered part of the larger category of nuisance alarms, i.e., alarms with no associated operator action that do not affect the process or indicate any abnormal situation.
[0070] Another category of alarms of the type shown in Figure 4B are alarms that have been previously indicated to prompt one or more operator actions to restore the corresponding process, component, device, or environmental condition to normal. In other words, when this type of alarm occurs, an operator is expected to take a defined set of timely corrective actions to normalize them. Operators' routine response to such alarms with a well-defined and consistent set of actions indicates that these alarm events are significant and therefore true alarms.
[0071] 5A is a flowchart illustrating a method for automatic alarm rationalization in accordance with the teachings of at least one example embodiment. In one embodiment, the method of FIG. 5A may be implemented in an alarm system, or in a process control system, or in a server in an alarm system or process control system.
[0072] Step 502A includes obtaining a set of historical data including alarm and event (A&E) log data. The set of historical data may be obtained by the server 106 from a database 108 in the process control system. The alarm and event log data may include data corresponding to any of: (i) one or more process, component, or device conditions that triggered an alarm event; (ii) the alarm event itself, such as data related to the type of alarm event, the cause of the alarm event, the time the alarm event occurred, the alarm recovery time, the duration of the alarm event, the cause of the alarm recovery, etc.; and (iii) one or more operator actions initiated in connection with the alarm event, such as operator actions initiated to affect the alarm recovery. The alarm and event log data may be extracted from one or more alarm logs, event logs, and / or alarm and event logs generated and stored by the process control system in the database 108.
[0073] Step 504A includes identifying chattering alarm data within the set of obtained historical data. Chattering alarm data may be identified in any number of different ways as will be apparent to those skilled in the art. Exemplary embodiments of steps for identifying chattering alarm data are discussed below in connection with FIG. 6.
[0074] Step 506A includes deleting, excluding, or excluding the identified chattering alarm data from the set of retrieved historical data to generate a reduced set of historical data that includes alarm and event data from the set of retrieved historical data other than the identified chattering alarm data.
[0075] Subsequently, step 508A includes generating a plurality of alarm event patterns based on the reduced set of historical data. The alarm event patterns may be generated according to several different methods, embodiments of which are discussed below. The time sequence of steps 504A-508A has proven to be important to the effectiveness of identifying a plurality of alarm event patterns; i.e., prior removal of chattering alarm data by step 504A and / or step 506A has been found to significantly improve the accuracy of the subsequent identification of alarm event patterns in step 508A. This is because chattering alarms are considered "noise" in the alarm and event log data, and removal of such chattering alarms may help identify alarm event patterns of true alarms in the log data.
[0076] Step 510A includes associating with each identified alarm event pattern a corresponding alarm event pattern response, which may include a set or sequence of instructions, actions, or steps intended to be implemented in response to future detection of the alarm event pattern associated with the alarm event pattern response. Exemplary alarm event pattern responses may include any of the following: (i) generating an alarm, alert, or notification corresponding to each alarm event in the detected alarm event pattern; (ii) suppressing or deleting one or more alarm events in the detected alarm event pattern and generating alarms, alerts, or notifications corresponding to one or more other alarm events in the detected alarm event pattern; (iii) presenting a standardized operating procedure, standardized guidance, or other instructions to an operator, for example, to achieve alarm recovery or otherwise respond to one or more alarm events in the alarm event pattern; and / or (iv) notifying an operator of one or more future events or conditions (e.g., a fault condition or event, a deviation condition or event, an alarm condition or event, etc.) in the industrial environment that are predicted to occur or are determined to have a high probability of occurrence based on one or more alarm events in the detected alarm event pattern. The occurrence times of the predicted alarm events are also presented. More details about alarm event pattern responses are provided below.
[0077] Step 512A includes generating a data record including data related to the identified alarm event pattern and data related to the associated alarm event pattern response. The data record may additionally store information linking or associating the identified alarm event pattern data with the corresponding alarm event pattern response data. The data record may be stored in a database, for example, database 108 within or communicatively coupled to the process control system or alarm system.
[0078] Figure 5B is a flow chart illustrating a method of alarm processing based on the alarm rationalization process described in connection with the method of Figure 5 A. In one embodiment, the method of Figure 5B may be implemented in an alarm system, or in a process control system, or in a server 106 in an alarm system or process control system.
[0079] Step 502B includes detecting one or more alarm events. The one or more alarm events may be detected by the server 106 communicatively coupled to the process control system or alarm system and may be based on (i) condition data received from one or more sensors, or actuators, or other devices within the process control system or alarm system, and (ii) one or more alarm event detection rules, alarm event detection criteria, and / or alarm event detection models.
[0080] In step 504B, in response to determining that the detected alarm event or events match a stored alarm event pattern, an alarm event pattern response is retrieved from a database configured to store alarm event pattern responses or from a data record in such a database, The retrieved alarm event pattern response is an alarm event pattern response previously associated with the matched alarm event pattern according to method step 510A and / or method step 512A of FIG.
[0081] As discussed above, in exemplary embodiments, the obtained alarm event pattern response may include initiating control signals to implement one or more alarm response events, where the alarm response events include (i) generating an alarm, alert, or notification corresponding to each alarm event in the matched alarm event pattern; (ii) suppressing or removing one or more alarm events in the matched alarm event pattern and generating an alarm, alert, or notification corresponding to one or more other alarm events in the matched alarm event pattern; or (iii) taking standardized action, e.g., to achieve alarm recovery or otherwise respond to one or more alarm events in the matched alarm event pattern. (iv) sending and receiving data to and from one or more sensors, actuators, or other devices in the process control system to correct detected deviations from normal process, component, device, or environmental conditions; and / or (v) notifying an operator of one or more future events or conditions (e.g., fault condition or event, deviation condition or event, alarm condition or event, etc.) in the industrial environment that are predicted to occur or determined to have a high probability of occurrence based on the one or more alarm events in the matched alarm event pattern. The time of occurrence of the predicted alarm event is also presented.
[0082] Step 506B includes implementing one or more (preferably all) of the instructions, actions, or steps defined in the obtained alarm event pattern response.
[0083]
[0023] Figure 6 is a flow chart illustrating a method for identifying chattering alarm data in a set of alarm and event (A&E) data in accordance with the teachings of at least one example embodiment. The method of Figure 6 may be implemented for purposes of step 504A and / or step 506A of Figure 5A. In one embodiment, the method of Figure 6 may be implemented in an alarm system, or in a process control system, or in a server 106 in an alarm system or process control system.
[0084] Step 602 includes identifying a first set of alarm events within a set of acquired historical data including alarm and event log data (e.g., the set of acquired historical data from step 502A of FIG. 5A), wherein each alarm event within the first set of alarm events has an alarm gap less than or equal to a predefined first duration.
[0085] For example, referring to FIG. 7A, in which the predefined first duration is one minute, step 602 includes identifying all alarm events with an alarm gap of one minute or less. The predefined first duration may be selected to have a time value that represents a duration threshold at which all or a significant number of alarms tend to repeat. In other words, the predefined first duration may represent a time value that can be used to determine whether consecutive alarm events qualify as "repeated alarms," a type of alarm that may be considered a nuisance alarm. Thus, by identifying alarm events with an alarm gap of less than or equal to the predefined first duration, step 602 identifies a first set of repetitive alarm events that have a reasonable or high probability of being nuisance alarm events, i.e., repeat alarms.
[0086] Step 604 includes identifying a second set of alarm events within the set of acquired historical data, each alarm event in the second set of alarm events having an alarm life less than or equal to a predefined second duration.
[0087] For example, referring to FIG. 7B, in which the predefined second duration is one minute, step 604 includes identifying all alarm events having an alarm lifespan of one minute or less. The predefined second duration may be selected to have a time value representing a duration threshold at which the alarm lifespans of all or a significant number of short-lived alarms tend to repeat. In other words, the predefined second duration may represent a time value that can be used to determine whether an alarm event qualifies as a "short-lived alarm," a type that may be considered a nuisance alarm. Thus, by identifying alarm events having an alarm lifespan less than the predefined second duration, step 604 identifies a second set of short-lived alarm events that have a reasonable or high probability of being nuisance alarm events, i.e., short-lived alarms.
[0088] It will be appreciated that the first predefined duration and the second predefined duration may have the same time value or may have different time values.
[0089] Step 606 includes classifying the data corresponding to the alarm events in the first set of alarm events (or the alarm events themselves) and the data in the second set of alarm events (or the alarm events themselves) as chattering alarm data (or chattering alarm events). In a more specific embodiment, step 606 may include classifying the data corresponding to the alarm events in the first set of alarm events (or the alarm events themselves) as recurring alarm data (or recurring alarms) and / or classifying the data corresponding to the alarm events in the second set of alarm events (or the alarm events themselves) as short-lived alarm data (or short-lived alarms). The classification information may be stored in the alarm system, or the process control system, or a database communicatively coupled to a server in the alarm system or the process control system.
[0090] It will be appreciated that once chattering alarm data in a set of historical data (including alarm and event log data) is identified based on the method steps of Figure 6, the chattering alarm data may be removed from the set of historical data, and the remaining data items in the set of historical data may be used to generate a reduced set of historical data (e.g., according to step 506A of Figure 5A), which may be analyzed to detect alarm event patterns according to the remaining method steps of Figure 5A.
[0091] 8 is a flowchart illustrating a method for generating an alarm event pattern response for association with an alarm event pattern that includes one or more alarm events that do not require any associated operator action to return them to a normal state. The method of FIG. 8 may be implemented as part of step 506B of FIG. 5B if the obtained alarm event pattern response includes suppressing one or more detected alarms or alarm events. In one embodiment, the method of FIG. 8 may be implemented in an alarm system, or in a process control system, or in a server 106 in an alarm system or process control system.
[0092] Step 802 includes identifying, for (or within) the alarm event pattern, one or more alarm events that do not have a corresponding concurrent operator action. It should be understood that the term "concurrent" does not only refer to a situation in which a corresponding operator action occurs simultaneously with its associated alarm event. The term also refers to a situation in which a corresponding operator action is performed after its associated alarm event. Identifying one or more alarm events that do not have a corresponding concurrent operator action may be performed by analyzing historical data, including alarm and event logs, to determine and identify alarm events that do not consistently (i.e., at least twice, preferably three or more times) have a corresponding operator action for alarm recovery or to restore the associated alarm to a normal state. In one embodiment, identifying candidate alarm events that do not have a corresponding concurrent operator action includes: determining one or more probabilities of a co-occurrence of at least one detected operator action and a candidate alarm event; identifying the candidate alarm event as an alarm event that does not have a corresponding coincident operator action in response to the determined one or more probabilities of coincidence being less than a predefined value; may include:
[0093] Step 804 includes including, within the alarm event pattern response corresponding to the identified alarm event pattern, an alarm suppression process flow for addressing instances of one or more alarm events found to have no corresponding operator action. In certain embodiments, the alarm suppression process flow may include instructions for preventing an alert, alarm, or notification associated with the one or more alarm events from being submitted or presented to an operator, which may include, for example, postponing the occurrence of the one or more alarm events or downplaying the one or more alarm events.
[0094] By incorporating within the alarm event pattern response (corresponding to the alarm event pattern) an alarm suppression process flow related to alarm events that are not associated with any relevant operator action, the method of FIG. 8 ensures that such alarm events are suppressed when they occur.
[0095] 9 is a flowchart illustrating a method for presenting an operator with a standardized set of operator actions to implement in response to a detected alarm event, in accordance with the teachings of at least one exemplary embodiment. The method of FIG. 9 may be implemented as part of step 506B of FIG. 5B when the obtained alarm event pattern response includes presenting an operator with a standardized set of operator actions to implement. In one embodiment, the method of FIG. 9 may be implemented within an alarm system, or within a process control system, or within a server 106 within an alarm system or process control system.
[0096] Step 902 includes identifying one or more alarm events having a corresponding set of concurrent operator actions for (or within) the alarm event pattern, each operator action occurring concurrently in response to one or more alarm events. Step 902 of identifying alarm events having a set of concurrent operator actions may be implemented by analyzing historical data including alarm and event logs to determine and identify alarm events having corresponding concurrent operator actions for alarm recovery or for restoring the associated alarm to a normal state. In one embodiment, identifying candidate alarm events having a corresponding set of concurrent operator actions includes: determining one or more probabilities of coincidence of one or more operator actions with a reference alarm event; identifying the one or more operator actions as concurrent operator actions that occur concurrently with the reference alarm event in response to the determined one or more probabilities of the co-occurrence of the one or more operator actions being greater than a predefined value; identifying a timestamp associated with each of the one or more concurrent operator actions; ordering each of the plurality of concurrent operator actions in a sequence, wherein the position of the concurrent operator action sought to be ordered in the sequence is determined based on (i) a median time difference between a timestamp associated with the concurrent operator action and a timestamp associated with a reference alarm event, and (ii) optionally based on a median absolute deviation from the median time difference determined for the operator action; may include:
[0097] Step 904 includes classifying the corresponding set of operator actions that have been performed consistently (i.e., at least twice, preferably three or more times) as a standardized set of operator actions for responding to the detection of one or more alarm events.
[0098] Step 906 includes associating a standardized set of operator actions with one or more alarm events in the alarm event pattern.
[0099] Step 908 includes initiating a control signal or instruction to present a standardized set of operator actions to an operator in response to detecting one or more alarm events or in response to detecting an alarm event pattern within an alarm event pattern response corresponding to the alarm event pattern.
[0100] By incorporating instructions for presenting a standardized set of operator actions within an alarm event pattern response (corresponding to the alarm event pattern), the method of FIG. 9 ensures that an operator can receive or be prompted with standardized guidance for responding to one or more detected alarm events within the alarm event pattern when they occur.
[0101] Figure 10 is an example sequence of operator actions observed in response to a flow high alarm event to illustrate the type of alarm event that has corresponding concurrent operator actions that have been triggered in the past, as described in connection with Figure 9. The example sequence is based on an example analysis of data from a petrochemical plant.
[0102] In particular, a sequence of consistently observed / recorded operator actions was detected between the activation and recovery of a "flow high" alarm. As shown in Figure 10, a consistent procedure for responding to the detection of a "high" alarm in a flow variable includes (i) first manipulating / adjusting / fine-tuning the manipulated variable (MV) of the flow loop in manual (MAN) mode, (ii) subsequently placing the flow loop in automatic (AUT) mode, and (iii) further adjusting the setpoint (SV) of the flow variable. These consistently observed operator actions can be used to generate a standardized set of operator actions that can be presented to the operator as standardized guidance for responding to a "flow high" alarm event within the detected alarm event pattern.
[0103] FIG. 11 is a flowchart illustrating a method for generating an alarm event pattern response for association with an alarm event pattern that includes one or more clusters of redundant alarm events, in accordance with the teachings of at least one example embodiment.
[0104] The method of Figure 11 may be implemented as part of step 506B of Figure 5B if the obtained alarm event pattern response includes suppressing one or more detected alarms or alarm events. In one embodiment, the method of Figure 11 may be implemented in an alarm system, or in a process control system, or in a server 106 in an alarm system or process control system.
[0105] Step 1102 includes identifying one or more clusters of redundant alarm events for (or within) the alarm event pattern, each cluster including multiple concurrent alarm events occurring in sequence within a very short time period. It should be understood that the term "concurrent" does not necessarily refer only to a situation where multiple alarm events occur at the same time. It may also refer to a situation where multiple alarm events occur at different times, but all within a short time period. Identifying clusters of redundant alarm events may be implemented by analyzing historical data including alarm and event logs to determine and identify one or more clusters of redundant alarm events. In one embodiment, identifying clusters of concurrent alarm events occurring in sequence within a very short time period includes: determining one or more probabilities of co-occurrence of at least one candidate alarm event with a reference alarm event; identifying the candidate alarm event as a concurrent alarm event that occurs concurrently with the reference alarm event in response to the determined one or more probabilities of co-occurrence being greater than a predefined value; generating a cluster of alarm events comprising a reference alarm event and one or more identified co-occurring alarm events, i.e., one or more identified candidate alarm events; identifying a timestamp associated with each of a plurality of alarm events in the cluster of generated alarm events; The method may include a step of clustering and ordering the plurality of alarm events based on: ordering each of the plurality of alarm events in the cluster of generated alarm events in a sequence, wherein the position of the alarm event to be ordered in the sequence is determined based on (i) a median time difference between a timestamp associated with a reference alarm event in the cluster of alarm events and a timestamp associated with a candidate alarm event; and (ii) optionally based on a median absolute deviation from the median time difference determined for the candidate alarm event.
[0106] Step 1104 includes incorporating or including an alarm suppression process flow with one or more alarm events in each identified cluster of redundant alarm events within an alarm event pattern response corresponding to the alarm event pattern. The alarm suppression process flow is intended to address instances of one or more alarm events in each identified cluster of redundant alarm events. In certain embodiments, the alarm suppression process flow may include instructions for preventing the submission or presentation to an operator of alerts, alarms, or notifications associated with one or more alarm events in each identified cluster of redundant alarm events, and the instructions may include, for example, postponing the occurrence of one or more alarm events or downplaying one or more alarm events. In certain embodiments, the alarm suppression process flow may include (i) rejecting, downplaying, deferring, ignoring, or refusing an instruction to generate an alert, alarm, or notification associated with one or more alarm events in each identified cluster of redundant alarm events, and (ii) generating an alert, alarm, or notification associated with at least one other alarm event in each identified cluster of redundant alarm events.
[0107] The method of FIG. 11 eliminates one or more redundant alarms or alarm events as they occur by identifying clusters of redundant alarm events and generating an alarm suppression workflow for suppressing one or more alarm events within each identified cluster of redundant alarm events.
[0108] FIG. 12 is a flow chart illustrating a method for identifying redundant alarm events for clustering the redundant alarm events according to the teachings of the method of FIG.
[0109] The method of Figure 12 may be implemented as part of step 1102 of Figure 11. In one embodiment, the method of Figure 12 may be implemented in an alarm system, or in a process control system, or in a server 106 in an alarm system or process control system.
[0110] Step 1202 involves identifying clusters of sequentially occurring alarm events for (or within) an alarm event pattern.
[0111] Step 1204 includes determining the occurrence time of each alarm event in the cluster. The occurrence time associated with the alarm event may be determined based on a recorded timestamp associated with the alarm event.
[0112] Step 1206 includes responding to a determination that the occurrence time of one or more (preferably each) alarm event in the cluster is separated from the immediately preceding alarm event in the cluster by less than a defined duration / interval / time value by classifying the cluster of alarm events as a cluster of redundant alarm events.
[0113] By enabling the identification of redundant alarm event clusters, the method of FIG. 12 prepares for the subsequent occurrence of an alarm suppression workflow to suppress one or more alarm events within each identified cluster of redundant alarm events (e.g., by the method of FIG. 11).
[0114] FIG. 13 is an example diagram of a group of alarms that includes one or more redundant alarm events of a type that may be subject to alarm suppression in accordance with the teachings of at least one example embodiment.
[0115] This figure is based on an exemplary analysis of data from a petrochemical plant and shows an example of a redundant alarm group identified after analyzing alarm and event logs from the petrochemical plant. As shown, three individual alarms were consistently observed to occur together within a very short time period (approximately 22 seconds): a "Pump P1 Motor Stopped" alarm, a "Vessel V2 / 3 Pressure Drop" alarm, and a "Pump P1 Running" alarm. Analysis reveals that all three alarms are triggered by the same root cause and are therefore classified as a cluster of redundant alarm events according to the methodology of Figure 12.
[0116]
[0013] Figure 14 is a flowchart illustrating a method for predictive alarm event detection in accordance with the teachings of at least one example embodiment. The method of Figure 14 may be implemented as part of step 506B of Figure 5B when the obtained alarm event pattern response includes a step of predictive alarm event detection based on one or more detected alarms or alarm events. In one embodiment, the method of Figure 14 may be implemented within an alarm system, or within a process control system, or within a server 106 within an alarm system or process control system.
[0117] Step 1402 includes identifying one or more clusters of resulting alarm events for (or within) the alarm event pattern, which clusters include multiple alarm events occurring in a sequence. Identifying alarm events occurring in a sequence may be implemented by analyzing historical data, including alarm and event logs, to determine and identify alarm events that occur consistently (i.e., at least twice, preferably three or more times) in the same sequence. In one embodiment, identifying clusters of alarm events occurring in a sequence includes: determining one or more probabilities of co-occurrence of at least one candidate alarm event with a reference alarm event; identifying the candidate alarm event as a concurrent alarm event that occurs concurrently with the reference alarm event in response to the determined one or more probabilities of co-occurrence being greater than a predefined value; generating a cluster of alarm events comprising a reference alarm event and the identified one or more co-occurring alarm events, i.e., one or more candidate alarm events; identifying a timestamp associated with each of a plurality of alarm events; The method may include a step of clustering and ordering the plurality of alarm events based on a step of ordering each of the plurality of alarm events in the sequence, wherein the position of a candidate alarm event to be ordered in the sequence is determined based on (i) a median time difference between a timestamp associated with a reference alarm event in the cluster of alarm events and a timestamp associated with the candidate alarm event, and (ii) optionally based on a median absolute deviation from the median time difference determined for the candidate alarm event.
[0118] Step 1404 includes including an alarm prediction process flow together with the identified cluster of consequential alarms in an alarm event pattern response corresponding to the identified alarm event pattern. The alarm prediction process flow may include responding (or initiating a response) to the detection of an occurring instance of one or more preceding alarm events in the identified cluster of consequential alarm events by presenting to an operator information that predicts a future occurrence of one or more instances of subsequent alarm events in said identified cluster of consequential alarm events before said instances of the subsequent alarm events are detected.
[0119] In certain embodiments, the information predicting the future occurrence of one or more instances of a subsequent alarm event within the identified cluster of consequential alarm events includes information identifying one or more predicted time values representing estimated time values or time windows within which the future occurrence of one or more instances of the subsequent alarm event is predicted to occur.
[0120] The identified cluster of resulting alarm events includes one or more preceding alarm events and one or more subsequent alarm events, and each of the one or more subsequent alarm events includes an alarm event that occurs subsequent to the corresponding one or more preceding alarm events within the identified cluster of resulting alarm events.
[0121] By incorporating instructions within an alarm event pattern response (corresponding to an alarm event pattern) for predicting the occurrence of one or more instances of a subsequent alarm event based on the detection of one or more instances of a preceding alarm event, the method of Figure 14 enables an operator to receive advance notification of a likely alarm event so that the operator can take appropriate action to remedy or eliminate abnormal or deviant process, component, device, or environmental conditions that may cause the subsequent alarm event, even before the subsequent alarm event occurs. The operator is also notified of the probability and time of the predicted occurrence of the subsequent alarm.
[0122] Figure 15 is a flow chart illustrating a method for clustering related consequential alarm events according to the teachings of the method of Figure 14. The method of Figure 15 may be implemented as part of step 1402 of Figure 14. In one embodiment, the method of Figure 15 may be implemented in an alarm system, or in a process control system, or in a server 106 in an alarm system or process control system.
[0123] Step 1502 includes identifying clusters of sequentially occurring alarm events for (or within) an alarm event pattern. Identifying sequentially occurring alarm events may be implemented by analyzing historical data, including alarm and event logs, to determine and identify sequences of alarm events that occur consistently (i.e., occur at least twice, and preferably three or more times).
[0124] Step 1504 includes determining a time of occurrence of each alarm event in the identified cluster of sequentially occurring alarm events. The time of occurrence associated with the alarm event may be determined based on a recorded timestamp associated with the alarm event.
[0125] Step 1506 includes responding to a determination that the occurrence times associated with one or more (or preferably each) alarm event in the cluster are separated from the immediately preceding alarm event in the cluster by more than a defined duration, e.g., five minutes, or more than five minutes, by classifying the cluster of alarm events as a cluster of consequential alarm events.
[0126] By enabling the identification of consequential alarm event clusters, the method of FIG. 15 enables the subsequent generation and implementation of an alarm prediction process flow for predicting future occurrences of one or more alarm events in a consequential alarm event cluster in response to detecting the occurrence of one or more earlier-occurring alarm events in the same consequential alarm event cluster (e.g., by the method of FIG. 14).
[0127] FIG. 16 is an example diagram of a group of alarm events including one or more related alarm events of a type that may be used for predictive alarm event detection in accordance with the teachings of at least one example embodiment.
[0128] This figure is based on an exemplary analysis of data from a petrochemical plant and shows an example of identified consequential alarm event clusters obtained after analyzing historical alarm and event logs from the petrochemical plant. As shown in Figure 16, the alarm "Vessel 1 Low Pressure" alarm can be predicted approximately 52 minutes before its actual occurrence based on the occurrence of its precursor "Vessel 1 Low Level" alarm.
[0129] The following paragraphs discuss an exemplary implementation of the method of at least one exemplary embodiment.
[0130] Exemplary Method for Identifying Redundant and / or Consequential Alarms Preliminary Nuisance Alarm Rejection: As discussed above, the identification of alarm event patterns may optionally be preceded by the identification and rejection of nuisance alarm events from the analysis. For example, low priority alarms and chattering alarms, such as short-lived and recurring alarms, can be identified from the retrieved alarm and event logs, and the alarm and event log data set can be reduced for further analysis by first removing these unwanted alarm events.
[0131] Generating a List of Unique Alarms: The method may then include creating a list of unique alarm events and determining a total number of occurrences of each unique alarm event in the list. Unique alarm events within the alarm and event log are first identified. In one embodiment, a combination of two fields within the alarm and event log is used to define the identified unique alarm event. For example, a unique combination of a tag name (e.g., 29-FIC-1111.PV) field and a condition field (e.g., High) can be used to identify a unique alarm event as "29-FIC-1111.PV_High." Once all unique alarm events within the alarm and event log have been identified, the total number of occurrences of each unique alarm event within the entire alarm and event log can be determined, and a list can be generated that includes all unique alarm events found within the entire alarm and event log and the corresponding number of occurrences.
[0132] Time Window Generation: A time window corresponding to each unique alarm event can then be generated. For example, it may be assumed that for each unique alarm event, a time window of a total duration of 2L hours is considered, beginning L hours before the occurrence of the unique alarm and ending L hours after the occurrence of the unique alarm. Thus, considering the parameter L=2 hours, if the occurrence time of a unique alarm event C is 8:00:00 AM on December 12, 2019, a time window of a total duration of 4 hours is considered, from 6:00:00 AM to 10:00:00 AM on December 12, 2019. Because a unique alarm event can occur multiple times within the alarm and event log, the time window discussed above is considered for each occurrence of each unique alarm event. For example, if a unique alarm event C occurs a total of five times, five such time windows are created. An exemplary time window of the type discussed in this paragraph is shown in FIG. 17.
[0133] Continuing with the above example, if there are multiple occurrences of the same unique alarm event C within the time window of alarm C, the time window of a total duration of 2L hours is truncated immediately after or immediately before the other occurrences of alarm event C such that no other occurrences of the same alarm event C are included within the time window of an instance of alarm C. The exemplary time windows discussed in this paragraph are shown in FIG.
[0134] Creating the Actual Matrix: After a time window is created for each unique alarm event, other alarm events that occur within each time window are identified and the number of their occurrences within each time window is calculated. This information is sometimes called the Actual Matrix. C ×N unique It can be stored in the form of a matrix.
[0135] The total number of rows in the actual matrix (N C ) = total number of time windows for C = total number of occurrences of alarm event C in the log. Total number of columns in the actual matrix (N unique) = number of unique alarm events in the log. Element N in the actual matrix ij represents the total number of occurrences of the jth unique alarm event within the ith time window. An illustrative example of an actual matrix that may be used to store data corresponding to the occurrence of alarm events within a particular time window is provided in Table 2, shown in FIG.
[0136] Generation of the time difference matrix: The time difference matrix is again calculated as (N C ×N unique ) matrix, which stores the time difference between the unique alarm event C and other alarm events in each time window for the alarm event C. The element N in the time difference matrix ij represents the time difference between the jth unique alarm event and the unique alarm event C in the i-th time window of the alarm event C.
[0137] As shown in Figure 20, which illustrates the principle of time difference determination when there are multiple occurrences of a second alarm event within a time window associated with a first alarm event, if there are multiple occurrences of a second alarm event A within the same time window, the occurrence of alarm event A that is closest to the occurrence of alarm event C within that time window is considered to calculate the time difference between the first alarm event C and the second alarm event A. Table 3 in Figure 21 shows the time difference matrix used to store time difference data when there are multiple occurrences of a second alarm event within a time window associated with a first alarm event.
[0138] Generate the timestamp matrix: Then, (N C ×N unique A timestamp matrix may be generated that includes a _ , ...
[0139] A next step for implementing at least one exemplary embodiment includes checking whether the timestamp values of a given associated alarm are exactly the same for two consecutive time windows. If the timestamp values of a given associated alarm in the two time windows are exactly the same, then both the actual matrix and the time difference matrix can be updated as follows: In the time difference matrix for the associated alarm event, the entry for the time window with the smallest absolute time difference is retained, and other time window entries are set to null values. In the actual matrix for the associated alarm event, the entry for the time window with the smallest absolute time difference is kept, and the other time window entries set are set to zero.
[0140] The procedure for updating both the actual matrix in Table 2 and the time difference matrix in Table 3 based on the timestamp matrix in Table 4 is described in more detail below with reference to the figures. It can be seen that the timestamp matrix in Table 4 has exactly the same timestamps (rows #1 and #2, ie, the highlighted cells in Table 4) in two consecutive windows. In the time difference matrix in Table 3, the entry (row #2) of the time window with the smallest absolute time difference (1050 seconds) is kept. The entries (row #1) of the other time windows are set to null. Similarly, the actual matrix in Table 2 was also updated. Here, the entry (row #2) for the time window with the smallest absolute time difference (1050) is retained. The other time window entries (row #1) are set to zero. The updated actual matrix and updated time difference matrix are shown as Table 5 (see FIG. 23) and Table 6 (see FIG. 24), respectively. The updated cells in each table are highlighted.
[0141] Generating a Binary Matrix: The next step for an implementation involves generating a binary matrix from the actual matrix. To generate the binary matrix, all positive valued entries in the actual matrix are set to 1. An example binary matrix derived from the updated actual matrix in Table 5 (see FIG. 23) is shown as Table 7 in FIG. 25.
[0142] Generating forward probability: The forward probability of alarm event A with respect to alarm event C indicates the probability of simultaneous occurrence of alarm event A with alarm event C, given the occurrence of alarm event C. The forward probability of alarm event A with respect to alarm event C can be calculated from the binary matrix of alarm event C as follows: Forward probability = (number of time windows in which alarm event A exists) / (total number of time windows in which alarm event C exists) (Equation 1) The numerator of the above forward probability formula is simply the column sum of the binary matrix described above, and the denominator is the total number of occurrences of alarm event C within the entire alarm and event log.
[0143] Generating backward probability: The backward probability of alarm event A with respect to alarm event C indicates the probability of simultaneous occurrence of alarm event A with alarm event C, given that alarm event A occurs. The backward probability of alarm event A with respect to alarm event C is Backward probability = (total number of instances of alarm event A within all time windows of alarm event C) / (total number of instances of alarm event A within the entire log) (Equation 2) The numerator in the above backward probability formula can be obtained from the column sum of the updated actual matrix.
[0144] If both the forward probability value and backward probability value of alarm event A with respect to alarm event C exceed a predefined threshold (e.g., 80%) set by an operator or user, alarm event A is considered to be strongly associated with alarm event C, and vice versa, for example, alarm event A and alarm event C belong to a cluster of redundant alarm events.
[0145] Calculation of median time difference and median absolute deviation (MAD) from the median time difference: For all relevant alarms of alarm event C that meet the predefined criteria of forward probability and backward probability, both the median time difference and the median absolute deviation (MAD) from the median time difference are calculated based on the time difference matrix.
[0146] After identifying all associated alarm events with respect to alarm event C, the identified associated alarm events are sorted in ascending order based on their median time difference from alarm event C. This provides the sequence of alarm events in the alarm event cluster / alarm event group along with the median time difference between them.
[0147] Exemplary methods for identifying no-action alarms and alarm response procedures for true alarms Preliminary Nuisance Alarm Rejection: Identifying alarm event patterns may optionally be preceded by identifying and rejecting nuisance alarm events from the analysis. For example, low priority alarm events, as well as chattering alarm events, such as short-lived and recurring alarm events, may be identified from the retrieved alarm and event logs, and the data set in the alarm and event logs may be reduced for analysis by first removing these unnecessary alarms.
[0148] Creating a List of Unique Alarms: The method then includes creating a list of unique alarm events and determining the total number of occurrences of each unique alarm event in the list. Unique alarm events within the alarm and event log are first identified. In one embodiment, a combination of two fields within the alarm and event log is used to define the identified unique alarm event. For example, a unique combination of a tag name (e.g., 29-FIC-1111.PV) field and a condition field (e.g., High), i.e., "29-FIC-1111.PV_High," can be used to identify the unique alarm "29-FIC-1111.PV_High." Once all unique alarm events within the alarm and event log have been identified, the total number of occurrences of each unique alarm event within the entire alarm and event log can be determined, and a list can be generated that includes all unique alarm events found within the entire alarm and event log and the corresponding number of occurrences.
[0149] Generating a List of Unique Operator Actions and a Total Count of Occurrences of Each Unique Operator Action: Unique operator actions throughout the alarm and event log are first identified. A combination of two fields within the alarm and event log may define a unique operator action. For example, the unique combination of a tag name (e.g., 29-FIC-1111) field and a parameter field (e.g., AUT), which is 29-FIC-1111_AUT, defines a unique operator action in this example. Once all unique operator actions within the log have been identified, a total count of the occurrences of each unique operator action within the entire log is calculated, and a list containing all unique operator actions within the entire log and their corresponding number of occurrences is created.
[0150] Identifying all alarm notification and recovery timestamp pairs for each unique alarm event: For each unique alarm event C, all alarm notification timestamps and their corresponding alarm recovery timestamps are identified and stored. Then, all alarm recovery times for each unique alarm event C are calculated from the time difference between the alarm notification timestamps and their corresponding alarm recovery timestamps. Finally, the median time to recovery for each unique alarm event and the MAD from the median times to recovery are determined.
[0151] Creating a time window for each unique alarm: For each unique alarm event C, a time window spanning from its alarm notification time to its alarm recovery time is determined or generated, and then the time window is extended on both sides by adding and subtracting the time duration to consider after alarm recovery (e.g., l1 minutes) and the time duration to consider before alarm notification (e.g., l2 minutes), respectively. The parameters l1 and l2 should be configurable by the user or operator. The time window generated in this way is shown in Figure 26 below.
[0152] If the extended time window begins before the previous alarm recovery event, the time window extension may be truncated to the previous alarm recovery event, whereas if the extended time window ends after the next alarm notification event, the time window extension may be truncated to the next alarm notification event.
[0153] Creating the Actual Matrix: After all time windows have been created for each unique alarm event, the operator actions that occur within each time window are identified and the number of their occurrences within each time window is calculated. This information is then compiled into a matrix called the Actual Matrix. C ×N uniqueOP It is stored systematically in the form of a matrix.
[0154] The total number of rows in the actual matrix (N C) = total number of time windows for C = total number of occurrences of alarm event C in the log. Total number of columns in the actual matrix (N uniqueOP ) = number of unique operator actions in the log. N elements in the actual matrix ij represents the total number of occurrences of the jth unique operator action within the ith time window.
[0155] Generation of the time difference matrix: The time difference matrix is again generated by C ×N uniqueOP ) matrix that stores the time difference between an operator action within each time window of an alarm event C and a unique alarm event C. Elements D in the time difference matrix ij represents the time difference between the jth unique operator action and the unique alarm event C within the ith time window of the alarm event C.
[0156] If there are multiple occurrences of an associated operator action A within the same time window of an alarm event C, the occurrence of the operator action A closest to the notification of the alarm event C within that time window (i.e., the one with the smallest absolute time difference) is considered to calculate the time difference between the associated operator action A and the alarm event C. Figure 27 illustrates the time difference calculation between an associated operator action and a corresponding alarm event when there are two or more occurrences of the operator action within the same time window in accordance with the teachings of at least one exemplary embodiment.
[0157] Generation of the timestamp matrix: The timestamp matrix is again (N C ×N unique ) matrix, which stores the timestamp of the closest unique operator action to alarm event C within the time window of alarm event C.
[0158] In one embodiment, for two consecutive time windows, a check may be performed to determine whether the timestamp values of a given operator action in the two time windows are exactly the same. If the timestamp values of a given operator action in the two time windows are exactly the same, then both the actual matrix and the time difference matrix may be updated, where the update is as follows: In the time difference matrix, the entry for the time window with the smallest absolute time difference for that given operator action is kept, and other time window entries are set as null. In the actual matrix, the time window with the smallest absolute time difference for that given operator action is retained. Other time window entries are set as zero.
[0159] Generating a Binary Matrix: The next step for implementation involves generating a binary matrix from the actual matrix. To generate the binary matrix, all positive valued entries in the actual matrix are set as 1.
[0160] Generating forward probability: The forward probability of operator action A with respect to alarm event C indicates the probability of co-occurrence of operator action A with alarm event C, given that there is alarm event C. The forward probability of operator action A with respect to alarm event C can be calculated from the binary matrix of alarm event C as follows: Forward probability = (number of time windows of alarm event C in which operator action A exists) / (total number of time windows of alarm event C) (Equation 3)
[0161] Generating backward probability: The backward probability of operator action A with respect to alarm event C indicates the probability of operator action A co-occurring with alarm event C, given operator action A. The backward probability of operator action A with respect to alarm event C is Backward probability = (total number of instances of operator action A within all time windows of alarm event C) / (total number of instances of operator action A within the entire log) (Equation 4) It can be calculated as follows:
[0162] If both the forward probability value and backward probability value of operator action A with respect to alarm event C exceed a predefined threshold (e.g., 80%) set by the operator or user, then operator action A is considered to be strongly associated with alarm event C, and vice versa.
[0163] On the other hand, if for alarm event C there is no associated operator action with a forward probability value equal to or greater than a predefined threshold (e.g., 10%) set by the operator or user, alarm event C is considered an alarm without a consistent operator action.
[0164] Calculation of median time difference and median absolute deviation (MAD) from the median time difference: For all relevant operator actions of alarm event C that meet the predefined criteria of forward probability and backward probability, both the median time difference of the operator action from alarm event C and the median absolute deviation (MAD) from the median time difference are calculated based on the time difference matrix.
[0165] After identifying all relevant operator actions for alarm event C, the relevant actions are sorted in ascending order based on their median time difference from alarm event C. This provides the sequence of operator actions associated with the alarm, along with the median time difference between them.
[0166] FIG. 28 illustrates a server 2800 configured in accordance with the teachings of at least one example embodiment of the type that may be implemented within a process control system or alarm system.
[0167] The server 2800 may be part of a process control system or an alarm system and comprises one or more processor-implemented servers configured to implement one or more methods of at least one embodiment.
[0168] The server 2800 may include (i) an operator interface 2802 for allowing control instructions to be entered by an operator of the process control system or alarm system and for output data to be presented to the operator of the process control system or alarm system, (ii) a processor 2804 configured for data processing operations within the server 2800, (iii) a transceiver 2806 configured to enable transmission and reception of data network-based messages in the server 2800, and (iv) a memory 2808, which may include temporary memory and / or non-temporary memory.
[0169] In one embodiment, memory 2808 includes: (i) an operating system 2810 configured to manage device hardware and software resources and provide common services for software programs implemented in server 2800; (ii) a database interface 2812 configured to enable server 2800 to retrieve data from and store data in databases communicatively coupled to server 2800; (iii) a historical data acquisition controller 2814 configured to enable server 2800 to retrieve historical data, including one or more alarm data logs, event data logs, and / or alarm and event data logs, from one or more databases (e.g., for purposes of implementing the method of FIG. 5A ); (iv) a historical data parser 2816 configured to enable server 2800 to parse the retrieved historical data (e.g., for implementing one or more of the steps of the method of FIG. 5A ); and (v) a historical data processor 2818 configured to execute the historical data parsing (e.g., for purposes of implementing method steps 504A and 506A of FIG. 5A ). (vi) a chattering alarm identification controller 2818 configured for identifying chattering alarms or chattering alarm events based on the acquired historical data (e.g., for implementing one or more of the method steps of FIG. 6 ); (vi) a no-response alarm identification controller 2820 configured to enable the server 2800 (e.g., for implementing one or more of the method steps of FIG. 8 ) to identify, within the acquired historical data, alarm events that do not correspond to an operator response; (vii) a consistent response alarm identification controller 2822 configured to enable the server 2800 (e.g., for implementing one or more of the method steps of FIG. 9 ) to identify, within the acquired historical data, alarm events that consistently correspond to an operator response; and (viii) a redundant alarm identification controller 2824 configured to enable the server 2800 (e.g., for implementing one or more of the method steps of FIG. 11 or FIG. 12 ) to identify, within the acquired historical data, one or more clusters of redundant alarm events;(ix) a resulting alarm identification controller 2826 configured to enable the server 2800 (e.g., for implementing one or more of the method steps of FIG. 14 or FIG. 15) to identify, within the retrieved historical data, one or more clusters of resulting alarm events; (x) an alarm suppression controller 2828 configured to enable the server 2800 (e.g., for implementing one or more of the alarm event pattern responses associated with corresponding alarm event patterns, as described in any of the method steps of FIG. 5A, FIG. 5B, FIG. 8, and / or FIG. 11) to suppress one or more alarms; (xi) a SOP presentation controller 2830 configured to enable the server 2800 (e.g., for implementing one or more of the method steps of FIG. 9) to present a standardized operating procedure to an operator in response to detection of one or more detected alarm events or alarm event patterns; and (xii) an alarm prediction controller 2832 configured to enable the server 2800 (e.g., for implementing one or more method steps of FIG. 14) to present information to an operator that predicts the occurrence of one or more alarm events before said alarm events are detected or occur.
[0170] FIG. 29 illustrates an example computer system 2900 upon which various embodiments may be implemented.
[0171] System 2900 includes a computer system 2902, which in turn comprises one or more processors 2904 and at least one memory 2906. Processor 2904 is configured to execute program instructions and may be a real or virtual processor. It will be understood that computer system 2902 does not suggest any limitation as to the scope of use or functionality of the described embodiments. Computer system 2902 may include, but is not limited to, one or more of a general-purpose computer, a programmed microprocessor, a microcontroller, an integrated circuit, and other device or arrangement of devices capable of implementing steps constituting the method of at least one exemplary embodiment. Exemplary embodiments of computer system 2902 may include one or more servers, desktops, laptops, tablets, smartphones, mobile phones, mobile communication devices, tablets, phablets, and personal digital assistants. In one embodiment, memory 2906 may store software for implementing various embodiments. Computer system 2902 may have additional components. For example, computer system 2902 may include one or more communication channels 2908, one or more input devices 2910, one or more output devices 2912, and storage 2914. An interconnection mechanism (not shown), such as a bus, controller, or network, interconnects the components of computer system 2902. In various embodiments, operating system software (not shown) provides an operating environment for various software executing within computer system 2902 using processor 2904 and manages the various functionality of the components of computer system 2902.
[0172] The communication channel 2908 enables communication with various other computing entities over a communication medium that provides information, such as program instructions, or other data in the communication medium, including, but not limited to, wired or wireless methodologies implemented with electrical, optical, RF, infrared, acoustic, microwave, Bluetooth, or other communication media.
[0173] The input devices 2910 may include, but are not limited to, a touch screen, keyboard, mouse, pen, joystick, trackball, voice device, scanning device, or any other device capable of providing input to the computer system 2902. In one embodiment, the input devices 2910 may be a sound card or similar device that accepts audio input in analog or digital form. The output devices 2912 may include, but are not limited to, a CRT, LCD, LED display, or any other display-like user interface associated with a server, desktop, laptop, tablet, smartphone, cell phone, mobile communication device, tablet, phablet, and personal digital assistant, printer, speaker, CD / DVD writer, or any other device that provides output from the computer system 2902.
[0174] Storage 2914 may include, but is not limited to, magnetic disks, magnetic tapes, CD-ROMs, CD-RWs, DVDs, any type of computer memory, magnetic stripes, smart cards, printed bar codes, or any other transitory or non-transitory medium that can be used to store information and that can be accessed by computer system 2902. In various embodiments, storage 2914 may include program instructions for implementing any of the described embodiments.
[0175] In one embodiment, computer system 2902 is part of a distributed network or set of available cloud resources.
[0176] At least one exemplary embodiment can be implemented in numerous ways, including as a system, a method, or a computer program product such as a computer-readable storage medium or a computer network over which program instructions are communicated remotely.
[0177] The inventive concept may preferably be embodied as a computer program product for use with computer system 2902. The methods described herein are typically implemented as a computer program product including a set of program instructions executed by computer system 2902 or any other similar device. The set of program instructions may be stored on a tangible medium, such as a computer-readable storage medium (storage 2914), e.g., a diskette, CD-ROM, ROM, flash drive, or hard drive, or may be a series of computer-readable codes transmittable to computer system 2902 via any tangible medium, including, but not limited to, optical or analog communications channel 2908, via a modem or other interface device. The implementation of the inventive concept as a computer program product may also be intangible, using wireless techniques, including, but not limited to, microwave, infrared, Bluetooth, or other transmission techniques. These instructions may be preloaded onto the system, recorded on a storage medium such as a CD-ROM, or available for downloading over a network, such as the Internet or a cellular phone network. The series of computer-readable instructions may embody all or part of the functionality previously described herein.
[0178] 1-29 and the associated text above, it should be appreciated that the inventive concept provides a method including the steps of identifying an alarm event pattern in a log of alarm events occurring in a process control system, determining that a current alarm event in the process control system belongs to the alarm event pattern, determining one or more actions to resolve the current alarm event based on the alarm event pattern, and implementing the one or more actions to resolve the current alarm event. The alarm event pattern may be identified by identifying a first occurrence of a first alarm event in the log of alarm events, assigning a time window to the first occurrence of the first alarm event, identifying a first occurrence of a second alarm event that falls within the time window, performing one or more operations to calculate at least one probability value for the time window, and determining whether to associate the first and second alarm events with each other based on the at least one probability value calculated for the time window.
[0179] In one embodiment, the at least one probability value includes a forward probability value and a backward probability value. The first and second alarm events are determined to be associated with each other if the forward probability value exceeds a first threshold probability value and the backward probability value exceeds a second threshold probability value. As can be appreciated, the first threshold probability value may be the same as or different from the second threshold probability value.
[0180] In one embodiment, the one or more operations include an operation of determining a number of occurrences of a second alarm event within each time window of a plurality of time windows allocated to occurrences of the first alarm event, an operation of assigning a binary value to the number of occurrences of the second alarm event within each time window of the plurality of time windows, and an operation of calculating at least one probability value based on the binary value.
[0181] In one embodiment, the one or more operations include an operation of determining a first time difference between a first occurrence of a first alarm event and a first occurrence of a second alarm event within a time window; an operation of determining a second time difference between a second occurrence of the first alarm event and a second occurrence of the second alarm event within another time window allocated to the second occurrence of the first alarm event; an operation of obtaining a first timestamp of the first occurrence of the second alarm event within the time window; an operation of obtaining a second timestamp of the second occurrence of the second alarm event within another time window allocated to the second occurrence of the first alarm event; and an operation of calculating at least one probability based on at least one matrix created using the first timestamp and the second timestamp.
[0182] The method may include rendering on a display an ordered sequence of all associated alarms of the first alarm event occurring during the time window and another time window. The method may further include removing unwanted alarm events from the alarm event log before identifying the alarm event pattern. In one embodiment, the unwanted alarm events include one or more of chattering alarm events, recurring alarm events, or short-lived alarm events.
[0183] The step of determining one or more actions to resolve the current alarm event based on the alarm event pattern may include the steps of identifying an occurrence of a first alarm event in the alarm event log; assigning time windows to the occurrence of the first alarm event based on a notification time and a recovery time of the occurrence; identifying operator actions within each time window; performing one or more operations to calculate at least one probability value for the time window; and determining whether to associate the operator action and the first alarm event with each other based on the at least one probability value calculated for the time window.
[0184] In one embodiment, the at least one probability value includes a forward probability value and a backward probability value. The first alarm event and the operator action are determined to be associated with each other if the forward probability value exceeds a first threshold probability value and the backward probability value exceeds a second threshold probability value. As can be appreciated, the first threshold probability value may be the same as or different from the second threshold probability value.
[0185] The one or more operations may include an operation of determining a number of occurrences of the operator action within each time window, an operation of assigning a binary value to the number of occurrences of the operator action within each time window, and an operation of calculating at least one probability value based on the binary values.
[0186] The one or more operations may include an operation of determining a first time difference between a first occurrence of an operator action within a first time window and a notification of a first alarm event; an operation of determining a second time difference between a second occurrence of an operator action within a second time window and a second occurrence of the first alarm event; an operation of obtaining a first timestamp of the first occurrence of the operator action within the first time window; an operation of obtaining a second timestamp of the second occurrence of the operator action within the second time window; and an operation of calculating at least one probability based on at least one matrix created using the first timestamp and the second timestamp.
[0187] The method may include rendering on a display an ordered sequence of operator actions determined to be associated with the first alarm event.
[0188] At least one embodiment is directed to a device including a processing circuit configured to identify an alarm event pattern in a log of alarm events occurring within a process control system, determine that a current alarm event in the process control system belongs to the alarm event pattern, determine one or more operator actions for resolving the current alarm event based on the alarm event pattern, and generate control signals that cause the process control system to implement the one or more operator actions for resolving the current alarm event. The processing circuit can be configured to remove unwanted alarm events from the alarm event log before identifying the alarm event pattern. The unwanted alarm events can include one or more of channeling alarm events.
[0189] In one embodiment, the alarm event pattern includes an alarm event without operator action, an alarm event with consistent operator action, a group of redundant alarm events, a group of consequential alarm events, or a combination thereof.
[0190] At least one example embodiment is directed to a system comprising an output device and processing circuitry configured to identify an alarm event pattern in a log of alarm events occurring within the process control system, determine that a current alarm event in the process control system belongs to the alarm event pattern, determine one or more operator actions for resolving the current alarm event based on the alarm event pattern, and render an ordered sequence of operator actions for resolving the current alarm event to the output device.
[0191] Based on the above, it will be apparent that the inventive concepts offer significant advantages by providing efficient solutions for alarm rationalization, alarm prediction, and alarm handling, including, inter alia, intelligent alarm suppression, nuisance alarm elimination, alarm event pattern recognition, redundant alarm elimination or reduction, and predictive handling of alarm events that enable an operator to take preventative steps in connection with one or more alarm conditions even before an alarm condition is detected or an alarm event is triggered.
[0192] While exemplary embodiments have been described and illustrated herein, it will be understood that they are merely exemplary. It will be understood by those skilled in the art that various modifications in form and detail may be made therein without departing from or violating the spirit and scope of the inventive concepts as defined above and by the appended claims. In addition, the embodiments illustratively disclosed herein may preferably be practiced in the absence of any element not specifically disclosed herein, and in certain specifically contemplated embodiments, the inventive concepts are intended to be practiced in the absence of any one or more elements not specifically disclosed herein. [Explanation of symbols]
[0193] 100 Process Control Systems 102a Sensor, valve device, or actuator, sensor / valve device / actuator, sensor, valve device, and / or actuator, sensor, or device, or actuator 102b Sensor, valve device, or actuator, sensor / valve device / actuator, sensor, valve device, and / or actuator, sensor, or device, or actuator 102c Sensor, valve device, or actuator, sensor / valve device / actuator, sensor, valve device, and / or actuator, sensor, or device, or actuator 104a Controller 104b Controller 106 Server 108 databases 110 Operator Terminal 2800 Server 2802 Operator Interface 2804 processor 2806 transceiver 2808 memory 2810 Operating System 2812 Database Interface 2814 Historical Data Acquisition Controller 2816 Historical Data Parser 2818 Chattering Alarm Identification Controller 2820 Unresponsive Alarm Identification Controller 2822 Consistent Response Alarm Identification Controller 2824 Redundant Alarm Identification Controller 2826 Resulting Alarm Identification Controller 2828 Alarm Suppression Controller 2830 SOP Presentation Controller 2832 Alarm Predictive Controller 2900 Computer Systems, Systems 2902 Computer Systems 2904 processor 2906 memory 2908 Communication Channel 2910 Input Devices 2912 Output Device 2914 Storage
Claims
1. identifying alarm event patterns in a log of alarm events occurring within the process control system; determining that a current alarm event in the process control system belongs to the alarm event pattern; determining one or more actions to resolve the current alarm event based on the alarm event pattern; implementing the one or more actions to resolve the current alarm event; Including, The alarm event pattern is identifying a first occurrence of a first alarm event in the alarm event log; assigning a time window to the first occurrence of the first alarm event; identifying a first occurrence of a second alarm event within the time window; performing one or more operations to calculate at least one probability value for the time window, the one or more operations comprising: determining a number of occurrences of the second alarm event within each of a plurality of time windows allocated to occurrences of the first alarm event; assigning a binary value to the number of occurrences of the second alarm event within each time window of the plurality of time windows; calculating the at least one probability value based on the binary values; and determining whether the first and second alarm events are associated with each other as at least part of the alarm event pattern based on the at least one probability value calculated for the time window; 10. A computer-implemented method as identified by:
2. 2. The method of claim 1, wherein the at least one probability value includes a forward probability value and a backward probability value, and the first and second alarm events are determined to be associated with each other if the forward probability value exceeds a first threshold probability value and the backward probability value exceeds a second threshold probability value.
3. The method of claim 2 , wherein the first threshold probability value is the same as or different from the second threshold probability value.
4. The method of claim 1 , further comprising the step of removing unwanted alarm events from the alarm event log before identifying the alarm event pattern.
5. The method of claim 4 , wherein the unwanted alarm events include one or more of a chattering alarm event, a recurring alarm event, or a short-lived alarm event.
6. determining the one or more actions to resolve the current alarm event based on the alarm event pattern, identifying an occurrence of a third alarm event in the alarm event log; assigning a second time window to the occurrence of the third alarm event based on a notification time and a recovery time of the occurrence; identifying operator actions within each second time window; performing one or more second operations to calculate at least one second probability value for the second time window, the one or more second operations comprising: determining a number of occurrences of said operator action within each second time window; assigning a binary value to the number of occurrences of the operator action within each second time window; calculating the at least one second probability value based on the binary value; and determining whether to associate the operator action and the third alarm event with each other based on the at least one second probability value calculated for the second time window; 2. The method of claim 1, comprising:
7. 7. The method of claim 6, wherein the at least one second probability value includes a forward probability value and a backward probability value, and the third alarm event and the operator action are determined to be associated with each other if the forward probability value exceeds a first threshold probability value and the backward probability value exceeds a second threshold probability value.
8. The method of claim 7 , wherein the first threshold probability value is the same as or different from the second threshold probability value.
9. The method of claim 8, wherein determining the one or more actions to resolve the current alarm event based on the alarm event pattern comprises: identifying an occurrence of a third alarm event in the alarm event log; assigning a second time window to the occurrence of the third alarm event based on a notification time and a recovery time of the occurrence; identifying operator actions within each second time window; performing one or more second operations to calculate at least one second probability value for the second time window, the one or more second operations comprising: determining a first time difference between a first occurrence of an operator action within a second time window and notification of the third alarm event; determining a second time difference between a second occurrence of the operator action and a second occurrence of the third alarm event within another second time window; obtaining a first timestamp of the first occurrence of the operator action within the second time window; obtaining a second timestamp of the second occurrence of the operator action within the different second time window; calculating the at least one second probability value based on at least one matrix created using the first timestamp and the second timestamp; and determining whether to associate the operator action and the third alarm event with each other based on the at least one second probability value calculated for the second time window; 2. The method of claim 1, comprising:
10. 10. The method of claim 9, further comprising the step of rendering on a display the ordered sequence of the operator actions determined to be associated with the third alarm event.
11. An output device; Identifying alarm event patterns in a log of alarm events occurring within the process control system; determining that a current alarm event in the process control system belongs to the alarm event pattern; determining one or more operator actions for resolving the current alarm event based on the alarm event pattern; Rendering to the output device an ordered sequence of the operator actions that resolve the current alarm event. A processing circuit configured as follows: Equipped with The alarm event pattern is identifying a first occurrence of a first alarm event in the alarm event log; assigning a time window to the first occurrence of the first alarm event; identifying a first occurrence of a second alarm event within the time window; performing one or more operations to calculate at least one probability value for the time window, the one or more operations comprising: determining a number of occurrences of the second alarm event within each of a plurality of time windows allocated to occurrences of the first alarm event; assigning a binary value to the number of occurrences of the second alarm event within each time window of the plurality of time windows; calculating the at least one probability value based on the binary values; performing one or more operations, including: determining whether the first and second alarm events are associated with each other as at least part of the alarm event pattern based on the at least one probability value calculated for the time window; and Identified by the system.
Citation Information
Patent Citations
Alarm support apparatus and method for presenting measure when alarm is given
JP2010049519A
Plant monitoring control system
JP2013182547A
State monitoring apparatus, state monitoring system, and state monitoring method
JP2014203432A
Analysis system and analytical method
JP2017059152A
Process performance issues and alarm notification using data analytics
US20190129395A1