Universal reinforcement learning framework for process monitoring and anomaly / fault detection

By adopting reinforcement learning models in the process control system, training state-action mapping and utilizing metric-reward mapping, the problem of insufficient fault detection accuracy in the prior art is solved, and higher detection accuracy and factory safety are achieved.

CN120153376APending Publication Date: 2025-06-13FISHER ROSEMOUNT SYST INC
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202380076302.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-09-27
Filing Date
2023-09-25
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

Existing process control systems have insufficient accuracy in fault detection and abnormal detection, especially in case of changes in factory conditions and different factory types.

Method used

The reinforcement learning model is adopted to train state-action mapping, use metric-reward mapping to identify normal and abnormal states, and the Q table is dynamically updated to improve detection accuracy.

Benefits of technology

Improve the accuracy of abnormal/fault detection of process control systems, reduce the occurrence of correct and misunderstandings, and optimize the safety and environmental impact of the factory.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120153376A_ABST
    Figure CN120153376A_ABST
Patent Text Reader

Abstract

A method includes receiving a metric-award map; and training the state-action mapping using enhanced machine learning. A method includes receiving a set of metrics corresponding to an ongoing industrial control process; determining an abnormal / fault action value and a normal action value by referencing a state-action mapping determined by reinforcement learning; and causing a remedial action to occur. A process control system includes an anomaly / fault detection device that receives metrics, determines an anomaly / fault action value and a normal action value; and causing a remedial action to occur.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure generally relates to process plants and process control systems, and more particularly to using reinforcement learning and / or other machine learning techniques to achieve better anomaly / fault detection and / or other benefits within a process control system. Background Art

[0002] Distributed process control systems, such as those used in chemical, petroleum, industrial, or other process plants for manufacturing, refining, transforming, generating, or producing physical materials or products, typically include one or more process controllers communicatively coupled to one or more field devices via an analog bus, a digital bus, or a combined analog / digital bus or via a wireless communication link or network. These field devices, which can be, for example, valves, valve positioners, switches, transmitters, sensors, etc., are located within a process control environment and typically perform physical or process control functions such as opening or closing valves, measuring process parameters and / or environmental parameters (e.g., temperature or pressure), etc., to control one or more processes being performed within the process plant or system. Smart field devices, such as those compliant with well-known Fieldbus protocols, can also perform control calculations, alarm functions, and / or other control functions typically implemented within a controller.

[0003] Process controllers, which are typically also located within the plant environment, receive signals indicative of process measurements made by the field devices and / or other information regarding the field devices and execute controller applications that run, for example, different control modules. The control modules make process control decisions, generate control signals based on the received information, and coordinate with control modules or blocks implemented in field devices such as or Fieldbus field devices. The control modules implemented in the controller transmit control signals to the field devices via communication lines or links to control the operation of at least a portion of the process plant or system, e.g., to control at least a portion of one or more industrial processes operating or being performed within the plant or system. I / O devices, which are typically also located within the plant environment, are typically provided between the controller and one or more field devices to enable communication, e.g., by converting electrical signals to digital values and vice versa.

[0004] Information from field devices and controllers is typically provided via a communication network to one or more other hardware devices, such as operator workstations, personal computers or computing devices, data historians, report generators, centralized databases, or other centralized management computing devices that are typically located in a control room or other locations remote from the harsher field environments of the plant. These hardware devices run applications that can, for example, enable an operator to perform functions regarding controlling processes and / or operations and monitoring a process plant (e.g., changing settings of process control routines, modifying the operation of control modules within a controller or field device, viewing the current state of a process, viewing alarms generated by field devices and controllers, simulating the operation of a process for the purpose of training personnel or testing process control software, maintaining and updating a configuration database, etc.). The communication network utilized by the hardware devices, controllers, and field devices can include wired communication paths, wireless communication paths, or a combination of wired and wireless communication paths.

[0005] For example, the DeltaV TM control system sold by Emerson Automation Solutions includes multiple applications stored within and executed by different devices located at different positions within a process plant. A configuration application (which resides in one or more workstations or computing devices within the process control system or the plant's backend environment) enables a user to create or change process control modules and download these process control modules to dedicated distributed controllers via the communication network. Typically, these control modules are composed of communicatively interconnected function blocks, which are objects in an object-oriented programming protocol that perform functions within a control scheme based on inputs thereto and provide outputs to other function blocks within the control scheme. The configuration application may also allow a configuration designer to create or change an operator interface, view applications use this operator interface to display data to an operator and enable the operator to change settings within a process control routine, such as setpoints. Each dedicated controller and, in some cases, one or more field devices store and execute a corresponding controller application that runs the control modules allocated and downloaded thereto to implement actual process control functionality. A viewing application (which can execute on one or more operator workstations (or on one or more remote computing devices communicatively connected to the operator workstations and the communication network)) receives data from the controller application via the communication network and uses a user interface to display this data to a process control system designer, operator, or other user, and can provide any one of a plurality of different views, such as an operator's view, an engineer's view, a technician's view, etc. A data historian application typically stores the current process control routine configuration and associated data.

[0006] Process control systems used within process control plants generate, process, and store large amounts of data, in part due to the number of devices included in a typical plant installation (e.g., one thousand or more), and most of this data is continuous, streaming data. This data deluge has been complicated only by the recent trend of process control devices (e.g., wireless-enabled field devices and other devices) being connected at the time of initial installation or via retrofit installations.

[0007] Traditionally, process control systems have processed large amounts of data generated by process plants for many purposes, including for fault detection and / or anomaly detection. Processing this data to attempt to identify abnormal or fault conditions in the process plant can enable plant operators to preemptively avoid costly and potentially catastrophic consequences. For example, the DeltaV TM control system can include a continuous data analysis mode that detects process faults caused by changes in process variables or disturbances or actuator or sensor problems.

[0008] As discussed in “Hierarchical Distributed Monitoring or the Early Production of GasFlare Events,” Ind. Eng. Chem. Res. 2019, 58, 11352 - 11363 (which is hereby incorporated by reference in its entirety for all purposes), one use of such systems is to determine the necessary time to perform gas flaring. Gas flaring is an undesirable and unfortunately routine process involving the controlled combustion of waste gas. In some cases, gas flaring may be performed in a process plant to avoid overpressure events that could cause damage or injury. Because gas flaring involves undesirable environmental, economic, and regulatory consequences, process plant operators attempt to avoid gas flaring as much as possible. However, as discussed in that paper, there is a lack of monitoring strategies for providing early warning of potential flare events.

[0009] Traditionally, soft sensor models (e.g., principal component analysis (PCA)) can be used to assist process engineers in anomaly / fault detection. In such cases, predefined thresholds based on statistical calculations (e.g., T2 / Q limits) must be provided so that the control system can attempt to determine a fault or anomaly. However, a single statistical constant threshold not based on the historical knowledge of a given process may not be reliable. For example, such a constant may not distinguish between a process state change and an actual fault / anomaly. Additionally, PCA is not ideal, especially when there are process state changes or other conditions such as flow disturbances or state transitions. A single threshold is insufficient. Moreover, it is well known that PCA-based threshold anomaly / fault detection is plagued by false positives and thus by operator alert / decision fatigue. SUMMARY OF THE INVENTION

[0010] Techniques, systems, devices, components, equipment, and methods for improving the accuracy of anomaly / fault detection in a process control system are disclosed, even in the presence of varying plant conditions and potentially across different plant types. The techniques, systems, devices, components, equipment, and methods can be applied to industrial process control systems, environments, and / or plants (which may be interchangeably referred to herein as "industrial control," "process control," or "process" systems, environments, and / or "plants"), and can facilitate the development, modification, troubleshooting, stability, safety, and / or other aspects of such systems. Generally, such systems and plants control one or more processes in a distributed manner, and the one or more processes operate to manufacture, refine, transform, generate, or produce physical materials or products.

[0011] In one aspect of the present disclosure, a process control system includes a reinforcement learning model training aspect that receives a metric-reward mapping including one or more metrics, each metric corresponding to a respective reward. The mapping specifies a negative reward (i.e., penalty) associated with each metric such that the reinforcement learning model learns to associate certain metrics (and combinations of metrics, referred to herein as states) as being associated with a normal action state or an anomaly / fault state. During training, the reinforcement learning model can reside in a server remote from the process control system / plant and / or in a device (such as a field device) of the process control system / plant. Specifically, the training aspect can include using reinforcement learning techniques to process historical plant data (e.g., a time series of plant data corresponding to plant operations) to train a state-action mapping. In some aspects, the mapping can be a table or a trained machine learning model (e.g., an artificial neural network).

[0012] At each time step, the reinforcement learning model can calculate a net reward (i.e., the reward for a particular state, or the sum of individual rewards associated with multiple rewards) by cross-referencing a metric-reward mapping. The reinforcement learning model can then update the dynamic values of the state-action mapping. For example, the reinforcement learning model can update the Q-table for each unique combination of states such that the Q-table includes normal action values and abnormal / fault action values for each state. The Q-table can be updated over time using real-time plant data such that the Q-table action values are dynamically updated, enabling the abnormal / fault detection aspect of the present invention to continue learning over time. In some aspects, instead of updating Q-table values, the machine learning model can update the weights of an artificial neural network (e.g., a recurrent neural network).

[0013] Once the action-value mapping has been trained, time series plant data including each metric in the metrics can be fed into the reinforcement learning model, and the machine learning model can determine which unique state the in-field plant corresponds to and then determine whether the plant is in a normal condition or an abnormal / fault condition at that time step. These metrics can include suitable plant operating characteristics such as (i) predefined statistical thresholds, (ii) predefined constant control limits, (iii) setpoint indicators, (iv) load disturbance indicators, (v) operating level transition indicators, or (vi) combustion event indicators.

[0014] Time series plant data can be received from plant devices such as sensors, valves, transmitters, positioners, standard 4 mA to 20 mA devices, field devices, devices, Fieldbus devices, Profibus devices, DeviceNet devices, ControlNet devices, or Modbus devices.

[0015] When the reinforcement learning model determines that the plant is in an abnormal / fault state, the technology can include taking passive or proactive remedial actions. Generally speaking, the dynamic behavior of this technology exhibits significant improvements over traditional systems that do not learn over time and thus produce excessive false positives and false negatives.

[0016] In one aspect, a computer-implemented method for improving abnormal / fault detection and / or mitigation in a process control plant includes: (i) receiving a metric-reward mapping including one or more metrics, each metric corresponding to a respective reward; and (ii) using reinforcement machine learning to process a time series of historical plant data to train a state-action mapping, wherein the processing includes: for at least one time step in the time series of historical plant data, calculating a net reward corresponding to the time step by cross-referencing one or more metrics in the at least one time step with the metric-reward mapping.

[0017] In another aspect, a computer-implemented method for improving plant safety and environmental impact anomalies / failures includes: (i) receiving a set of metrics corresponding to an ongoing industrial control process; (ii) determining an anomaly / failure action value and a normal action value corresponding to the set of metrics by cross-referencing the set of metrics with reinforcement learning information; and (iii) causing a remedial action to occur based on the anomaly / failure action value.

[0018] In another aspect, a process control system includes: a plurality of process data generating devices including one or more field devices configured to generate data corresponding to an ongoing industrial control process implemented by the process control system; and an electronic network communicatively coupling at least some of the plurality of physical devices to an anomaly / failure detection device, where the anomaly / failure detection device includes a memory storing computer-executable instructions that, when executed by one or more processors of the anomaly / failure detection device, cause the anomaly / failure detection device to: (i) receive a set of metrics corresponding to the ongoing industrial control process; (ii) determine an anomaly / failure action value and a normal action value corresponding to the set of metrics by cross-referencing the set of metrics with reinforcement learning information; and (iii) cause a remedial action to occur based on the anomaly / failure action value.

[0019] In another aspect, one or more data generating devices are configured to: (i) receive reinforcement learning information; (ii) generate data corresponding to an ongoing industrial control process of a process control system; and (ii) process the generated data using the received reinforcement learning information. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] Figure 1 A block diagram depicting an exemplary process plant or process control system in accordance with some aspects is shown.

[0021] Figure 2 A block diagram depicting an exemplary process plant anomaly / failure detection computing environment in accordance with some aspects is shown.

[0022] Figure 3 An example of an exemplary anomaly / failure detection reinforcement learning pipeline in accordance with some aspects is shown.

[0023] Figure 4 A flowchart depicting an exemplary method for improving anomaly / failure detection and / or mitigation in a process control plant by using reinforcement learning to avoid false alarms in accordance with some aspects is shown.

[0024] Figure 5Depicts a flowchart illustrating an exemplary method according to some aspects for improving plant safety and environmental impact by processing reinforced learning information to determine plant anomalies / failures. Detailed Description

[0025] Figure 1 Depicts a block diagram of an example process control system 100 that can utilize one or more of the novel techniques described herein. Generally, the process control system 100 processes signals indicative of process measurements made by field devices to implement control routines and generates control signals that are transmitted via wired or wireless process control communication links or networks to other field devices to control the operation of physical processes in the process control system 100. Generally, at least one field device performs a physical function (e.g., opening or closing a valve, increasing or decreasing a temperature, making a measurement, sensing a condition, etc.) to control the physical process. Some types of field devices use I / O devices to communicate with other devices (e.g., controllers).

[0026] In Figure 1 the example, the process controller 111 is communicatively coupled to the wired field devices 115 - 122 via input / output (I / O) cards 126 and 128 and is communicatively coupled to the wireless field devices 140 - 146 via a wireless gateway 135 and a communication network 180.

[0027] The communication network 180 can include one or more wired and / or wireless communication links and can be implemented using any desired or suitable one or more communication protocols (such as, for example, Ethernet protocol). In some configurations (not shown) that include the controller 111, the controller 111 can be communicatively coupled to the wireless gateway 135 using one or more communication networks in addition to the network 180, such as by using any number of other wired communication links or wireless communication links that support one or more communication protocols (e.g., Wi-Fi or other IEEE 802.11 - compliant wireless local area network (WLAN) protocols, mobile communication protocols (e.g., WiMAX, LTE, or other ITU - R - compatible protocols), Profibus, Fieldbus, etc.).

[0028] One or more devices of the process control system 100 (which may include the controller 111) host control system services to implement a batch process or a continuous process (e.g., using at least some of the field devices 115 - 122 and 140 - 146). In one aspect, in addition to being communicatively coupled to the network 180, the controller 111 also uses, for example, standard 4 mA to 20 mA devices, I / O cards 126, 128, and / or any intelligent communication protocol (such as Fieldbus protocol, protocols, any desired hardware and software associated with the protocols, etc. are communicatively connected to at least some of the field devices 115 to 122 and 140 to 146. In Figure 1 the example, the controller 111, the field devices 115 to 122, and the I / O cards 126, 128 are wired devices, and the field devices 140 to 146 are wireless field devices. Of course, the wired field devices 115 to 122 and the wireless field devices 140 to 146 can conform to any other desired standards or protocols, such as any wired protocol or wireless protocol, including any standards or protocols developed in the future.

[0029] Figure 1 The process controller 111 in the example includes a processor 130 and a memory 132. The processor 130 is configured to communicate with the field devices 115 to 122 and 140 to 146 and with other nodes communicatively connected to the controller 111. The memory 132 (e.g., random access memory (RAM) and / or read-only memory (ROM)) can store computational memories (e.g., containers) executed by the processor 130 to provide certain control system services. Although not shown in Figure 1 it, Figure 1 any one or more of the other devices shown (e.g., the field devices 115 to 122 or 140 to 146, the wireless gateway 135, or any of the devices 117a, 117b, 117c, 118, or 112 discussed below) can also include a memory and a processor that enable these devices to similarly host one or more control system services.

[0030] In some aspects, the process control system 100 (i.e., the controller 111 and / or other devices) uses components commonly referred to as function blocks to implement control strategies, where each function block (via communication referred to as a link) operates together with other function blocks to implement a process control loop within the process control system 100. Function blocks based on control typically perform one of the following: an input function (e.g., associated with a transmitter, a sensor, or some other process parameter measurement device), a control function (e.g., associated with a control routine that performs PID, fuzzy logic, etc. control), or an output function that controls the operation of some device (e.g., a valve or a pump) to perform some physical function within the process control system 100. Of course, there are hybrid and other types of function blocks. The function blocks can be stored in the controller 111 and executed by the controller (this is typically the case when these function blocks are used with standard 4 mA to 20 mA devices and some types of smart field devices (such as devices) or are associated with them), or can be stored in the field device itself and implemented by it (this can be in the case of Fieldbus devices), and / or can be stored in and implemented by other devices of the process control system 100.

[0031] The wired field devices 115 to 122 can be any type of device, such as sensors, valves, transmitters, locators, etc., and the I / O cards 126 and 128 can be any type of I / O device that conforms to any desired communication or controller protocol. In Figure 1 , the field devices 115 to 118 are standard 4 mA to 20 mA devices or devices that communicate with the I / O card 126 via analog lines or a combination of analog and digital lines, while the field devices 119 to 122 are intelligent devices (such as Fieldbus field devices) that use the Fieldbus communication protocol to communicate with the I / O card 128 via a digital bus. However, in some aspects, at least some of the wired field devices among the wired field devices 115, 116, and 118 to 121 or at least one of the I / O cards 126, 128 use the network 180 and / or other suitable control system networks and protocols (e.g., Profibus, DeviceNet, Fieldbus, ControlNet, Modbus, etc.) to communicate with the controller 111 additionally or alternatively.

[0032] In Figure 1 the wireless field devices 140 to 146 use a wireless protocol (such as protocol) to communicate via the wireless process control communication network 170. Such wireless field devices 140 to 146 can directly communicate with one or more other devices or nodes of the wireless network 170, and the one or more other devices or nodes are also configured to communicate wirelessly (e.g., using a wireless protocol or another wireless protocol). To communicate with one or more other nodes that are not configured to communicate wirelessly, the wireless field devices 140 to 146 can utilize the wireless gateway 135 connected to the network 180 or connected to another process control communication network. The wireless gateway 135 provides access to various wireless devices 40 to 58 of the wireless communication network 170. Specifically, the wireless gateway 135 provides communication coupling between the wireless devices 40 to 58, the wired devices 15 to 28, and / or other nodes or devices of the process control system 100. For example, the wireless gateway 135 can provide communication coupling by using the network 180 and / or one or more other communication networks of the process control system 100.

[0033] Similar to the wired field devices 115 to 222, the wireless field devices 140 to 146 of the wireless network 170 perform physical control functions within the process control system 100, e.g., opening or closing valves or making measurements of process parameters. However, the wireless field devices 40 to 46 are configured to communicate using the wireless protocol of the network 170. Thus, the wireless field devices 40 to 46, the wireless gateway 135, and the other wireless nodes 52 to 58 of the wireless network 170 are producers and consumers of wireless communication packets.

[0034] In some configurations of the process control system 100, the wireless network 170 also includes non - wireless devices. For example, in Figure 1 the Figure 1 field device 148 is a traditional standard 4mA to 20mA device, and the field device 150 is a wired device. To communicate within the network 170, the field devices 148 and 150 are connected to the wireless communication network 170 via wireless adapters 152A, 152B. The wireless adapters 152A, 152B support wireless protocols (such as ) and may also support one or more other communication protocols (such as Fieldbus, Profibus, DeviceNet, etc.). Additionally, in some configurations, the wireless network 170 includes one or more network access points 155A, 155B, which may be separate physical devices that communicate with the wireless gateway 135 in a wired manner or may be provided as an integrated device with the wireless gateway 135. The wireless network 170 may also include one or more routers 158 for forwarding packets from one wireless device within the wireless communication network 170 to another wireless device. In Figure 1 the wireless devices 140 to 146 and 152 to 158 communicate with each other and with the wireless gateway 135 via the wireless links 160 of the wireless communication network 170 and / or via the network 180.

[0035] In Figure 1In this case, the process control system 100 includes one or more operator workstations or user interface devices 118, which are communicatively connected to the network 180. Via the operator workstation 118, an operator can view and monitor the runtime operation of the process control system 100 and take any diagnostic, corrective, maintenance, and / or other actions that may be required. At least some of the operator workstations in the operator workstation 118 may be located in various protected areas within or near the process control system 100, and in some cases, at least some of the operator workstations in the operator workstation 118 may be located remotely but are still communicatively connected to the process control system 100. The operator workstation 118 can be a wired or wireless computing device.

[0036] In some configurations, the process control system 100 includes one or more other wireless access points 117a that communicate with other devices using other wireless protocols (such as Wi-Fi or other IEEE 802.11-compliant WLAN protocols, mobile communication protocols (such as WiMAX (Worldwide Interoperability for Microwave Access), LTE (Long Term Evolution), or other ITU-R (International Telecommunication Union Radiocommunication Sector)-compatible protocols), short-wavelength radio communications (such as Near Field Communication (NFC) and protocols) or other suitable wireless communication protocols). Generally, the wireless access point 117a allows handheld or other portable computing devices to communicate through a corresponding wireless process control communication network that is different from the wireless network 170 and supports a wireless protocol different from the wireless network 170. For example, the wireless or portable user interface device 8 can be a mobile workstation or diagnostic test equipment utilized by an operator within the process control system 100. In some cases, in addition to the portable computing device, one or more process control devices (e.g., one or more of the controllers 111, field devices 115 to 122, and / or wireless devices 135, 140 to 158) also communicate using the wireless protocol supported by the access point 117a.

[0037] In some configurations, the process control system 100 includes one or more gateways 117b, 117c (also referred to herein as "edge gateways") to systems external to the process control system 100. Generally, such systems are associated with customers or suppliers of information generated or operated by the process control system 100. For example, the process control system 100 may include a gateway node 117b for communicatively connecting a process plant that includes the process control system 100 to another process plant. Additionally or alternatively, the process control system 100 may include a gateway node 117c for communicatively connecting the process control system 100 to external public or private systems such as a process control system of another supplier, a laboratory system (e.g., a laboratory information management system or LIMS), an operator loop database, a materials handling system, a maintenance management system, a product inventory control system, a production scheduling system, a weather data system, a shipping and disposal system, a packaging system, the Internet, and / or other external systems.

[0038] It should be noted that although Figure 1 illustrates a particular arrangement of a particular number of (particular types of) devices, this is merely illustrative and non-limiting. For example, the process control system 100 may omit the controller 111 as described above, or may include multiple controllers similar to the controller 111. As another example, the process control system 100 may include any number of wired and / or wireless field devices similar to the field devices 115 through 122 and / or 140 through 150, any number of other devices (e.g., devices 117a, 117b, 117c, 118, 112, 152a, 152b, 155a, 155b, etc.), and so on.

[0039] Furthermore, it should be noted that Figure 1 the process control system 100 may include a field environment (e.g., "process plant floor") and a backend environment (e.g., including the server 112), which are communicatively connected via a communication network 180. As Figure 1 shown, the field environment includes physical components (e.g., process control devices, networks, network elements, etc.) that are arranged, installed, and interconnected in the field environment to operate to control a process during runtime. For example, the controller 111, the I / O cards 126, 128, the field devices 115 through 122, and other devices and network components 135, 140 through 150, 152, 155, 158, and 170 are located, arranged, or otherwise included in the field environment of a plant that includes the process control system 100. Generally, in the field environment of the process control system 100, physical components provided in the field environment may be used to receive and process raw materials to produce one or more products.

[0040] The backend environment of a process plant that includes a process control system 100 includes various components that are shielded and / or protected from harsh conditions and materials in the field environment. For example, the backend environment can include an operator workstation 118, a server 112, and / or functionality that supports the runtime operation of the process control system 100. In some configurations, the various computing devices, databases, and other components and equipment included in the backend environment of a plant that includes the process control system 100 can be physically located in different physical locations, some of which can be local to the process plant and some of which can be remote.

[0041] Example process plant anomaly / fault detection computing environment

[0042] Figure 2 Depicts a block diagram of an example process plant anomaly / fault detection computing environment 200 that can be implemented, for example, in a process control system (such as Figure 1 process control system 100). Although the example process plant anomaly / fault detection computing environment 200 is described below with reference to the devices / components of the process control system 100, it should be understood that in some aspects, other systems and / or devices can alternatively implement the environment 200.

[0043] The example process plant anomaly / fault detection computing environment 200 includes an anomaly / fault detection device 202 and a plurality of process data generation devices 204-1 to 204-N, which can each perform corresponding data generation functions associated with the process control system. For example, in some aspects, the anomaly / fault detection device 202 does not need to be a dedicated device. As used herein, an "anomaly / fault detection device" can be a dedicated device or any other device that includes computer-executable instructions. For example, the anomaly / fault detection device can be a field device, a mobile computing device, a wearable device, etc.

[0044] In some aspects, the anomaly / fault detection device 202 can correspond to Figure 1 server 112. The data generation devices 204 can correspond to field devices 115 to 122, network components 135 to 170, etc. The process data generation devices 204 can include a plurality of devices integrated into the control of the physical process. The process data generation devices 204 can also or alternatively include a plurality of devices associated with the physical process in some other way.

[0045] The devices 204 can communicate with each other and / or with other components of the environment 200 via a computer network and / or other communication means (such as a direct API library, interprocess communication, intraprocess communication, remote procedure call, shared memory, etc.), and can include a wired and / or wireless (e.g., radio frequency) network or be layered thereon. For example, referring to Figure 1The process control system 100, for example, the device 204 can communicate via an electronic network 206, which can operate on top of the communication network 170 and / or network 180, independent of any protocols associated with those networks (e.g., IEEE 802.11 compliant WLAN, WiMAX, LTE, Profibus, Fieldbus, etc.), as discussed above.

[0046] The process data generation device 204 can provide and facilitate services within the process plant, such as monitoring services, diagnostic services, analysis services, and so on. For example, the data generation device 204 can facilitate operator console services, alarm management services, event management services, diagnostic services, remote access services, edge gateway services, input / output services, data historian services, external and / or peripheral input / output conversion services, key performance indicator services, data monitoring services, messaging services, security logic services, and / or any other suitable types of services related to the control system. Each of these services can generate service-specific data, which can be captured within the process plant anomaly / fault detection computing environment 200.

[0047] In some aspects, the anomaly / fault detection device 202 and the process data generation device 204 can be communicatively coupled via the electronic network 206 (e.g., Figure 1 the network 180 and / or network 170). In some aspects, the environment 200 may also include a process data database 208 (e.g., relational database, raw file database, transaction database, key-value store, structured query language (SQL) database, NoSQL database, flat file database, etc.). In some aspects, the database 208 can be stored in the device 202 and / or stored remotely.

[0048] The anomaly / fault detection device 202 can be any suitable computing device, such as a server, laptop computer, wearable device, mobile device, cloud computing instance, virtual machine, etc. The anomaly / fault detection device 202 can include one or more processors (e.g., one or more central processing units (CPUs), graphics processing units (GPUs), etc.) and one or more persistent and / or transient memories (e.g., one or more magnetic drives, one or more solid state drives, one or more random access memories, etc.) on which one or more sets of computer-executable instructions are stored. The one or more processors can access the one or more memories and the instructions stored thereon and execute the instructions stored thereon.

[0049] For example, in Figure 2In this case, the memory of the anomaly / fault detection device 202 may store the anomaly / fault detection application 230. The anomaly / fault detection application 230 may include multiple modules, each corresponding to a respective set of computer-executable instructions that, when executed by one or more processors of the anomaly / fault detection device 202, perform corresponding functions and operations.

[0050] In some aspects, the anomaly / fault detection application 230 includes a data processing module 232a, a machine learning training module 232b, a machine learning operation module 232c, and a process control remediation module 232d. Those of ordinary skill in the art will understand that, in some aspects, the module 232 may include more or fewer modules. For example, in some aspects, the machine learning training module 232b and the machine learning operation module 232c may be combined into a single module. Each of the modules 232 may be capable of communicating (e.g., exchanging information) with any of the other modules 232. Additionally, the anomaly / fault detection application 230 may include a set of database client binding instructions (not depicted) that enable any of the modules 232 to create, read, update, and / or delete information from one or more electronic databases (e.g., the process data database 208).

[0051] The data processing module 232a may include one or more sets of computer-executable instructions for performing data acquisition, data preprocessing, data buffering, and data storage. Specifically, the data processing module 232a may include instructions for receiving / retrieving process data from 4 mA to 20 mA standard devices, I / O cards 126, 128, and / or intelligent communication protocol devices (such as Fieldbus protocol, protocol, protocol, etc.). Process data generally includes data generated by various elements of the processing plant while the processing plant is operating. The data processing module 232a may include a software library that enables the anomaly / fault detection application 230 to read data from one or more devices located in a process control system (e.g., the process control system 100).

[0052] For example, the anomaly / fault detection application 230 may use one or more routines of the data processing module 232a to receive / retrieve data from one or more distributed devices in the process control system on a periodic basis, on a continuous or streaming (e.g., real-time) basis, on a batch basis (e.g., once or multiple times per day at predetermined intervals), and / or according to any other suitable frequency. In some aspects, the data processing module 232a may include instructions for reading time series data from one or more devices in the process control system and / or instructions for creating time series data based on data received from one or more devices of the process control system.

[0053] In some aspects, the data processing module 232a may include a thread pool or other parallel processing techniques that enable the data processing module to use distributed computing techniques to receive, process, buffer, and / or store data. In some aspects, multiple additional anomaly / fault detection devices 202 are included in the environment 200 to enable the data processing module 232a to scale horizontally via distributed computing. This technique can be particularly advantageous in situations where the data processing module 232a (such as in a production plant) processes large amounts of data and provides significant processing improvements.

[0054] In some aspects, the data processing module 232a may process data from the process plant and store the processed data (e.g., stored in an electronic database, the memory of the anomaly / fault detection device 202, etc.). At a later time, another module (e.g., the machine learning operation module 232c) may retrieve the stored processed data and further process the data. In some aspects, the data processing module 232a may operate in a pass-through mode, where the data processing module 232a does not store the data in a persistent form for later processing but directly provides the processed data to another module (e.g., the machine learning operation module 232c) directly (e.g., via shared / transient memory, network sockets, message queues, etc.). Those of ordinary skill in the art will understand that for various scenarios of data collection and processing, including several different and useful configurations in terms of distributed computing, are possible.

[0055] The machine learning training module 232b may include one or more sets of computer-executable instructions for creating, loading, training, and / or storing one or more machine learning models. The machine learning training module 232b may train any type of machine learning model, including supervised learning models, unsupervised learning models, and / or reinforcement learning models. In some aspects, the machine learning training module 232b may train one or more hybrid machine learning models, such as partially supervised reinforcement learning models. Examples of the types of machine learning modeling techniques that the machine learning training module 232b may use include, but are not limited to, principal component analysis, artificial neural networks, deep learning, regression, Markov decision processes, and the like.

[0056] The machine learning training module 232b may include libraries bound to client databases that enable the machine learning training module 232b to access an electronic database (e.g., the process data database 208). The machine learning training module 232b may retrieve / receive training data from the process data database 208 and / or from the data processing module 232a. The training data may correspond to historical data from a storage device (e.g., labeled historical data corresponding to the operation of a process plant) and / or updated data that has not yet been stored (e.g., process data from an operating process plant). For example, in some aspects, the labeled data may include historical knowledge of a process, which includes process stage changes, operating conditions, combustion events, and the like. The process data may be labeled according to a normal / abnormal state (e.g., according to whether the data corresponds to a combustion event) or a time period prior to a combustion event (e.g., T - 15 minutes, where T is the time of the combustion event).

[0057] New data that has not yet been stored may be referred to herein as streaming data, real-time data, or production data. The machine learning training module 232b may use historical and / or streaming data to train one or more machine learning models (e.g., reinforcement learning models). In some aspects, the machine learning training module 232b may use libraries bound to client databases to store information (such as hyperparameters / weights of a trained machine learning model, a serialized / pickled machine learning model, etc.) in an electronic database.

[0058] In some aspects, the machine learning training model may perform model-free training (e.g., Q-learning) by, for example, populating state-action values in a Q-table. In other aspects, the training may include training an artificial neural network (e.g., a recurrent neural network, an LTSM network, etc.) as a function approximator instead of a Q-table. Using a Q-table may advantageously simplify the programmer's programming task, while function approximation techniques may be appropriate if the state space is continuous and / or if there are a large number of states. For example, the machine learning training module 232b may perform training to populate as described below with respect to Figure 3The Q - table for a finite state space = |256| (i.e., a state space with a size / cardinality of 256) is discussed in further detail.

[0059] The machine learning operation module 232c may include one or more sets of computer - executable instructions for loading, parameterizing, deserializing, and / or operating one or more machine learning models (e.g., machine learning models trained and / or stored by the machine learning training module 232b). For example, the machine learning operation module 232c may load and operate one or more trained reinforcement learning models (e.g., Figure 3 the reinforcement learning model 300). The machine learning operation module 232c may include instructions for receiving / retrieving data (e.g., from the data processing module 232a and / or from the process data database 208), inputting the data into one or more machine learning models, and directing the output of these machine learning models. For example, the machine learning operation module 232c may include instructions for storing the output of one or more machine learning models in the memory of the anomaly / fault detection device 202 and / or the process data database 208.

[0060] As discussed below with respect to Figure 3 In some aspects, the machine learning operation module 232c may operate multiple machine learning models together and provide the output of the multiple models to another module (e.g., the process control remediation module 232d) for further processing / analysis.

[0061] The process control remediation module 232d may include one or more sets of computer - executable instructions for causing various remediation actions regarding the process plant in response to an input. Remediation actions may include proactive actions that directly modify the state of the process plant (e.g., actuating a valve, causing a chimney to release a gas flare, etc.) and / or passive actions that do not directly modify the state of the process plant (e.g., sounding an alarm, sending a notification, displaying a warning, etc.). Thus, the process control remediation module 232d may include computer - executable instructions for directly accessing one or more process control system devices (e.g., Figure 1 the field devices 115 to 122).

[0062] The input to the remediation module 232d may be the output of another module (e.g., the machine learning operation module 232c). For example, in some aspects, the process control remediation module 232d may receive one or more anomaly / fault indications as input and cause one or more remediation actions based on these indications. In some aspects, the process control remediation module 232d may not cause a remediation action. For example, the remediation module 232d may store one or more anomaly / fault indications in the process data database 208.

[0063] In an operation (e.g., for anomaly / fault detection), the data processing module 232a receives process control data (e.g., a time series of process control data) from one or more process data generating devices in the process data generating device 204. The data processing module 232a processes the process control data and / or stores the process control data (e.g., in the process data database 208). As discussed, the data processing module 232a may receive real-time data continuously and / or periodically. The machine learning operation module 232c receives / retrieves the processed process control data and inputs the process control data into a previously trained machine learning model (e.g., a reinforcement learning model trained using the machine learning training module 232b). The machine learning operation module 232c receives the output of the trained machine learning model and forwards the output to the process control remediation module 232d, and / or stores the output in the process data database 208. In response to determining that the output of the trained machine learning model received from the machine learning operation module 232c corresponds to an anomaly or fault condition (e.g., referring to the Q-table, as discussed herein), the process control remediation module 232d causes one or more proactive or passive actions regarding the process control system to occur.

[0064] In some aspects, in addition to or instead of forwarding the output of the trained machine learning model to the process control remediation module 232d, the machine learning operation module 232c may cause the trained machine learning model to be updated. For example, during an online training mode, the machine learning operation module 232c may cause a pre-trained (or untrained) reinforcement learning model to be trained based on the processed process control data. In some aspects, the data processing module 232a and / or the machine learning training module 232b may be omitted from the environment 200.

[0065] Exemplary machine learning-based dynamic threshold aspects

[0066] Turning to Figure 3 , according to some aspects, an anomaly / fault detection reinforcement learning pipeline 300 is depicted. The anomaly / fault detection reinforcement learning pipeline 300 includes a soft sensor model block 302, an anomaly / fault result block 304, and a reinforcement learning block 306. The soft sensor model may include a threshold-based statistical algorithm (e.g., principal component analysis). As discussed above, soft sensing for anomaly / fault detection has been used with static / fixed thresholds, but static thresholds are problematic for several reasons (e.g., the tendency for excessive false positives / negatives). In particular, for large amounts of process data, it has traditionally been difficult to define thresholds due to the presence of many potential triggering events. Thus, in the present technique, the reinforcement learning pipeline 300 is improved, for example, by adding reinforcement learning as will be described in more detail now.

[0067] In some aspects, the soft sensor model at block 302 is a DeltaV TM control system neural block. Pipeline 300 can use a pre-existing or traditional soft sensor model calculated, for example, by the DeltaV TM control system or another control system. For example, reinforcement learning at block 306 can be performed by one or more process data generation devices among the anomaly / fault detection device 202 and / or the process data generation device 204. For example, Figure 2 components of can exchange information with the soft sensor model via the electronic network 206 at block 302. Figure 2

[0068] At block 302, the soft sensor model includes a principal component analysis algorithm that performs process detection. For example, the soft sensor model can perform principal component analysis at block 302. In this case, given a new observation x t (1×m, where m is the number of parameters, and where x t represents process measurements collected at time t, such as temperature, flow rate, pressure, concentration, etc.), the principal component analysis performed at block 302 can include preprocessing the data based on the following formula:

[0069]

[0070] where and σ i are the mean and standard deviation of the i-th training parameter x i respectively, where

[0071]

[0072] and

[0073]

[0074] Two control charts can be used in process fault detection, namely Hotelling's T 2 and the squared prediction error (SPE) or Q statistic:

[0075]

[0076] where t t (1×a) is the scoring vector of x t (1×m), P(m×a) contains the loading vectors associated with the first a principal components, is the predicted value, e t (1×m) is the prediction residual, and λ i is the standard deviation of the scores of the i-th component.

[0077] At the reinforcement learning block 306, the agent 310 can learn based on performing actions and receiving rewards from the unknown process control plant environment 312. In some aspects, the environment 312 can correspond to a process control plant, such as Figure 1 the plant. In some aspects, the environment 312 can correspond to another non-process control plant environment. The agent 310 can include a policy 320 and a reinforcement learning algorithm 322. The agent 310 learns to take an action A t , which may affect the state of the environment 312. In response to the action A t , the environment 312 can generate a reward R t . The reinforcement learning algorithm 322 can process the reward R t and generate a policy update that is received and stored by the policy 320. Meanwhile, the agent 310 can receive one or more observations O t from the environment 312, and generate and save additional policy updates based on the observations O t . In some aspects, the agent 310 attempts to maximize the reward R t and / or the cumulative reward Q t .

[0078] It should be understood that in some cases, the best action is to do nothing. For example, this corresponds to Table 2 (action = 0). During the agent training process, the agent is learning when to label a state as "fault / anomaly" (action = 1). Doing nothing (action = 0) does not necessarily result in poor performance, especially when there is actually no fault / anomaly in the process.

[0079] The potential actions of the agent are typically described by the conditional probability π(a,o) = p(A t = a|O t = o) (i.e., the probability of action o given the observation o). For example, in some aspects, the reinforcement learning algorithm 322 defines A = {0,1}, where 1 means the state O t corresponds to an abnormal / fault state and 0 corresponds to a normal state.

[0080] The performance of the policy 320 can be measured based on the following formula:

[0081]

[0082] where d π (o) is the probability that the target system is in state o when using the policy mapping π, and Q(o,a) represents the expected value from the observation o with action a and Q(o t ,a t ) = Ε[R t+1 +γQ(o t+1 ,a t+1) Cumulative rewards starting from this point. Here, "E" represents "expectation", which is equal to calculating the average value of future rewards R t+1 and taking action a t+1 followed by the learning rate γ and the product of the future cumulative reward Q(o t+1 , a t+1 ).

[0083] d π (o) can be a constant that depends only on the number of states, that is, where N is the number of states.

[0084] The optimal policy mapping π satisfies the following equation:

[0085]

[0086] Since d π (p) is approximately the same for all states p, so if then π(o,a) can be defined as 1. In other words, the optimal policy mapping π * can be completely determined by the cumulative reward function Q(o,a).

[0087] The experience E can include a set of tuples, where each tuple is defined as <o,a,r,o′> and records all the behaviors of the policy mapping π (i.e., policy 320). The values o and o′ can respectively indicate the states of the target system (e.g., environment 312) before and after taking action a. In some aspects, r is the immediate reward obtained by using action a in state s. In an anomaly and fault detection system, actions can be decided based on the policy mapping π (i.e., policy 320). The reinforcement learning at block 306 can improve policy 320 by learning from the experience E. The experience E and the process control plant environment 312 can include real-time data and / or historical data. "E" represents experience, and the agent will learn from this experience to perform better in the fault detection task. Such experience can include new plant data with known operating states for training the agent.

[0088] Generally speaking, the reinforcement learning at block 306 can include training (e.g., through Figure 2 the machine learning training module 232b) to learn when the conditions in the environment 312 correspond to a fault or abnormal state, rather than representing a correct or incorrect recognition situation. In this way, the reinforcement learning at block 306 can include generating a dynamic threshold 330 that can be received by the soft sensor model and using it instead of the previous (e.g., static) threshold 310. Therefore, the agent 310 learns to take action A by interacting with the environment 312 t .

[0089] Advantages of the reinforcement learning pipeline 300 include: First, process engineers are provided with more options to evaluate the status of process operations and define normal and abnormal operations during training data preparation; and second, once training data is obtained and no assumptions are made about anomaly / fault detection (i.e., whether the data is outside control limits / standard deviation limits), the reinforcement learning pipeline 300 can continuously learn to utilize new experience E to improve its policy. Further, the present technology is capable of considering multiple criteria related to the process.

[0090] Traditionally, constant control limits and Q statistic values for T 2 are used for anomaly / fault detection and are calculated as follows:

[0091]

[0092]

[0093] where v and m are the sample mean and variance of the Q samples, and is the critical value of a chi-square variable with 2M 2 / v degrees of freedom at the significance level α.

[0094] When preparing the training data set (labeling normal data versus anomalies / faults), in addition to such control limits, historical knowledge of the process includes process stage changes and available operating conditions. For example, as discussed above, combustion events can be used to label process data as normal or abnormal.

[0095] Once the training data is prepared with the correct labels, the reinforcement learning (RL) framework for anomaly / fault detection will evolve with experience E to obtain a better estimate of the cumulative reward Q(o,a). Since the policy mapping π (i.e., Figure 3 policy 320 of ) does not make assumptions about anomaly / fault detection alone, it can consistently improve its ability with new experience E and dynamically improve anomaly / fault detection performance.

[0096] Those of ordinary skill in the art will understand that the present technology is flexible and contemplates many additional use cases. For example, the present technology can be used for predicting and detecting combustion events. The present technology can also be used to detect and diagnose suspicious valve behavior (e.g., regarding faulty valves). The present technology can be used in control applications with non-linearity (such as pH control systems and viscosity control systems). Further, the present technology can be used to perform measurements using cameras for learning normal and abnormal patterns in a safety context (e.g., network traffic patterns), as well as for measurement verification (e.g., to detect uncertain or incorrect measurement results).

[0097] Exemplary reinforcement learning aspects based on a metric-reward table and a Q-table

[0098] Table 1 depicts an exemplary metric-reward table that maps metrics to corresponding rewards according to some aspects.

[0099]

[0100]

[0101] Table 1 - Metric-Reward Table

[0102] Table 1 includes a metric column that includes multiple possible states, each state representing a specific metric related to the operation of a process plant.

[0103] In some aspects, Table 1 can be generated by Figure 2 the anomaly / fault detection application 230. For example, the machine learning training module 232b of the anomaly / fault detection application 230 can establish a database table schema in the process data database 208 that represents Table 1, and insert and update reward values into this database table during the training phase. Additionally, the machine learning operation module 232c can retrieve the reward values by issuing a SELECT query, such as one with a WHEREIN clause that includes values corresponding to the metrics. For example, the query "SELECT reward FROM metric_reward WHERE loadDisturbance = True" will return the value -1.

[0104] Similarly, with respect to Figure 3 , the reinforcement learning algorithm 322 can retrieve the reward R in a similar manner t . In some aspects, the reinforcement learning algorithm 322 can read the values in Table 1 into memory before operation to reduce network bandwidth usage.

[0105] The first four rows of Table 1 relate to evaluating the constant control limits for T 2 as discussed above. Table 1 associates the corresponding reward values with the corresponding values of T 2 . The next four rows of Table 1 relate to evaluating the Q statistic values discussed above, and associate the corresponding reward values with the corresponding Q values (i.e., perform mapping).

[0106] The next two rows of Table 1 hold the reward values for the application when there has been a setpoint change (α), and map the corresponding reward values to whether a setpoint change has occurred. The next two rows of the metric - reward Table 1 relate to evaluating whether there is a load perturbation (β), and map the corresponding reward values to whether a load perturbation has occurred. The next two rows of Table 1 hold the reward values for instances where there has been an operational - level transition (γ), and map the corresponding reward values to whether an operational - level transition has occurred. The next two rows of the metric - reward Table 1 hold the reward values for instances where there has been a gas combustion event (δ), and map the corresponding reward values to whether a gas combustion event has occurred.

[0107] Based on the size and possible values of the metric - reward Table 1, there can be 256 possible states. The reward for each state (o t ) can be the sum of the corresponding rewards for each individual metric in the metric - reward Table 1 for the set of metric values. For example, state o t can be defined as the set of where α = setpoint change; β = load perturbation; γ = operational - level transition; and δ = gas combustion event. In some aspects, the reward may be evaluated only when action = 1; that is, state o t is marked as abnormal / faulty. For example, in the state of ( no setpoint change; no load perturbation; no operational - level transition; no combustion event), the total reward is calculated as: - 1 - 1+0 + 0+0 + 0=-2.

[0108] In some aspects, the present technology may include using the historical process - plant data of o t for each of the 256 possible states (from index 0 to index 255) to train a fault - detection agent (e.g., Figure 3 agent 310) to populate a Q - table that can be used for (i.e., with respect to real - time process - plant data) future fault detection.

[0109] Table 2 depicts an exemplary Q - table according to some aspects.

[0110]

[0111] Table 2

[0112] Similar to Table 1, Table 2 can correspond to, for example, Figure 2The tables in the process data database 208. The machine learning training module 232b can issue UPDATE queries to the database 208 during training to update the values in Table 2. For example, "UPDATE q SET actionFault = 10.5, actionNormal = 0 where state = 1". Then, the machine learning operation module 232c can retrieve those stored values via "SELECT actionFault, actionNormal FROM q WHERE state = 1". Figure 3 The reinforcement learning algorithm 322 can perform similar SQL updates and select values.

[0113] Each row of Table 2 corresponds to a possible state, where each state includes an index, the corresponding abnormal limit (action = 1) value, and the corresponding normal limit (action = 0). Each state corresponds to a unique set of metrics in the metric-reward Table 1. An agent in the reinforcement learning model (e.g., Figure 3 the agent 310) can use Table 2 (i.e., the Q-table) to determine which state o the factory is in at a specific time t t , denoted as Q(o t ,a t ) = Ε[R t+1 + γQ(o t+1 ,a t+1 )], where "E" represents "expectation", which is equal to calculating the average value of the future reward R t+1 , the learning rate γ after taking the action a t+1 , and the product of the future cumulative reward Q(o t+1 ,a t+1 ), and based on this reward, select the action value from the Q-table (e.g., in some aspects, by choosing a higher reward). The action value (i.e., 1 or 0) can respectively correspond to whether the state of a given combination of values is in an abnormal / fault state or a normal state.

[0114] For example, consider state [0]: Q t / Q UCL ∈ (0,1], no setpoint change, no load disturbance, no operation level transfer, no combustion event. Since Q(o t ,a t ) is -35.5 for labeling the state as abnormal / fault, the agent will be more inclined to label the state as normal rather than because (Q = 0 > -35.5). In another example, consider state [1]: Q t / Q UCL∈(0,1], no setpoint change, no load disturbance, no operation level transfer, no combustion event. Since (Q = 10.5 > 0), in this case, the agent will mark the state as abnormal / faulty. In another example, consider state [2]: Q t / Q UCL ∈(0,1], no setpoint change, no load disturbance, no operation level transfer, no combustion event. In this example, the agent will mark the state as abnormal / faulty because Q = 20.5 > 0. In another example, consider state

[255] : Q t / Q UCL ∈(3,∞), setpoint change, load disturbance, operation level transfer, combustion event. Here, since Q(o t ,a t ) = 150 and 150 > 0, the agent will mark the state as abnormal / faulty.

[0115] Table 2 reflects the advantageous dynamic properties of the present technology. Consider state [1] again. Over time, training can adjust the values of the abnormal / faulty values and / or the normal values such that the normal value exceeds the abnormal / faulty value. In this case, state [1] will no longer be an abnormal / faulty state and will switch to a normal state. It should be understood that in some aspects, ties may occur. In such cases, the state can be resolved to have action 1 (i.e., abnormal / faulty). This dynamic behavior is contrary to traditional techniques that only allow the selection of static values. As Figure 3 shown, the Q-table values can correspond to dynamic thresholds (e.g., dynamic threshold 330) that can be learned over time to improve various aspects of abnormal / faulty detection (e.g., to avoid false positives). The user can provide feedback to the training to guide the system to better corresponding action values.

[0116] As discussed above, in some aspects, the present technology can include using one or more cameras in conjunction with reinforcement learning aspects. For example, process control system 100 can include one or more camera devices. One or more camera devices can measure aspects of the process plant (e.g., level in a tank, flare stack, furnace burner, etc.). For example, one or more cameras can include invisible spectrum sensors that measure the spectrum or the shape of a curve. In this case, the present technology can include, for example, training an artificial neural network and / or a convolutional neural network by Figure 2 the machine learning training module 232b to determine the abnormal / faulty state.

[0117] Example edge computing device

[0118] It should be understood that in some aspects, it may be advantageous to locate the trained machine learning model and associated data (e.g., the metric-reward table and / or the Q-table) deeper (i.e., at a lower level) within a process plant, such as one or more of the process data generation devices in process data generation device 204. Doing so can bring several advantages. First, the computational load can be distributed from the anomaly / fault detection device 202 to the process data generation devices 204. Second, by locating the trained machine learning model near the anomaly / fault detection device 202 where data is generated, the network resource consumption typically required to move data to the backend is eliminated. For large process plants, this can result in a significant release of network bandwidth. Third, by locally processing data at the process data generation devices 204, the present technique avoids the round trips that can add latency to the anomaly / fault detection and prediction techniques discussed herein.

[0119] In such aspects, one or more of the process data generation devices in process data generation device 204 can receive machine learning information (e.g., a trained machine learning model, an artificial neural network, a metric-reward table, a Q-table, etc.) from the anomaly / fault detection application 230. The anomaly / fault detection application 230 can include instructions for activating local processing in one or more of the process data generation devices 204. In this way, the anomaly / fault detection application 230 can configure the process data generation devices 204 to perform load balancing of process data. The technique can also advantageously enable the present technique to selectively allocate the load only to those devices that are capable of performing local processing (e.g., those process data generation devices 204 that have been upgraded and include the software and hardware necessary to execute the machine learning model and / or other custom instruction sets at the edge).

[0120] Exemplary method

[0121] Figure 4 A flowchart depicting an exemplary method 400 for improving anomaly / fault detection and / or mitigation in a process control plant by using reinforcement learning to avoid false alarms, according to some aspects, is shown. For example, method 400 can be performed by one or more physical devices of a process control system that implements a physical process, such as Figure 1 process control system 100). For example, method 400 can be implemented by the controller 111 and / or by one or more other physical devices in the process control system 100 (e.g., one or more of the operator workstations 108, servers 112, field devices 115 to 122 and / or 140 to 150, I / O devices 126 and / or 128, network devices 135, etc.). In some aspects, method 400 can use Figure 2by components of the environment 200 (e.g., the anomaly / fault detection device 202 and the multiple process data generation devices 204).

[0122] Method 400 may include: receiving a metric-reward mapping including one or more metrics, each metric corresponding to a respective reward (block 402). For example, method 400 may receive a metric-reward table (such as Table 1). Of course, in some aspects, there may be more or fewer metrics. The rewards may also be different, and the metrics themselves may also be different. For example, some types of processes may lack level transitions, and in such cases, the γ variable may be omitted. In some aspects, method 400 may receive metric-rewards in different formats (e.g., JSON or XML format). The table is chosen for convenience, but any suitable data structure / data format may be used to represent data in a computer.

[0123] Method 400 may include: using reinforcement machine learning to process historical plant data time series to train a state-action mapping (block 404). Specifically, as discussed with respect to Figure 3 this reinforcement machine learning may be performed at block 306, where the training attempts to maximize the reward. For example, method 400 may include: taking an action A that affects the state of the environment (e.g., a process plant) t ; receiving a reward R based on the action t (e.g., a reward associated with one or more metrics at the next time step); observing the reward R t ; and updating the policy to maximize the cumulative reward, where the reward R t is part of the cumulative reward.

[0124] In some aspects, the processing at block 404 includes: for at least one time step in the historical plant data time series, calculating a net reward corresponding to the time step by cross-referencing one or more metrics in the at least one time step with the metric-reward mapping. In other words, the reinforcement learning at block 404 may include: calculating the respective reward values for each of the metrics in Table 1, and then the values of each respective reward value may be summed to arrive at the net reward. As discussed above, for the state of ( no setpoint change; no load perturbation; no operating level transition; no combustion event), the respective reward is equal to (-1 - 1 + 0 + 0 + 0 + 0), and the net reward is equal to -2.

[0125] In some aspects, method 400 may include: generating a Q-table including multiple states, each state having a respective anomaly / fault action value and a respective normal action value. Specifically, in some aspects, a side effect or result of the training at block 404 may be a Q-table (such as Table 2 above).

[0126] As in the examples of Tables 1 and 2, the cardinality of the Q-table (i.e., |Q-table|) can be determined by multiplying together the number of possible metric states using the product rule. Thus, since there are 4 possible T2 states, 4 possible Q states, 2 possible setpoint states, 2 possible load disturbance states, 2 possible operating level transition states, and 2 possible combustion states, the cardinality of the Q-table in this example is calculated as 4 × 4 × 2 × 2 × 2 × 2, which equals the cardinality |256|. Thus, in some aspects, the size of the Q-table can be 256.

[0127] As noted, when the Q-table becomes larger (e.g., due to a larger number of metrics, continuous metrics, etc.), function approximation approaches (e.g., using artificial neural networks such as recurrent neural networks) may be more appropriate. In such cases, method 400 may include: training an artificial neural network to act as a function approximator such that the artificial neural network processes state inputs and directly predicts (i) anomaly / fault action values and (ii) normal action values; or a boolean action value corresponding to whether the state is an anomaly / fault state.

[0128] In some aspects, the metrics of method 400 may include: i) predefined statistical thresholds, ii) predefined constant control limits, iii) setpoint indicators, iv) load disturbance indicators, v) operating level transition indicators, or vi) combustion event indicators. Reward values are typically represented as integers (e.g., negative or positive integers), where positive values are associated with faults / anomalies and negative values are associated with non-fault / non-anomaly behavior. For example, a metric indicating a human-initiated and thus potentially intentional action (e.g., a setpoint change) has a negative reward value, which effectively teaches the agent that such data points should not be classified as faults / anomalies (and thus should not result in passive or active mitigation / remediation). Similarly, given that an event associated with such a metric is considered to always be a significant and error-free indicator of a fault / anomaly, the reward value associated with the combustion metric can be relatively large. Of course, those of ordinary skill in the art will understand that, in some aspects, the sign of the integer reward can be easily reversed.

[0129] In some aspects, the predefined statistical threshold and the predefined constant control limit of method 400 are defined as

[0130]

[0131] and

[0132]

[0133] where T 2 is the Hotelling T-squared distribution and Q is the squared prediction error (SPE) statistic value;

[0134] where ν and m are the sample mean and variance of the Q samples, respectively; and

[0135] where is the critical value of a chi - square variable with 2m 2 / v degrees of freedom at the significance level α.

[0136] In some aspects, each time step in the historical plant data time series is labeled as corresponding to one of i) a fault state or ii) a normal state. In this way, the training process of method 400 can learn to associate certain states with abnormal / faulty actions or normal actions. As discussed above, method 400 can include: pre - processing the historical plant data time series. This can be performed, for example, when the historical plant data includes data generated by heterogeneous devices that must be reconciled or combined before processing. Method 400 can include: storing some of the historical plant data in the historical plant data in, for example, a memory (such as the memory of the anomaly / fault detection device 202, the process data database 208, etc.).

[0137] Figure 5 depicts a flowchart representing an exemplary method 500 for improving plant safety and environmental impact by processing reinforcement learning information to determine plant anomalies / faults, according to some aspects. For example, method 500 can be performed by one or more physical devices of a process control system that implements a physical process (such as Figure 1 process control system 100). For example, method 500 can be implemented by the controller 111 and / or by one or more other physical devices in the process control system 100 (such as one or more of the operator workstations 108, servers 112, field devices 115 to 122 and / or 140 to 150, I / O devices 126 and / or 128, network devices 135, etc.). In some aspects, method 500 can use Figure 2 components of the environment 200 (such as the anomaly / fault detection device 202 and the plurality of process data generating devices 204) to perform.

[0138] Method 500 can include: receiving a set of metrics corresponding to an ongoing industrial control process (block 502). For example, the set of metrics can be included in a set of time - series data corresponding to the operation of a process plant. These metrics can include at least one of the following: i) a predefined statistical threshold, ii) a predefined constant control limit, iii) a set - point indicator, iv) a load disturbance indicator, v) an operating - level transition indicator, or vi) a combustion event indicator. In some aspects, method 500 can include: by pre - processing data from one or more devices in a process plant (such as sensors, valves, transmitters, locators, standard 4 mA to 20 mA devices, field devices, devices, generate a set of metrics from process control data generated by a Fieldbus device, a Profibus device, a DeviceNet device, a ControlNet device, and / or a Modbus device).

[0139] Method 500 may include: determining an abnormal / fault action value and a normal action value corresponding to the set of metrics by cross-referencing the set of metrics with the reinforcement learning information (block 504). In some aspects, the reinforcement learning information in method 500 may include a reinforcement learning Q-table (e.g., Table 2 above). In some aspects, the reinforcement learning information in method 500 may include the function approximation output of a trained artificial neural network (e.g., a recurrent neural network), as described herein.

[0140] Method 500 may include: causing a remedial action to occur based on the abnormal / fault action value (block 506). In some aspects, the remedial action of method 500 is a passive remedial action such that the action does not directly cause any physical change to the process plant. For example, such passive or indirect remedies may include (i) sounding an alarm, (ii) sending a notification, and / or (iii) displaying a warning.

[0141] In some aspects, the remedial action of method 500 may be an active remedial action, such as (i) actuating a valve, (ii) causing a chimney to release a gas flare, (iii) performing an action on the plant, etc.

[0142] Aspects of the techniques described in this disclosure may individually or in combination include any number of the following aspects:

[0143] 1. A computer-implemented method for improving abnormal / fault detection and / or mitigation in a process control plant, the computer-implemented method comprising: receiving a metric-reward mapping including one or more metrics, each metric corresponding to a respective reward; and using reinforcement machine learning to process a historical plant data time series to train a state-action mapping, wherein the processing includes: for at least one time step in the historical plant data time series, calculating a net reward corresponding to the time step by cross-referencing one or more metrics in the at least one time step with the metric-reward mapping.

[0144] 2. The computer-implemented method according to aspect 1, wherein using reinforcement machine learning to process the historical plant data time series to train the state-action mapping includes: generating a Q-table including a plurality of states, each state having a respective abnormal / fault action value and a respective normal action value.

[0145] 3. The computer-implemented method according to any one of aspect 2, wherein the cardinality of the Q-table is defined via a product rule regarding the number of possible different rewards for each of the metrics.

[0146] 4. The computer-implemented method according to any one of aspect 3, wherein the size of the Q-table is 256.

[0147] 5. The computer-implemented method according to aspect 4, wherein using reinforcement machine learning to process the historical plant data time series to train the state-action mapping includes: training an artificial neural network to act as a function approximator for classifying an input into one of (i) an abnormal / fault action value and (ii) a normal action value.

[0148] 6. The computer-implemented method according to aspect 5, wherein the artificial neural network is a recurrent neural network.

[0149] 7. The computer-implemented method according to aspect 1, wherein the metrics include at least one of the following: (i) a predefined statistical threshold, (ii) a predefined constant control limit, (iii) a setpoint indicator, (iv) a load disturbance indicator, (v) an operating level transition indicator, or (vi) a combustion event indicator.

[0150] 8. The computer-implemented method according to aspect 7, wherein the predefined statistical threshold and the predefined constant control limit are respectively defined as

[0151]

[0152] and

[0153]

[0154] where T 2 is the Hotelling T-squared distribution and Q is the squared prediction error (SPE) statistic value;

[0155] where v and m are respectively the sample mean and variance of the Q samples; and

[0156] where is the critical value of a chi-square variable with 2m 2 / v degrees of freedom at the significance level α.

[0157] 9. The computer-implemented method according to aspect 1, wherein each respective reward value is represented as an integer.

[0158] 10. The computer-implemented method according to aspect 1, wherein each time step in the historical plant data time series is labeled as corresponding to one of i) a fault state or ii) a normal state.

[0159] 11. The computer-implemented method according to any one of aspects 1 to 10, the computer-implemented method further comprising: preprocessing the historical plant data time series.

[0160] 12. The computer-implemented method according to any one of aspects 1 to 11, the computer-implemented method further comprising: storing at least some of the historical plant data time series in the historical plant data time series in an electronic database.

[0161] 13. The computer-implemented method according to aspect 1, wherein using reinforcement machine learning to process the historical plant data time series to train the state-action mapping includes: taking an action A that affects the state of the environment t ; receiving a reward R based on the action t ; observing the reward R t ; and updating the policy to maximize the cumulative reward, the reward R t being part of the cumulative reward.

[0162] 14. A computer-implemented method for improving plant safety and environmental impact anomalies / failures, the method comprising: receiving a set of metrics corresponding to an ongoing industrial control process; determining anomaly / failure action values and normal action values corresponding to the set of metrics by cross-referencing the set of metrics with reinforcement learning information; and causing a remedial action to occur based on the anomaly / failure action values.

[0163] 15. The computer-implemented method according to aspect 14, wherein the set of metrics corresponding to the ongoing industrial control process includes at least one of the following: (i) a predefined statistical threshold, (ii) a predefined constant control limit, (iii) a set point indicator, (iv) a load disturbance indicator, (v) an operating level transition indicator, or (vi) a combustion event indicator.

[0164] 16. The computer-implemented method according to any one of aspects 14 to 15, the computer-implemented method further comprising: generating the set of metrics by preprocessing process control data generated by one or more devices in a process plant.

[0165] 17. The computer-implemented method according to aspect 16, wherein the one or more devices in the process plant include at least one of the following: a sensor, a valve, a transmitter, a locator, a standard 4 mA to 20 mA device, a field device, a device, a Fieldbus device, a Profibus device, a DeviceNet device, a ControlNet device, or a Modbus device.

[0166] 18. The method according to aspect 14, wherein the reinforcement learning information includes a reinforcement learning Q-table.

[0167] 19. The method according to aspect 14, wherein the reinforcement learning information includes the function approximation output of a trained artificial neural network.

[0168] 20. The method according to aspect 19, wherein the trained artificial neural network is a recurrent neural network.

[0169] 21. The computer-implemented method according to aspect 14, wherein the remedial action is a passive remedial action.

[0170] 22. The computer-implemented method according to aspect 21, wherein the passive remedial action includes at least one of the following: (i) sounding an alarm, (ii) sending a notification, or (iii) displaying a warning.

[0171] 23. The computer-implemented method according to aspect 14, wherein the remedial action is an active remedial action.

[0172] 24. The computer-implemented method according to aspect 23, wherein the active remedial action includes at least one of the following: (i) actuating a valve, (ii) causing a chimney to release a gas flare, or (iii) performing an action on the plant.

[0173] 25. A process control system, the process control system comprising: a plurality of process data generating devices, the plurality of process data generating devices including one or more field devices configured to generate data corresponding to an ongoing industrial control process implemented by the process control system; and an electronic network that communicatively couples at least some of the plurality of process data generating devices to an anomaly / fault detection device, wherein the anomaly / fault detection device includes a memory storing computer-executable instructions that, when executed by one or more processors of the anomaly / fault detection device, cause the anomaly / fault detection device to perform the following operations: receive a set of metrics corresponding to the ongoing industrial control process; determine an anomaly / fault action value and a normal action value corresponding to the set of metrics by cross-referencing the set of metrics with reinforcement learning information; and cause a remedial action to occur based on the anomaly / fault action value.

[0174] 26. The process control system according to aspect 25, wherein the set of metrics corresponding to the ongoing industrial control process includes at least one of the following: (i) a predefined statistical threshold, (ii) a predefined constant control limit, (iii) a setpoint indicator, (iv) a load disturbance indicator, (v) an operating level transition indicator, or (vi) a combustion event indicator.

[0175] 27. The process control system according to any one of aspects 25 to 26, wherein the anomaly / fault detection device includes a memory storing computer-executable instructions that, when executed by one or more processors of the anomaly / fault detection device, cause the anomaly / fault detection device to perform the following operations: generate the set of metrics by preprocessing process control data generated by the plurality of process data generating devices.

[0176] 28. The process control system according to aspect 25, wherein the plurality of process data generating devices in the process plant includes at least one of the following: sensors, valves, transmitters, positioners, standard 4 mA to 20 mA devices, field devices, devices, Fieldbus devices, Profibus devices, DeviceNet devices, ControlNet devices, or Modbus devices.

[0177] 29. The process control system according to aspect 25, wherein the reinforcement learning information includes a reinforcement learning Q-table.

[0178] 30. The process control system according to aspect 25, wherein the reinforcement learning information includes a function approximation output of a trained artificial neural network.

[0179] 31. The process control system according to aspect 30, wherein the trained artificial neural network is a recurrent neural network.

[0180] 32. The process control system according to aspect 25, wherein the remedial action is a passive remedial action.

[0181] 33. The process control system according to aspect 32, wherein the passive remedial action includes at least one of the following: (i) sounding an alarm, (ii) sending a notification, or (iii) displaying a warning.

[0182] 34. The process control system according to aspect 25, wherein the remedial action is an active remedial action.

[0183] 35. The process control system according to aspect 34, wherein the active remedial action includes at least one of the following: (i) actuating a valve, or (ii) causing a chimney to release a gas flare, or (iii) performing an action on the plant.

[0184] 36. One or more data generation devices configured to: receive reinforcement learning information; generate data corresponding to an ongoing industrial control process of a process control system; and process the generated data using the received reinforcement learning information.

[0185] 37. The one or more data generation devices according to aspect 36, wherein each data generation device among the data generation devices is at least one of the following: a sensor, a valve, a transmitter, a locator, a standard 4 mA to 20 mA device, a field device, a device, a Fieldbus device, a Profibus device, a DeviceNet device, a ControlNet device, or a Modbus device.

[0186] 38. The one or more data generation devices according to aspect 36, wherein the device is further configured to: generate the data corresponding to the ongoing industrial control process of the process control system in response to an activation instruction received from a remote anomaly / fault detection application.

[0187] 39. The one or more data generation devices according to aspect 36, wherein the received reinforcement learning information includes a metric-reward and a Q-table.

[0188] 40. The one or more data generation devices according to aspect 36, wherein the reinforcement learning information includes the function approximation output of a trained artificial neural network.

[0189] 41. The one or more data generation devices according to aspect 40, wherein the trained artificial neural network is a recurrent neural network.

[0190] 42. The one or more data generation devices according to any one of aspects 36 to 41, wherein the device is further configured to: calculate a set of metrics corresponding to the ongoing industrial process; determine abnormal / fault action values and normal action values corresponding to the set of metrics by cross-referencing the set of metrics with the reinforcement learning information; and cause a remedial action to occur based on the abnormal / fault action values.

[0191] 43. The one or more data generation devices according to aspect 42, wherein the remedial action is a passive remedial action.

[0192] 44. The one or more data generation devices according to aspect 43, wherein the passive remedial action includes at least one of the following: (i) sounding an alarm, (ii) sending a notification, or (iii) displaying a warning.

[0193] 45. The one or more data generation devices according to aspect 44, wherein the remedial action is an active remedial action.

[0194] 46. The one or more data generation devices according to aspect 45, wherein the active remedial action includes at least one of the following: (i) actuating a valve, or (ii) causing a chimney to release a gas flare, or (iii) performing an action on the plant.

[0195] 47. The one or more data generation devices according to aspect 42, wherein the remedial action includes causing at least one of the one or more data generation devices to stop generating the data corresponding to the ongoing industrial control process of the process control system.

[0196] 48. The one or more data generation devices according to any one of aspects 36 to 47, the one or more data generation devices are further configured to: send the generated data to one or both of (i) another data generation device and (iii) a remote abnormal / fault detection device.

[0197] 49. The one or more data generation devices according to any one of aspects 36 to 48, the one or more data generation devices are further configured to: receive updated reinforcement learning information in response to the sending.

[0198] When implemented in software, any of the applications, services, and engines described herein can be stored in any tangible non-transitory computer-readable memory, such as on a disk, a laserdisc, a solid-state memory device, a molecular memory storage device, or other storage media, in the RAM or ROM of a computer or processor, and the like. Although the example systems disclosed herein are disclosed as including software and / or firmware and other components that execute on hardware, it should be noted that such systems are merely illustrative and should not be considered restrictive. For example, it is contemplated that any one or all of these hardware, software, and firmware components can be embodied specifically in hardware, specifically in software, or in any combination of hardware and software. Thus, while the example systems described herein are described as being implemented in software that executes on the processors of one or more computer devices, those of ordinary skill in the art will readily understand that the examples provided are not the only way to implement such systems.

[0199] Accordingly, while the invention has been described with reference to specific examples, the specific examples are intended to be illustrative only and not limiting of the invention, and it will be apparent to those of ordinary skill in the art that changes, additions, or deletions can be made to the disclosed aspects without departing from the spirit and scope of the invention.

Claims

1. A computer-implemented method for improving anomaly / fault detection and / or mitigation in a process control plant, the computer-implemented method comprises: receiving a metric-reward mapping including one or more metrics, each metric corresponding to a respective reward; and using reinforcement machine learning to process a historical plant data time series to train a state-action mapping, wherein the processing includes: for at least one time step in the historical plant data time series, calculating a net reward corresponding to the time step by cross-referencing one or more metrics in the at least one time step with the metric-reward mapping.

2. The computer-implemented method according to claim 1, wherein reinforcement machine learning is used to process the historical plant data time series to train the state-action mapping comprises: generating a Q-table including a plurality of states, each state having a respective anomaly / fault action value and a respective normal action value.

3. The computer-implemented method according to claim 2, wherein the cardinality of the Q-table is defined via a product rule regarding the number of possible different rewards for each of the metrics.

4. The computer-implemented method according to claim 3, wherein the size of the Q-table is 256.

5. The computer-implemented method according to claim 1, wherein reinforcement machine learning is used to process the historical plant data time series to train the state-action mapping comprises: training an artificial neural network to act as a function approximator for classifying an input as corresponding to one of (i) an anomaly / fault action value and (ii) a normal action value.

6. The computer-implemented method according to claim 5, wherein the artificial neural network is a recurrent neural network.

7. The computer-implemented method according to claim 1, wherein the metrics include at least one of the following: (i) a predefined statistical threshold, (ii) a predefined constant control limit, (iii) a setpoint indicator, (iv) a load disturbance indicator, (v) an operating level transition indicator; or (vi) a combustion event indicator.

8. The computer-implemented method according to claim 7, wherein the predefined statistical threshold and the predefined constant control limit are respectively defined as and where T 2 is the Hotelling's T-squared distribution and Q is the squared prediction error (SPE) statistic; where ν and m are respectively the sample mean and variance of the Q samples; and wherein is the critical value of a chi-square variable with 2m 2 / v degrees of freedom at the significance level α.

9. The computer-implemented method according to claim 1, wherein each respective reward value is represented as an integer.

10. The computer-implemented method according to claim 1, wherein each time step in the historical plant data time series is labeled as corresponding to one of i) a fault state or ii) a normal state.

11. The computer-implemented method according to claim 1, the computer-implemented method further comprises: preprocessing the historical plant data time series.

12. The computer-implemented method according to claim 1, the computer-implemented method further comprises: storing at least some of the historical plant data time series in an electronic database.

13. The computer-implemented method according to claim 1, wherein reinforcement machine learning is used to process the historical plant data time series to train the state-action mapping comprising: Take action A that affects the state of the environment t ; Receive a reward R based on the action t ; Observe the reward R t ; and Update the strategy to maximize the cumulative reward, where the reward R t is part of the cumulative reward.

14. A computer-implemented method for improving plant safety and environmental impact anomalies / failures, the method comprising: receiving a set of metrics corresponding to an ongoing industrial control process; determining anomaly / failure action values and normal action values corresponding to the set of metrics by cross-referencing the set of metrics with reinforcement learning information; and causing a remedial action to occur based on the anomaly / failure action values.

15. The computer-implemented method according to claim 14, wherein the set of metrics corresponding to the ongoing industrial control process comprises at least one of the following: (i) predefined statistical thresholds, (ii) predefined constant control limits, (iii) setpoint indicators, (iv) load disturbance indicators, (v) operation-level transition indicators; or (vi) combustion event indicators.

16. The computer-implemented method according to claim 14, the computer-implemented method further comprising: generating the set of metrics by preprocessing process control data generated by one or more devices in a process plant.

17. The computer-implemented method according to claim 16, wherein the one or more devices in the process plant comprise at least one of the following: sensors, valves, transmitters, positioners, standard 4 mA to 20 mA devices, field devices, Device Fieldbus device, Profibus devices, DeviceNet devices, ControlNet devices; or Modbus devices.

18. The method according to claim 14, wherein the reinforcement learning information comprises a reinforcement learning Q-table.

19. The method according to claim 14, wherein the reinforcement learning information comprises the function approximation output of a trained artificial neural network.

20. The method according to claim 19, wherein the trained artificial neural network is a recurrent neural network.

21. The computer-implemented method according to claim 14, wherein the remedial action is a passive remedial action.

22. The computer-implemented method according to claim 21, wherein the passive remedial action comprises at least one of the following: (i) sounding an alarm, (ii) sending a notification, or (iii) displaying a warning.

23. The computer-implemented method according to claim 14, wherein the remedial action is an active remedial action.

24. The computer-implemented method according to claim 23, wherein the active remedial action comprises at least one of the following: (i) actuating a valve, (ii) causing a chimney to release a gas flare, or (iii) performing an action on the plant.

25. A process control system, the process control system comprising: A plurality of process data generation devices, the plurality of process data generation devices including one or more field devices configured to generate data corresponding to an ongoing industrial control process implemented by the process control system; and An electronic network that communicatively couples at least some of the plurality of process data generation devices to an anomaly / fault detection device, wherein the anomaly / fault detection device includes a memory having computer-executable instructions stored thereon, the computer-executable instructions, when executed by one or more processors of the anomaly / fault detection device, cause the anomaly / fault detection device to perform the following operations: Receive a set of metrics corresponding to the ongoing industrial control process; Determine an anomaly / fault action value and a normal action value corresponding to the set of metrics by cross-referencing the set of metrics with reinforcement learning information; and Cause a remedial action to occur based on the anomaly / fault action value.

26. The process control system according to claim 25, wherein the set of metrics corresponding to the ongoing industrial control process includes at least one of the following: (i) A predefined statistical threshold, (ii) A predefined constant control limit, (iii) A setpoint indicator, (iv) A load disturbance indicator, (v) An operating level transition indicator; or (vi) A combustion event indicator.

27. The process control system according to claim 25, wherein the anomaly / fault detection device includes a memory having computer-executable instructions stored thereon, the computer-executable instructions, when executed by one or more processors of the anomaly / fault detection device, cause the anomaly / fault detection device to perform the following operation: Generate the set of metrics by preprocessing process control data generated by the plurality of process data generation devices.

28. The process control system according to claim 25, wherein the plurality of process data generation devices in the process plant includes at least one of the following: Sensors, Valves, Transmitters, Positioners, Standard 4 mA to 20 mA devices, Field devices, Device Fieldbus device, Profibus devices, DeviceNet devices, ControlNet devices; or Modbus devices.

29. The process control system according to claim 25, wherein the reinforcement learning information includes a reinforcement learning Q-table.

30. The process control system according to claim 25, wherein the reinforcement learning information includes a function approximation output of a trained artificial neural network.

31. The process control system according to claim 30, wherein the trained artificial neural network is a recurrent neural network.

32. The process control system according to claim 25, wherein the remedial action is a passive remedial action.

33. The process control system according to claim 32, wherein the passive remedial action includes at least one of the following: (i) sounding an alarm, (ii) sending a notification, or (iii) displaying a warning.

34. The process control system according to claim 25, wherein the remedial action is a proactive remedial action.

35. The process control system according to claim 34, wherein the proactive remedial action includes at least one of the following: (i) actuating a valve, or (ii) causing a chimney to release a gas flare or (iii) performing an action regarding the plant.

36. One or more data generation devices, the one or more data generation devices being configured to: Receive enhanced learning information; Generate data corresponding to an ongoing industrial control process of a process control system; And Use the received enhanced learning information to process the generated data.

37. The one or more data generation devices according to claim 36, wherein each data generation device among the data generation devices is at least one of the following: A sensor, A valve, A transmitter, A locator, A standard 4 mA to 20 mA device, A field device, Device Fieldbus device, A Profibus device, A DeviceNet device, A ControlNet device; or A Modbus device.

38. The one or more data generation devices according to claim 36, wherein the device is further configured to: In response to receiving an activation instruction from a remote anomaly / fault detection application, generate the data corresponding to the ongoing industrial control process of the process control system.

39. The one or more data generation devices according to claim 36, wherein the received enhanced learning information includes metric-rewards and a Q-table.

40. The one or more data generation devices according to claim 36, wherein the enhanced learning information includes the function approximation output of a trained artificial neural network.

41. The one or more data generation devices according to claim 40, wherein the trained artificial neural network is a recurrent neural network.

42. The one or more data generation devices according to claim 39 or 40, wherein the device is further configured to: Calculate a set of metrics corresponding to the ongoing industrial process; Determine an anomaly / fault action value and a normal action value corresponding to the set of metrics by cross-referencing the set of metrics with the enhanced learning information; and Based on the anomaly / fault action value, cause a remedial action to occur.

43. The one or more data generation devices according to claim 42, wherein the remedial action is a passive remedial action.

44. The one or more data generation devices according to claim 43, wherein the passive remedial action includes at least one of the following: (i) sounding an alarm, (ii) sending a notification, or (iii) displaying a warning.

45. The one or more data generation devices according to claim 44, wherein the remedial action is a proactive remedial action.

46. The one or more data generation devices according to claim 45, wherein the proactive remedial action includes at least one of the following: (i) actuating a valve, or (ii) causing a chimney to release a gas flare or (iii) performing an action regarding the plant.

47. The one or more data generation devices according to claim 42, wherein the remedial action includes causing at least one of the one or more data generation devices to stop generating the data corresponding to the ongoing industrial control process of the process control system.

48. The one or more data generation devices according to claim 36, the one or more data generation devices being further configured to: Send the generated data to one or both of (i) another data generation device and (iii) a remote anomaly / fault detection device.

49. The one or more data generation devices according to claim 48, the one or more data generation devices being further configured to: Receive updated reinforcement learning information in response to the sending.

Citation Information

Cited By

  • Industrial quality inspection method and system based on general quality inspection large model

    CN122067059A