Systems and methods for detecting and predicting faults in industrial process automation systems

Through machine learning technology, the problem of large data volume and high professionalism is solved, and the faults and root causes are quickly identified and the efficiency and accuracy of fault detection are improved.

CN113614666BActive Publication Date: 2025-08-26SCHNEIDER ELECTRIC SYSTEMS USA INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202080022643.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-05-03
Filing Date
2020-03-24
Publication Date
2025-08-26
Estimated Expiration
2040-03-24

AI Technical Summary

Technical Problem

Modern industrial process automation systems generate a large amount of technical and system-specific data, making it difficult to effectively monitor and maintain system health, and require professional knowledge, which increases the difficulty of fault detection.

Method used

Using machine learning technology, automated detection and prediction of faults through data extraction, feature extraction and network topology construction, provides intuitive contextual format display, reduces dependence on expertise and improves the efficiency of fault detection and prediction.

Benefits of technology

System-level health monitoring of industrial process automation systems is realized, and potential failures, root causes and corrective measures are quickly identified, reducing the demand for specialized subject matter experts and improving analysis and response speed.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113614666B_ABST
    Figure CN113614666B_ABST
Patent Text Reader

Abstract

A system and method for detecting and predicting faults in industrial process automation systems uses trend data to predict warnings and allow action to be taken before problems occur. The system and method provide improved fault / failure predictions over time as more empirical data is collected for related sets of system components. The system and method can identify relationships between components of a process automation system; identify and collect changes to system configurations; identify and collect data to inform reliability and predictive models; develop domain-specific predictive models for one or more components that allow component-based failure or degradation predictions; develop system predictive models that utilize reliability and criticality relationships, component-based predictions, and operating parameters to predict the health of a portion or the entire process automation system; provide a prioritized warning system; and identify root causes of component failures.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0002] This patent application claims the benefit of priority to and is incorporated herein by reference in its entirety by reference to U.S. Provisional Application No. 62 / 823,377, filed on March 25, 2019, entitled “Systems and Methods for Detecting and Predicting Faults in an Industrial Process Automation System”; U.S. Provisional Application No. 62 / 823,469, filed on March 25, 2019, entitled “Systems and Methods for Performing Industrial Plant Diagnostics and Operations”; and U.S. Provisional Application No. 62 / 842,929, filed on May 3, 2019, entitled “Systems and Methods for Performing Industrial Plant Diagnostics.” Technical Field

[0003] Aspects of the present disclosure relate generally to industrial process automation and control systems. More specifically, aspects of the present disclosure relate to systems and methods for detecting and predicting faults in industrial process automation systems. Background Art

[0004] A typical industrial plant uses numerous interrelated and interconnected process automation systems to control and operate plant processes. Each system typically generates data in the form of log files specific to that system's operations. Log files provide a record of events occurring within the system (including the date and time of occurrence), as well as messages and communications between the system's various components. These log files allow personnel to monitor various systems for malfunctions, trace the root cause of any failures, and take appropriate corrective action.

[0005] Modern process automation systems generate data and error messages at an extremely high rate, resulting in massive amounts of data generated in a short period of time. This sheer volume of data can often overwhelm plant personnel and those attempting to monitor and interpret the data. Furthermore, each system generates data in a system-specific format that often differs from other systems, making it difficult to interpret data and error messages. Furthermore, the data and error messages generated by each system are often highly technical, requiring plant personnel to possess specialized knowledge of that particular system. Further complicating matters, each system maintains data and error messages in separate, often indistinguishable locations.

[0006] Therefore, there is a need for improvements in the field of industrial process automation, particularly in monitoring and maintaining the health of industrial process automation systems. Summary of the Invention

[0007] Embodiments of the present disclosure provide systems and methods for detecting and predicting faults in industrial process automation systems. This embodiment is particularly useful in industrial process automation systems that employ distributed control systems. In some embodiments, the system and method use trend data to predict warnings and allow action to be taken before problems occur. The system and method provide improved fault / failure predictions over time as more empirical data is collected for related sets of system components. The system and method can identify interrelationships between components of a process automation system; identify and collect changes to system configurations; identify and collect data to inform reliability and predictive models; develop domain-specific predictive models for one or more components that allow component-based failure or degradation predictions; develop system predictive models that use reliability and criticality relationships, component-based predictions, and operating parameters to predict the health of part or all of a process automation system; provide a prioritized warning system; and identify the root cause of component failure.

[0008] A fault associated with the first one or more devices can be brought to the user's attention by displaying faults associated with other devices in the system. In some embodiments, the systems and methods herein for detecting faults in an industrial process automation system can determine and display a root cause of the displayed fault (i.e., display an indication that the fault associated with the first one or more devices is a root cause of faults associated with other devices). In some embodiments, the systems and methods can track past and / or current solutions to system health issues to verify the effectiveness of proposed future solutions.

[0009] In some embodiments, the systems and methods herein for detecting faults in process automation systems can encode and automate subject matter expertise in diagnosing and predicting system problems. This can reduce the demand for dedicated subject matter experts and improve the speed of analysis and response. In some embodiments, the systems and methods herein can place separate warnings in context based on reliability, system interaction, and criticality. The systems and methods can automate root failure cause detection and additionally identify root failure causes that originate from another component or configuration change of the system. In some embodiments, the systems and methods can map one or more log messages and / or warnings, system data, context, or relationships to human-readable text excerpts.

[0010] In some embodiments, the systems and methods herein for detecting faults in process automation systems can generate customized system reliability and warning models for one or more process automation system components; integrate system reliability and warning models based on relationships between components; perform trend-based warnings; perform solution effect predictions based on historical actions and system impacts; and perform root cause identification at the system level.

[0011] In some embodiments, the systems and methods herein for detecting faults in a process automation system may use a structural view of the process automation system and its components to identify one or more critical components and connections and generate a relational database; build a reliability / warning model for one or more components and connections based on subject matter expertise; identify and capture relevant data for one or more components; adjust the model based on specific characteristics of the component / system; identify operational and trend data for one or more components; detect when one or more entities have an abnormal condition or a predicted abnormal condition; evaluate the root cause of the abnormal condition (e.g., the entity itself, another related entity, or a configuration change) and evaluate the impact of the condition on the system; convert the identified condition into a human-readable text excerpt; and record one or more corrective actions and associate them with previous warnings, patterns, and corrective actions to predict the effect of one or more actions.

[0012] In general, in one aspect, embodiments of the present disclosure relate to a monitoring system for an industrial plant. The monitoring system includes, among other things, one or more processors and a memory unit communicatively coupled to the one or more processors. The memory unit stores processor-executable instructions that, when executed by the one or more processors, cause the monitoring system to execute a process for inputting a data file to the industrial plant, the data file containing data related to nodes in the industrial plant, the data in each data file being in a different data format. The processor-executable instructions further cause the monitoring system to execute a process for extracting data from the data file, the extracted data including a timestamp, a device name, a device health status, and message content, and for converting the timestamp, device name, device health status, and message content from the data file into a homogeneous format. The processor-executable instructions further cause the monitoring system to execute a process for extracting features from the converted timestamp, device name, device health status, and message content using machine learning to identify the features, and for identifying nodes in the industrial plant that are experiencing an alarm, the alarm indicating that the node has failed or will fail within a specified time.

[0013] According to any one or more of the foregoing embodiments, the processor executable instructions further cause the monitoring system to execute a process of constructing a network topology for the node in the industrial plant, the network topology establishing a hierarchical structure for the node in the industrial plant. According to any one or more of the foregoing embodiments, the processor executable instructions further cause the monitoring system to execute a process of identifying the root cause of the alarm using machine learning to identify the root cause, executing a process of estimating the probability of the root cause using machine learning to calculate the probability, executing a process of displaying a time-to-failure for the alarm based on the probability of the root cause, and / or executing a process of graphically displaying a severity level of the alarm based on the time-to-failure of the alarm and / or the impact of the alarm on plant operations. According to any one or more of the foregoing embodiments, the processor executable instructions further cause the monitoring system to execute a process of graphically displaying all data within a specified time period for the node in the industrial plant that is experiencing the alarm, and / or executing a process of graphically displaying all nodes in the industrial plant that are experiencing the alarm using machine learning to identify the node. According to any one or more of the foregoing embodiments, the processor-executable instructions further cause the monitoring system to execute a process that uses machine learning to identify a corrective action for the alarm based on captured knowledge to extract the corrective action, the captured knowledge comprising a maintenance log for the industrial plant. According to any one or more of the foregoing embodiments, the processor-executable instructions further cause the monitoring system to extract features by executing a process that uses machine learning to apply feature extraction rules to the converted timestamp, device name, device health status, and message content.

[0014] In general, in another aspect, embodiments of the present disclosure relate to a method for monitoring an industrial plant. The method includes, among other things, inputting a data file for the industrial plant, the data file containing data related to nodes in the industrial plant, the data in each data file being in a different data format; and extracting the data from the data file, the extracted data including a timestamp, a device name, a device health status, and a message content. The method also includes converting the timestamp, device name, device health status, and message content from the data file into a homogeneous format; and extracting features from the converted timestamp, device name, device health status, and message content using machine learning to identify the features. The method further includes identifying a node in the industrial plant that is experiencing an alarm using machine learning to identify the node that is experiencing an alarm, the alarm indicating that the node has failed or will fail within a specified time.

[0015] According to any one or more of the preceding embodiments, the method further includes constructing a network topology for the nodes in the industrial plant, the network topology establishing a hierarchical structure for the nodes in the industrial plant. According to any one or more of the preceding embodiments, the method further includes identifying a root cause of the alarm using machine learning to identify the root cause. According to any one or more of the preceding embodiments, the method further includes estimating a probability of the root cause using machine learning to calculate the probability; displaying a time to failure of the alarm based on the probability of the root cause; and / or graphically displaying a severity level of the alarm based on the time to failure of the alarm and / or the impact of the alarm on plant operations. According to any one or more of the preceding embodiments, the method further includes graphically displaying all data within a specified time period for the node in the industrial plant experiencing the alarm, and / or graphically displaying all nodes in the industrial plant experiencing the alarm using machine learning to identify the node. According to any one or more of the preceding embodiments, the method further includes identifying a corrective action for the alarm based on captured knowledge including maintenance logs for the industrial plant to extract the corrective action and / or applying feature extraction rules using machine learning to the converted timestamps, device names, device health status, and message content.

[0016] In general, in yet another aspect, embodiments of the present disclosure relate to a computer-readable medium storing computer-readable instructions for causing the one or more processors to perform a method according to any one or more of the preceding embodiments. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] A more detailed description of the disclosure, briefly summarized above, can be obtained by reference to various embodiments, some of which are illustrated in the accompanying drawings. Although the drawings illustrate selected embodiments of the disclosure, these drawings are not to be considered limiting of its scope, as the disclosure may admit to other equally effective embodiments.

[0018] Figure 1 An exemplary industrial plant monitoring system according to an embodiment of the present disclosure is shown;

[0019] Figure 2 An exemplary machine learning process that may be used with embodiments of the present disclosure is shown;

[0020] Figure 3 An exemplary method of monitoring an industrial plant according to an embodiment of the present disclosure is shown;

[0021] Figure 4 shows an exemplary network topology according to an embodiment of the present disclosure;

[0022] Figures 5A to 5Bshows an exemplary failure analysis screen of an HMI for an exemplary industrial plant monitor according to an embodiment of the present disclosure;

[0023] Figure 6 shows an equipment error trend screen of an HMI for an exemplary industrial plant monitor according to an embodiment of the present disclosure; and

[0024] Figure 7 shows an aggregate alarm screen of an HMI for an exemplary industrial plant monitor according to an embodiment of the present disclosure;

[0025] Figure 8 shows a detailed alarm screen of an HMI for an exemplary industrial plant monitor according to an embodiment of the present disclosure; and

[0026] Figure 9 An aggregated log message screen of an HMI for an exemplary industrial plant monitor is shown in accordance with an embodiment of the present disclosure.

[0027] Where appropriate, identical reference numerals are used to designate identical elements that are common to the figures. However, elements disclosed in one embodiment may be beneficially utilized on other embodiments without specific recitation. DETAILED DESCRIPTION

[0028] This specification and the accompanying drawings illustrate exemplary embodiments of the present disclosure and should not be considered restrictive, wherein the claims define the scope of the present disclosure (including equivalents). Various mechanical, composition, structure, electrical and operational changes may be made without departing from the scope of this specification and the claims (including equivalents). In some cases, well-known structures and technologies are not shown or described in detail in order to avoid blurring the present disclosure. In addition, elements described in detail with reference to an embodiment and their associated aspects (as long as they are tried) may be included in other embodiments that are not specifically shown or described. For example, if an element is described in detail with reference to an embodiment and the element is not described in detail with reference to a second embodiment, the element can still be claimed to be included in the second embodiment.

[0029] Now refer to Figure 1, an industrial process automation system 100 is shown for executing an industrial process, such as an oil refining process, a chemical processing process, a manufacturing process, or the like. The specific industrial process automation system 100 described herein is generally referred to as a distributed control system (DCS). The DCS 100 generally includes a plurality of field devices, one of which is designated 102, that execute some sub-process within the industrial process. These field devices 102 may include, for example, sensors, actuators, motors, valves, and the like. Several of the field devices 102 are connected to a device bus, one of which is designated 104, that allows messages to be sent to and from the devices 102. Each device bus 104 is connected to an I / O module, one of which is designated 106, that links the devices 102 to a control processor 108. The I / O module 106 is responsible for regulating messages to and from the devices 102 and may be, for example, a fieldbus module (FBM). The control processor 108 operates substantially autonomously to control the operation of the device 102 and may be, for example, a programmable logic controller (PLC), a remote terminal unit (RTU), a programmable automation controller (PAC), etc. The control network 110 allows the control processors 108 to communicate with each other and with other systems on the network 110. The control network 110 may be implemented using Ethernet switches, gigabit interface converters (GBICs), fiber optic cables, cables, etc., and / or may be wireless.

[0030] One or more data servers 112 and one or more application workstations / workstation processors (AW / WP) 114, among other components, are also connected to the network 110. The application workstations / workstation processors 114 allow plant personnel to perform various tasks related to the industrial process, such as running tests, configuring hardware, installing software, modifying process parameters, etc., typically manually through a graphical user interface. The data servers 112 automatically provide data used to execute the industrial process (or portions thereof) to the control processor 108 and also automatically retrieve data from the control processor 108 for monitoring the industrial process. Some of the data retrieved by the data servers 112 may include events occurring at the control processor 108 (including the date and time of occurrence) and messages and communications sent and received by the control processor 108. These events and messages are typically retrieved by a process monitoring application running on the data servers 112, which typically records the data in log files and system health data.

[0031] The log files and system health data can then be sent by the data server 112 over the industrial network 116 to an industrial plant diagnostic system 120 that monitors the entire industrial process. The industrial plant diagnostic system 120 provides, among other things, high-level monitoring of the individual process controllers 108, which allows plant personnel to oversee the industrial process and coordinate the operations of the various process controllers 108. In the example shown, the exemplary industrial plant diagnostic system 120 has a typical system architecture that includes one or more processors 122, internal input and / or output ("I / O") interfaces 124, and memory 126, all of which are communicatively coupled and / or electrically connected to one another. The operation of these components of the system 120 is generally well known in the art and, therefore, will only be briefly mentioned here.

[0032] In general, one or more processors 122 are adapted to execute processor-executable instructions stored in memory 126. I / O interface 124 allows processor 122 to interact and communicate with external systems (and users). This communication can be accomplished via one or more communication networks (such as industrial network 110, wide area network (WAN), local area network (LAN), etc.). Memory 126 is adapted to provide processor-executable instructions to processor 122 upon request. Other computing components known to those skilled in the art may also be included in industrial plant diagnostic system 120 within the scope of the present disclosure.

[0033] As described above, the data server 112, or more specifically, the process monitoring applications running thereon, can generate log files and system health data at an extremely high rate, thereby generating a large amount of data in a short period of time. In addition, the data generated can be highly technical and system-specific, requiring plant personnel to have specialized knowledge of that particular system. Examples of process monitoring applications that can be run on the data server 112 include SMON (System Monitor), Wireshark, Multicast, Netsight, Syslog, System Auditor, as well as data history and archiving applications, counters from local observation databases, application workstation space reporting, control processor load reporting, CRC (Cyclic Redundancy Check) and other error checking applications, and system and network trap and interrupt routines known to those skilled in the art. The sheer volume of data and the highly technical nature of the data can often overwhelm plant personnel who need to use the data to detect faults in the DCS 100 and determine their root causes.

[0034] Thus, according to an embodiment of the present disclosure, memory 126 stores an industrial plant monitor 130 that can automatically process data generated by data servers 112 to detect and predict faults in DCS 100 (and similar industrial process automation systems). In particular, industrial plant monitor 130 can automatically aggregate log files and system health data from various data servers 112 and analyze the impact and interactions between different components to identify possible component failures, estimate the time until failure, issue alarms for failures, and determine the likely root causes of the failures. Industrial plant monitor 130 can also present data in an intuitive, context-based format that allows plant personnel to quickly assess potential fault conditions, their criticality, their root causes, and determine possible corrective actions. In short, industrial plant monitor 130 can provide a system-level health monitor for the entire DCS 100, not just for individual devices therein.

[0035] In some embodiments, the industrial plant monitor 130 includes multiple functional modules that work together to provide the above-mentioned system-level health status monitoring. These functional modules may include, for example, a data extractor 132, a network topology builder 134, a feature extractor 136, a root cause identifier 138, a root cause probability estimator 140, a time before failure estimator 142, a pattern matching module 144, one or more ML algorithms 146, and an HMI application 148. Although the functional modules are shown here as discrete blocks, those skilled in the art will understand that two or more blocks can be combined into a single block and any single block can be divided into several constituent blocks without departing from the scope of this disclosure. A more detailed description of the operation of the industrial plant monitor 130 is described later herein.

[0036] Next reference Figure 2 , shows an exemplary machine learning process 200 that can be used with the industrial plant monitor 130 in some embodiments. The machine learning process 200 can be applied by any of the functional modules 132 to 144 that require some form of machine learning to analyze log files and system health data from the various data servers 112. In the example shown, the machine learning process 200 includes a data input component 202, one or more machine learning models or algorithms 204, an automatic feedback / correction component 206, a user application 208, a manual feedback / correction component 210, and an analyzer 212. Examples of algorithms that can be used with the machine learning process 200 include Long Short-Term Memory (LSTM), random forests, decision trees, natural language processing, and the like.

[0037] In general operation, the data input component 202 receives data (e.g., log files, system health data, etc.) and, after appropriate preprocessing, feeds the data to one or more machine learning models 204. The machine learning models 204 use machine learning and neural network processing techniques to extract relevant features from the data (e.g., events, date and time of occurrence, error messages, etc.). The automatic feedback / correction component 206 applies rules and algorithms configured to detect errors in the output received from the machine learning models 204. These errors are used to automatically correct the model output and are fed back to the machine learning models 204 via the analyzer 212 to update the processing of the machine learning models 204. The processed output from the automatic feedback / correction component 206 is then displayed to the user for verification via the user application 208. Corrections made by the user are captured by the manual feedback / correction component 210 and fed back to the machine learning models 204 via the analyzer 212. This allows the machine learning models 204 to continuously improve their evaluation and extraction of relevant features from the input data.

[0038] Figure 3 300 is a flowchart illustrating an exemplary operation of the industrial plant monitor 130 according to an embodiment of the present disclosure. Flowchart 300 has two main phases: a configuration phase 302 and an application phase 304. During the configuration phase 302, the industrial plant monitor 130 builds the network topology for the DCS 100 and performs other software and hardware configuration tasks. During the application phase 304, the industrial plant monitor 130 performs analysis on log files and system health data to detect faults, predict time to failure, and identify root causes. The industrial plant monitor 130 continuously executes both phases 302 and 304 as needed, so they run substantially in parallel with each other.

[0039] During the configuration phase 302, output from a network discovery application is obtained for the DCS 100 and stored in a component and module inventory database 306. The information stored in the component and module inventory database 306 can be provided using any suitable network discovery application that can search a network such as the DCS 100 and discover network nodes, connectivity, routing protocols, etc. The industrial plant monitor 130 then uses this inventory to build a network topology for the DCS 100 at block 308. In some embodiments, building the network topology involves aggregating all nodes in the DCS 100 and positioning the nodes in a hierarchy based on their relationships and connectivity to one another. This information may include information that uniquely identifies each node, such as each node's IP address, MAC address, letter code (e.g., an alphanumeric identifier), etc., as well as any hardware, software, and firmware numbers for each node and the network routing protocols used by each node. In the case of fault-tolerant devices, information related to primary devices and secondary or backup devices may also be collected. The industrial plant monitor 130 can then use this information to build a network topology that details how the nodes are connected to one another and how data is transferred between nodes in the network.

[0040] Figure 4 An example of a network topology 400 for a DCS 100 is shown, constructed by the industrial plant monitor 130 based on the discovery output during the configuration phase 302. The network topology 400 in this example includes multiple branches 402 connected to each other to form the overall network topology 400. Each branch includes one or more root devices 404 (e.g., root bridges) and one or more switches 406 connected together via a mesh network 408. One or more workstations 410 are connected to one or more of the switches 406 via the mesh network 408, as are one or more control processors 412. If a backup or shadow control processor is provided, the shadow control processor may also be connected to one or more of the switches 406 via the mesh network 408. The one or more control processors 412 are in turn connected to one or more fieldbuses 414, which link the one or more control processors 412 to one or more field devices 416 via one or more fieldbus modules 418 (FBMs), such as field device system integrator modules (FDSIs). This network topology 400 may then be stored in the network database 310 for subsequent use in the application phase 306 to detect faults, predict time to failure, and determine root causes for failures.

[0041] Application phase 306 ( Figure 3 ) typically begins by inputting data sources into the industrial process monitor 130 at box 312. The data sources may include log files and system health data generated by process monitoring applications (e.g., SMON, Wireshark, Syslog, System Auditor, CRC, etc.) running on the aforementioned data servers, as well as other time-series data. At box 314, the industrial process monitor 130 extracts relevant data from the various data sources. This data extraction may involve reading the time-series data from the various data sources and loading the data into memory, and then converting the data from the various sources into a homogeneous format. An exemplary homogeneous format may include a timestamp field, a data source field, a field for the device that generated the data, a field for a message associated with the data, a field for the name of the device that is the subject of the message, and the like. The extracted data may then be used to dynamically update the network topology with any devices identified by the data that are not yet included in the network topology.

[0042] At block 316, the industrial process monitor 130 extracts relevant features from the extracted data to be used in conjunction with the aforementioned machine learning process 200 ( Figure 2 ) together. This feature extraction can involve applying various rules to the extracted data to identify relevant features based on the data. The data can then be resampled at regular intervals (e.g., every 10 minutes, 20 minutes, 30 minutes, etc.) and relevant features can be extracted from the data again. Examples of rules that can be applied to the extracted data are shown in Table 1 below.

[0043]

[0044]

[0045]

[0046] Table 1: Feature extractor rules

[0047] The above rules can be applied to the data extracted (block 314) from the various data sources to identify relevant features for machine learning purposes. Exemplary types of features that can be extracted include the following: device type, daily ARP count, ARP search device, daily total GBIC error count, daily intermittent GBIC error count, GBIC trend count, ReadLM error count, percentage of control processors showing errors, equipment failures, daily topology change count, intermittent topology count, daily bus error count, intermittent bus errors, analog error count, intermittent analog error count, marry-remarried intermittent pattern, etc.

[0048] At block 318, the industrial process monitor 130 identifies potential failures and the root causes of the failures based on the extracted features (e.g., using the machine learning process 200). This failure / root cause identification may involve training a machine learning model, such as a random forest or decision tree, to identify the root causes using historical log file data. The network topology information used for root cause identification is provided by the network database 310. In one example, nine months of historical log file data from an industrial process automation system (e.g., DCS 100) was used. Six weeks of data were selected from this data, with four weeks of data used for training and two weeks of data used to validate the training.

[0049] Training involves creating a feature matrix using real data (e.g., messages) and feature labels. The dimension of this matrix is ​​N×M, where N is the number of features and M is the number of feature extraction rules (e.g., Table 1). Feature labels are derived from plant maintenance logs and input from subject matter experts. Intermediate labels are created for logically related groups of features. Thus, features related to topology changes, increased ARP patterns, ReadLM errors, etc. are given an intermediate label such as "switch hardware problem". Similarly, features related to increased topology changes, GBIC errors, increased ARP searches, etc. are labeled as, for example, "switch GBIC problem". Bus access errors are labeled as, for example, "bus access error", and A-to-D faults are labeled as, for example, "A-to-D device failure", while control processor reconnection failure and module reset errors are labeled as, for example, "control processor hardware error". Errors such as intermittent ReadLM errors and intermittent ARP messages are labeled as, for example, "dirty fiber between switch and control processor", while intermittent ARP messages for all connected devices are labeled as, for example, "slow response time".

[0050] At block 320, the industrial process monitor 130 identifies the probability of the failure / root cause based on the identified failure / root cause (e.g., using the machine learning process 200). This probabilistic identification may involve a similar process of training a machine learning model (e.g., a random forest, a decision tree, etc.) using the historical log file data as described above with respect to block 318. In some embodiments, the device associated with the identified root cause may also be provided as input for training the machine learning model.

[0051] At block 322, the industrial process monitor 132 predicts a time to failure based on the failure / root cause probability determined in block 318 (e.g., using the machine learning process 200). This time to failure prediction may involve assigning a predefined time interval to a given failure / root cause based on the probability of the given failure / root cause. The duration of the time to failure interval may be based on, for example, historical log files, system health data, and error data. For example, if the probability of a given failure is greater than 99%, then the failure has occurred and may be assigned a time to failure of zero days by the industrial process monitor 130. If the probability of a given failure is 90% or higher, then the failure has occurred or is about to occur and may be assigned a time to failure of 24 hours.

[0052] If the probability of a given failure is greater than 30% but less than 90%, the industrial process monitor 130 can predict that the features extracted for that failure will occur within 1 day (i.e., project or interpolate the data forward 1 day). The industrial process monitor 130 can then rerun the root cause and probability identification with these features to see if the probability has reached 90%. If so, the 1-day pre-failure time is retained for that failure. If not, the industrial process monitor 130 increases the pre-failure time by another day and repeats the process until the probability reaches 90%. If the number of days increased exceeds 5 days, no pre-failure time is assigned.

[0053] In some embodiments, an industrial process monitor can use a Random Forest Regression (RFR) model to find the pre-failure time interval (while Random Forest Classification (RFC) can be used to find the root cause). Constructing the RFR involves setting the actual day of the failure (e.g., as reported by the field engineer) as day 0 for that failure in the training data, and then looking at features in the data going back, for example, to the previous 5 days. Table 2 shows an example training data set for the RFR model. The data is for a component problem that causes communication problems between the control processor and the fieldbus module. Two trends can be seen related to PIO bus access errors and fault-tolerant MAC reset counts.

[0054] PIO bus access error Fault-tolerant MAC reset error Remaining hours 214 595 0 162 400 24 162 200 48 100 95 72 35 22 96

[0055] Table 2: Example RFR training data

[0056] At block 324, the industrial process monitor 130 performs pattern matching on the data from the knowledge capture database 326 to determine whether the same or similar failure / root cause has occurred previously and what corrective actions were taken to resolve the failure. The data stored in the knowledge capture database 326 typically includes maintenance logs and records of actions previously taken by plant personnel to correct various errors that occurred over time in the DCS 100. These maintenance logs and records, which may include text documents, spreadsheets, etc., are typically maintained by plant personnel using common words and phrases. Therefore, the industrial process monitor 130 uses natural language processing (NLP) via the machine learning process 200 to extract relevant information from the maintenance logs and records. Natural language processing allows the industrial process monitor 130 to quickly filter out irrelevant words and phrases and focus on key information. Thus, for example, if multiple different corrective actions A, B, and C are taken to resolve a particular failure because the immediately preceding actions were ineffective, the industrial process monitor 130 can directly return to the final corrective action (Action C) that fixed the failure.

[0057] At block 328, the industrial process monitor 130 provides the above-described analysis to plant personnel in the form of an HMI (referred to herein as a dashboard). A dashboard is essentially a collection of screens that the industrial process monitor 130 can generate and display to a user, providing data to the user in an intuitive, context-based format that allows the user to quickly assess potential fault conditions, their criticality, their root causes, and determine possible corrective actions. The dashboard graphically visualizes the contents of large quantities (potentially millions) of log files and system health data that have been aggregated and transformed into usable, actionable information. From the HMI / dashboard, a user might be able to quickly see, for example, that a particular switch (e.g., switch TT2061) has a problem in a component (e.g., GBIC 17) that may cause the switch to fail soon (e.g., within the next five days).

[0058] Figures 5A to 5B An exemplary failure analysis screen 500 of a dashboard of an exemplary industrial plant monitor is shown. Screen 500 includes a switch 502 connected to a control processor 506 via one of its ports 508 via an optical fiber 504. The control processor 506 is in turn connected to a fieldbus module 512 via one of its ports 510 via a bus 514 (e.g., HDLC). Figure 5A In the example of , the industrial plant monitor has determined that a fault condition exists between one of the ports 510 of the control processor 506 and the fieldbus module 512 based on log messages from the control processor 506. Additionally, based on these log messages, the industrial plant monitor has determined that the fault condition is likely caused by the bus 514. In contrast, in Figure 5B In the example of , the industrial plant monitor has determined that the fault condition is likely caused by fiber optic cable 504 based on log messages from control processor 506 and switch 502 .

[0059] Figure 6 An exemplary device error trend screen 600 of a dashboard that can be generated and displayed by an industrial plant monitor is shown. The screen provides a device error trend indicated at 602, which shows a graph of error counts for each log source associated with the device. The device in this example is a TT2061, and the log source from which the device data is obtained is indicated at 604. In some embodiments, the log source can be color-coded and / or symbol-coded to distinguish between different log sources. From this screen, a user can quickly view the time interval in which the error count of a device begins to increase, thereby alerting the user to potential problems with the device within that time interval. Date and time information is provided at 606, and zoom options are provided at 608 (e.g., 1 hour, 3 hours, 6 hours, 1 day, 3 days, 1 week, etc.).

[0060] Figure 7 An exemplary aggregated alarm screen 700 of a dashboard that can be generated and displayed by an industrial plant monitor is shown. The primary purpose of this screen is to provide a list of all alarms and potential fault conditions in the DCS (indicated at 702). In the illustrated embodiment, the list 702 includes a date, device name, severity indicator, and a message field containing available and actionable information for each alarm in the list. In this example, the message field contains an identification of the root cause of the alarm along with a probability estimate for the alarm. Based on the probability estimate (e.g., the probability of the first alarm is 39%), the industrial plant monitor predicts a time to failure for the device in question (e.g., within three days). Consequently, the severity indicator may provide a severity of "high" for the device. In some embodiments, the severity indicator may be color-based (e.g., red for critical, yellow for high, orange for low, etc.), symbol-based (e.g., exclamation point for critical, question mark for high, dash for low, etc.), or a combination of both. In some embodiments, additional and / or alternative information, such as the date and time of the current analysis, may be included, indicated at 704. In some embodiments, a search box 706 may be included to allow a user to search for alarms and fault conditions using natural language queries. Selecting one of the alarms in list 702 (such as the alarm for device 01CP21) (eg, by tapping, double-clicking, etc.) takes the user to a detailed alarm screen for that alarm.

[0061] In some embodiments, the industrial plant monitor assigns severity levels (e.g., critical, high, low, etc.) to alarms based on the impact the alarm will have on the continuity of the plant and / or business operations. Alarms with more significant impact (e.g., potential process shutdown) are assigned a higher severity relative to alarms with lesser impact (e.g., throughput reduction). Thus, for example, switches and control processors can be assigned devices with higher criticality relative to application workstations, FBMs, field devices, etc. Similarly, areas controlled by controlled processors can be assigned high priority or medium / low priority based on the functions performed by the areas of the processor and the impact on business operations. In some cases, severity assignment can be performed manually by an operator during system configuration, and / or severity assignment can be performed continuously by the system using a machine learning algorithm trained with historical alarm training data. In either case, the ability to assign different severity levels to various alarms allows the industrial plant monitor to provide the operator with context for the alarms so that higher priority can be shifted to processes / areas in the plant affected by critical equipment.

[0062] Figure 8 An exemplary detailed alarm screen 800 of a dashboard that may be generated and displayed by an industrial plant monitor is shown. As the name implies, the screen displays detailed information about the selected alarm, including error details such as date, severity, analysis details, etc. indicated at 802, and device details such as device identification code, device description, software / hardware / firmware versions, etc. indicated at 804. The screen may also provide information related to the selected alarm at 806. Figure 7 , which shows a graph 808 of error counts for each log source associated with the device. To provide context, this screen also shows the relevant network portion 810 of the network topology where the device resides, so the user can see where the device resides within the DCS. In some embodiments, the screen also provides the option to zoom in and out 812 as needed.

[0063] Figure 9 An aggregated log message screen 900 of a dashboard that can be generated and displayed by an industrial plant monitor is shown. This screen aggregates log error messages and groups them by log file, system health data, and time. Thus, the error message indicated at 902 is from one log file, while the error message indicated at 904 is from a different log file, and so on. In some embodiments, the error messages can be color-coded for easier viewing. From this screen, a user can quickly view error messages for all devices in the DCS currently experiencing a fault condition.

[0064] Thus, as described herein, embodiments of the present disclosure provide systems and methods for detecting and predicting faults in industrial process automation systems.Such embodiments may include a special-purpose computer that includes various computer hardware as described in more detail below.

[0065] Embodiments within the scope of the present disclosure also include computer-readable media for carrying or having computer-executable instructions or data structures stored thereon.Such computer-readable media can be any available media that can be accessed by a dedicated computer, and include computer storage media and communication media.By way of example and not limitation, computer storage media include both volatile and non-volatile, removable and non-removable media for storing information (such as computer-readable instructions, data structures, program modules or other data) implemented in any method or technology.Computer storage media is non-temporary, and includes but is not limited to random access memory (RAM), read-only memory (ROM), electrically erasable programmable ROM (EEPROM), optical disc ROM (CD-ROM), digital versatile disk (DVD) or other optical disc storage devices, solid-state drive (SSD), cassette, magnetic tape, magnetic disk storage device or other magnetic storage device, or can be used for carrying or storing desired non-temporary information in the form of computer-executable instructions or data structures and any other medium that can be accessed by a computer.When information is delivered or provided to a computer via a network or another communication connection (hard-wired, wireless or a combination of hard-wired or wireless), the computer appropriately regards the connection as a computer-readable medium. Thus, any such connection is properly termed a computer-readable medium. The above combinations should also be included within the scope of computer-readable media. Computer-executable instructions include, for example, instructions and data that cause a general-purpose computer, a special-purpose computer, or a special-purpose processing device to perform a certain function or group of functions.

[0066] The following discussion is intended to provide a brief, general description of a suitable computing environment in which various aspects of the present disclosure may be implemented. Although not required, various aspects of the present disclosure will be described in the general context of computer-executable instructions (such as program modules) executed by computers in a network environment. Typically, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. Computer-executable instructions, associated data structures, and program modules represent examples of program code components for performing the steps of the methods disclosed herein. A specific sequence of such executable instructions or associated data structures represents an example of corresponding actions for implementing the functions described in such steps.

[0067] Those skilled in the art will appreciate that aspects of the present disclosure can be practiced in a network computing environment having many types of computer system configurations, including personal computers, handheld devices, multiprocessor systems, microprocessor-based or programmable consumer electronics, network PCs, minicomputers, mainframe computers, and the like. Aspects of the present disclosure can also be practiced in a distributed computing environment where tasks are performed by local and remote processing devices that are linked (by hardwired links, wireless links, or a combination of hardwired or wireless links) through a communications network. In a distributed computing environment, program modules can be located in both local and remote memory storage devices.

[0068] The exemplary system for implementing various aspects of the present disclosure includes a special-purpose computing device in the form of a conventional computer, including a processing unit, a system memory and a system bus that couples various system components including the system memory to the processing unit. The system bus can be any of a variety of bus structures, including a memory bus or a memory controller, a peripheral bus and a local bus using any of a variety of bus architectures. The system memory includes computer storage media, including non-volatile and volatile memory types. A basic input / output system (BIOS) can be stored in ROM, which includes basic routines that help transfer information between elements in the computer (such as during startup). In addition, the computer can include any device (e.g., a computer, a laptop computer, a tablet computer, a PDA, a mobile phone, a mobile phone, a smart TV, etc.) that can wirelessly receive an IP address from the Internet or send an IP address wirelessly to the Internet.

[0069] The computer may also include a magnetic hard drive for reading from and writing to a magnetic hard disk, a magnetic disk drive for reading from or writing to a removable disk, and an optical disk drive for reading from or writing to a removable optical disk (such as a CD-ROM or other optical media). The magnetic hard drive, the magnetic disk drive, and the optical disk drive are connected to the system bus via a hard drive interface, a magnetic disk drive interface, and an optical drive interface, respectively. The drive and its associated computer-readable medium provide the computer with non-volatile storage of computer-executable instructions, data structures, program modules, and other data. Although the exemplary environment described herein employs a magnetic hard disk, a removable disk, and a removable optical disk, other types of computer-readable media for storing data may also be used, including magnetic tape, flash memory cards, digital video disks, Bernoulli cassettes, RAM, ROM, SSDs, etc.

[0070] Communication media typically embodies computer readable instructions, data structures, program modules or other data in a modulated data signal such as a carrier wave or other transport mechanism and includes any information delivery media.

[0071] The program code device that comprises one or more program modules can be stored on hard disk, disk, CD, ROM and / or RAM, comprises operating system, one or more application programs, other program modules and program data.The user can input command and information into the computer by keyboard, pointing device or other input device (such as microphone, joystick, game controller, satellite dish, scanner etc.).These and other input devices are connected to processing unit by the serial port interface that is coupled to system bus usually.Alternatively, input device can be connected by other interface such as parallel port, game port or universal serial bus (USB).Monitor or another kind of display device also is connected to system bus via interface such as video adapter.Except monitor, personal computer also comprises other peripheral output devices (not shown), such as loudspeaker and printer usually.

[0072] One or more aspects of the present disclosure may be embodied in computer executable instructions (i.e., software), routines or functions stored in system memory or non-volatile memory as applications, program modules and / or program data. Alternatively, the software may be stored remotely, such as on a remote computer with a remote application. Typically, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types when executed by a processor in a computer or other device. Computer executable instructions may be stored on one or more tangible, non-temporary computer-readable media (e.g., hard disks, optical disks, removable storage media, solid-state memory, RAM, etc.) and executed by one or more processors or other devices. As will be appreciated by those skilled in the art, in various embodiments, the functionality of program modules may be combined or distributed as needed. In addition, the functionality may be embodied in whole or in part in firmware or hardware equivalents (such as integrated circuits, application specific integrated circuits, field programmable gate arrays (FPGAs), etc.).

[0073] A computer can operate in a networked environment using a logical connection to one or more remote computers. The remote computers can each be another personal computer, tablet computer, PDA, server, router, network PC, peer device or other public network node, and typically include many or all of the elements described above relative to the computer. Logical connections include local area networks (LANs) and wide area networks (WANs) presented here by way of example and not limitation. Such networked environments are common in office-wide or enterprise-wide computer networks, intranets and the Internet.

[0074] When used in a LAN networking environment, the computer is connected to the local network via a network interface or adapter. When used in a WAN networking environment, the computer may include a modem, a wireless link, or other components for establishing communications over a wide area network (such as the Internet). The modem (which may be internal or external) is connected to the system bus via a serial port interface. In a networked environment, program modules or portions thereof depicted relative to the computer may be stored in a remote memory storage device. It should be understood that the network connections shown are exemplary and that other components for establishing communications over a wide area network may be used.

[0075] Preferably, the computer executable instructions are stored in a memory such as a hard drive and executed by the computer. Advantageously, the computer processor has the capability to perform all operations (eg, execute the computer executable instructions) in real time.

[0076] Unless otherwise stated, the execution (execution / performance) order of the operations in the embodiments of the present disclosure shown and described herein is not necessary. That is, unless otherwise stated, operations can be performed in any order, and embodiments of the present disclosure may include more or fewer operations than those disclosed herein. For example, it is contemplated that performing a particular operation before, simultaneously with, or after another operation is within the scope of the various aspects of the present disclosure.

[0077] Embodiments of the present disclosure may be implemented using computer-executable instructions. Computer-executable instructions may be organized into one or more computer-executable components or modules. Aspects of the present disclosure may be implemented using any number and organization of such components or modules. For example, aspects of the present disclosure are not limited to specific computer-executable instructions or the specific components or modules shown in the figures and described herein. Other embodiments of the present disclosure may include different computer-executable instructions or components having more or less functionality than shown and described herein.

[0078] When introducing elements of aspects of the present disclosure or embodiments thereof, the articles "a," "an," "the," and "said" are intended to mean that there are one or more elements. The terms "comprising," "including," and "having" are intended to be inclusive and mean that there may be additional elements other than the listed elements.

[0079] Having described various aspects of the disclosure in detail, it will be apparent that numerous modifications and changes may be made without departing from the scope of the present disclosure as defined in the appended claims. As various changes may be made in the above-described structures, products, and methods without departing from the scope of the present disclosure in its various aspects, it is intended that all matter contained in the above description and shown in the accompanying drawings shall be interpreted as illustrative and not in a limiting sense.

Claims

1. A monitoring system for an industrial plant, the monitoring system comprising: one or more processors; a memory unit communicatively coupled to the one or more processors and having stored thereon processor-executable instructions that, when executed by the one or more processors, cause the monitoring system to: executing a process of inputting data files for the industrial plant, the data files containing data associated with nodes in the industrial plant, the data files for a given node including a log file for the given node, the data in each data file being in a different data format; executing a process of extracting the data from the data file, the extracted data including a timestamp, a device name, a device health status, and a message content, wherein the message content of the given node indicates that a device error has occurred at the given node; performing a process of converting said timestamp, device name, device health, and message content from said data file into a homogeneous format; performing a process of extracting features from the converted timestamp, device name, device health, and message content using machine learning to identify features comprising a count of errors experienced by the node at each given time interval; performing a process of identifying nodes in the industrial plant that are experiencing an alarm using machine learning and the extracted features to identify the nodes that are experiencing an alarm, the alarm indicating that the node has failed or will fail within a specified time; performing a process of graphically displaying on a dashboard an identification of a node experiencing an alert, and further graphically displaying on the dashboard a screen containing detailed information regarding the alert in response to a user tapping the identification of the node; as well as A process is performed to determine a time to failure of the alarm using machine learning, and if the time to failure is less than or equal to a preset threshold, graphically displaying the time to failure and an identification of the node on the dashboard.

2. The monitoring system of claim 1, wherein the processor-executable instructions further cause the monitoring system to execute a process of constructing a network topology for the nodes in the industrial plant, the network topology establishing a hierarchical structure for the nodes in the industrial plant.

3. The monitoring system of claim 1 , wherein the processor-executable instructions further cause the monitoring system to perform a process of identifying a root cause of the alarm using machine learning to identify the root cause.

4. The monitoring system of claim 3, wherein the processor-executable instructions further cause the monitoring system to perform a process of estimating the probability of the root cause using machine learning to calculate the probability.

5. The monitoring system of claim 4, wherein the processor-executable instructions further cause the monitoring system to perform a process of determining a time to failure of the alarm based on the probability of the root cause using machine learning.

6. The monitoring system of claim 5, wherein the processor-executable instructions further cause the monitoring system to perform a process for graphically displaying a severity level of the alarm based on the time to failure of the alarm and / or the impact of the alarm on plant operations.

7. The monitoring system of claim 1, wherein the processor-executable instructions further cause the monitoring system to perform a process of graphically displaying all data within a specified time period for the node in the industrial plant that is experiencing the alarm.

8. The monitoring system of claim 1, wherein the processor-executable instructions further cause the monitoring system to perform a process of using machine learning to graphically display all nodes in the industrial plant that are experiencing the alarm to identify the node.

9. The monitoring system of claim 1 , wherein the processor-executable instructions further cause the monitoring system to perform a process of using machine learning to identify a corrective action for the alarm based on captured knowledge to extract the corrective action, the captured knowledge comprising a maintenance log for the industrial plant.

10. The monitoring system of claim 1, wherein the processor-executable instructions further cause the monitoring system to extract features by performing a process of applying feature extraction rules to the converted timestamps, device names, device health status, and message content using machine learning.

11. A method for monitoring an industrial plant, the method comprising: inputting data files for the industrial plant, the data files containing data related to nodes in the industrial plant, the data files for a given node including a log file for the given node, the data in each data file being in a different data format; extracting the data from the data file, the extracted data including a timestamp, a device name, a device health status, and message content, wherein the message content of the given node indicates that a device error has occurred at the given node; converting the timestamp, device name, device health, and message content from the data file into a homogeneous format; extracting features from the converted timestamp, device name, device health, and message content using machine learning to identify the features, the features comprising a count of errors experienced by the node at each given time interval; identifying nodes in the industrial plant that are experiencing an alarm using machine learning and the extracted features to identify the nodes that are experiencing an alarm, the alarm indicating that the node has failed or will fail within a specified time; graphically displaying on a dashboard an identification of a node experiencing an alert, and further graphically displaying on the dashboard a screen containing detailed information regarding the alert in response to a user tapping the identification of the node; as well as A time to failure of the alarm is determined using machine learning, and if the time to failure is less than or equal to a preset threshold, the time to failure and an identification of the node are graphically displayed on the dashboard. 12 . The method of claim 11 , further comprising constructing a network topology for the nodes in the industrial plant, the network topology establishing a hierarchical structure for the nodes in the industrial plant.

13. The method of claim 11, further comprising identifying a root cause of the alarm using machine learning to identify the root cause.

14. The method of claim 13, further comprising estimating the probability of the root cause using machine learning to calculate the probability.

15. The method of claim 14, wherein determining the time to failure of the alarm using machine learning is based on the probability of the root cause. 16 . The method of claim 15 , further comprising graphically displaying a severity level of the alarm based on the time to failure of the alarm and / or the impact of the alarm on plant operations. 17 . The method of claim 11 , further comprising graphically displaying all data within a specified time period for the node in the industrial plant that is experiencing the alarm.

18. The method of claim 11, further comprising using machine learning to graphically display all nodes in the industrial plant that are experiencing the alarm to identify the node.

19. The method of claim 11, further comprising using machine learning to identify a corrective action for the alarm based on captured knowledge to extract the corrective action, the captured knowledge comprising maintenance logs for the industrial plant.

20. The method of claim 11, further comprising applying feature extraction rules to the converted timestamp, device name, device health, and message content using machine learning.

21. A computer-readable medium storing computer-readable instructions for causing one or more processors to execute the method according to any one of claims 11 to 20.