Intelligent alarm management method for use in industrial processes
By training an artificial neural network (ANN) model, abnormal behaviors in industrial processes can be identified and managed, solving the problem of excessive alarms in complex industrial processes and improving the accuracy and efficiency of alarm management.
Patent Information
- Application Number
- CN202180028760.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-04-16
- Filing Date
- 2021-04-13
- Publication Date
- 2025-11-18
- Estimated Expiration
- 2041-04-13
AI Technical Summary
In complex industrial processes, identifying abnormal behavior is difficult and prone to errors, leading to excessive alarms and making it difficult for operators to understand the criticality of the current process.
By training an artificial neural network (ANN) model, using input and score data, abnormal behaviors in industrial processes can be identified and managed, outputting alarms and providing scenario numbers and predicted process values, thereby reducing the over- or under-alarm nature of alarms.
Effectively identify abnormal behavior in industrial processes, reduce operator workload, improve the accuracy and efficiency of alarm management, and provide faster response and understanding of the current process.
Smart Images

Figure CN115427907B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of intelligent alarm management, particularly to intelligent alarm management in industrial processes. The invention also relates to computer program products, computer-readable storage media, and the use of such methods. Background Technology
[0002] At least some industrial processes (e.g., within a facility) can be so complex that operators and / or service personnel are not always aware of their behavior. In particular, for at least some cases, identifying anomalous behavior can be difficult and / or error-prone. This can become even more complicated because sometimes too many alerts may be issued, making it impossible to clearly understand the criticality of the current industrial process. Summary of the Invention
[0003] Therefore, the object of the present invention is to provide improved alarm management for industrial processes. This object is achieved through the subject matter of the independent claims. Further embodiments will be apparent from the dependent claims and the following description.
[0004] One aspect relates to a method for detecting anomalous behavior in industrial processes. The method includes the following steps:
[0005] A machine learning model is trained using input data and score data, where the machine learning model is an artificial neural network (ANN).
[0006] The input data includes a first time series of at least one observable process value of the industrial process, a second time series of at least one manipulated variable affecting the industrial process, and a third time series of at least one internal variable of the industrial process.
[0007] Furthermore, the scoring data includes: a first critical value for each of at least one observable process value indicating anomalous behavior of the industrial process, and a fourth time series of at least one predicted observable process value of the industrial process;
[0008] The trained machine learning model is run by applying the first time series data to it; and
[0009] The output value is generated by a trained machine learning model, which includes at least a second critical value that indicates at least one predicted observable process value indicating anomalous behavior of an industrial process within a predefined time interval.
[0010] Industrial processes can operate within industrial facilities, as used in, for example, chemical and process engineering. Industrial processes can be configured to produce and / or manufacture substances, such as materials and / or compounds. Anomalous behavior can be an intentional deviation from the industrial process and / or the industrial facility. Anomalous behavior can be indicated by observable (“external”) values, such as temperature or pressure within the facility's containers, and / or by unobservable (“internal”) values, such as unintentional internal interference with the mixing of compounds. Anomalous behavior may trigger alarms, whether short-term, such as immediately, and / or over a time interval, such as within seconds, minutes, and / or other time spans.
[0011] Machine learning models are artificial neural networks (ANNs) used after training and / or the training phase. Training can be performed once or repeated during the model's use. Training can be performed using input data and score data; however, additional data can also be used for training. The time series of input data and / or score data can be based on past data records. Therefore, the "future behavior" of an industrial process may be "known," i.e., it may be part of a time series. For example, a rapid change in one process value may lead to a critical situation within minutes, while a rapid change in another process value may have already proven to be non-critical, even if an alarm has been issued. Input data may include: (multiple) observable process values, (multiple) unobservable internal variables (possibly from simulation and / or, for example, from unobservable disturbances), and / or (multiple) manipulated variables (e.g., from operators responding to alarms and / or other behaviors of the industrial process).
[0012] The scoring data can be the rewards and / or penalties of the ANN. The scoring data can include: a first critical value, which can be a function of (multiple) observable process values and / or (multiple) internal variables. This function can be a composite function and / or a combined function of one or more variables. The function can be a simple function, such as: "If the temperature is below 32°C, then alarm5 = true" (alarm 5 = true). The predicted observable process values can be based on historical data that may show "developing" process values, such as: "If the temperature is above 76°C, then alarm8 = true" (alarm 8 = true) because the process behavior becomes critical within 2 minutes. The time interval can be fixed, such as 5 minutes, and it may be more than one interval, and / or a variable interval, possibly influenced by at least one historical time series.
[0013] The trained machine learning model can be run after the training phase. During the operation of the process, at least the first time series data is applied to the trained machine learning model; however, more data may be applied to the model.
[0014] The output of a trained machine learning model can include alerts and / or additional alerts. Alerts may be based on or may include a second threshold. The model's output may be similar to or different from other alerts, for example, through other components and / or subsystems of the industrial process. In some cases, the model's output may lead to a reassessment of alerts, such as potentially causing an alert to be "overweighted" and / or underweighted. This "correction" may contribute to more effective alert management and / or reduce the operator's burden in industrial processes. In particular, it can facilitate the identification of anomalous behavior in industrial processes.
[0015] In various embodiments, the output value also includes a scenario number for the industrial process, depending on at least one of a first time series, a second time series, and / or a third time series. If a scenario number cannot be found, an "undefined" scenario number may be output. Scenario numbers can advantageously contribute to a better understanding of what is currently happening in the industrial process, leading to a faster response to the current situation and / or relevant historical circumstances, and / or further investigation.
[0016] In various embodiments, the output values also include a fifth time series that depends on at least one of the first, second, and / or third time series. The fifth time series may resemble a fourth time series of predicted observable process values. The fifth time series may include “bypassing” the fourth time series, thereby advantageously utilizing a knowledge base provided by multiple historical time series. Therefore, based on past and current measurements and / or data of a given process, and also based on planned future operator actions when considering manipulated variables, the method can facilitate or contribute to predicting facility behavior. Furthermore, this can be used to train machine learning algorithms to create fast and accurate proxy models for the methods described above used in online deployment.
[0017] In various embodiments, the output value also includes a first critical value of at least one observable process value (i.e., the current value). Inputs can also be used as outputs, for example, simply "forwarded" and / or as a "shortcut" to the first critical value. This further enhances the understanding of the current process behavior.
[0018] In various embodiments, the method further includes the step of outputting a manipulated variable that depends on at least one of a first time series and / or a third time series. This implementation may include a “simple forwarding or bypassing” of the value (i.e., observable and / or unobservable values). This can advantageously achieve a seamless integration of sensor values and simulation results. This may further contribute to insight if a viable solution or “standard solution” exists for this situation. Furthermore, new scenarios may be communicated to the operator as indications of some particular concern regarding that situation and / or scenario.
[0019] In various embodiments, the method further includes a step of determining the time distance to a second threshold exceeding a predefined threshold. This can advantageously answer questions such as, "In this situation / scenario, when will the next alarm occur?" or "Will there be an alarm for this situation / scenario?" This can be advantageously used in certain situations to "discharge" alarm messages to personnel. Therefore, depending on the rate of increase of the threshold, a time buffer may be inserted for that particular alarm, for example, in the future. This can further improve alarm management.
[0020] In various embodiments, the method further includes the steps of: determining the rate of increase of a second threshold; and outputting an alarm when the rate of increase exceeds a predefined threshold. This may trigger an alarm as a response to an acceleration of some process value (e.g., temperature rise).
[0021] One aspect relates to a computer program product comprising instructions that, when executed by a computer and / or an artificial neural network (ANN), cause the computer and / or ANN to perform the methods described above and / or below.
[0022] One aspect relates to a computer-readable storage medium on which a computer program or computer program product as described above is stored.
[0023] One aspect involves machine learning models, particularly trained machine learning models, which are configured to perform the methods described above and / or below.
[0024] One aspect involves the use of machine learning models in monitoring and / or controlling industrial processes.
[0025] One aspect relates to an industrial facility that includes a computer and / or an ANN (Artificial Neural Network), wherein instructions are stored on the industrial facility, and when the program is executed by the computer and / or the ANN, it causes the computer or industrial facility to perform the methods described above and / or below.
[0026] To further clarify, the present invention is described with reference to embodiments shown in the accompanying drawings. These embodiments are considered examples only and not limitations. Attached Figure Description
[0027] These accompanying figures depict:
[0028] Figure 1 The illustration schematically depicts the collaboration between an industrial process and a machine learning model according to an embodiment;
[0029] Figure 2 A lookup table including input data and threshold values is schematically illustrated according to an embodiment;
[0030] Figure 3 Some elements and input data according to an embodiment are illustrated schematically;
[0031] Figure 4 The training process according to an embodiment is illustrated schematically;
[0032] Figure 5 The prediction process according to an embodiment is illustrated schematically;
[0033] Figure 6 The proxy model training workflow according to an embodiment is illustrated schematically;
[0034] Figure 7 An example of using a proxy model for predicting alarms, according to an embodiment, is illustrated schematically;
[0035] Figure 8 The intervals of an alarm management system corresponding to the time evolution of process variables according to an embodiment are schematically shown;
[0036] Figure 9 An interval or alarm limit according to another embodiment is illustrated schematically;
[0037] Figure 10 The training process of a machine learning model according to an embodiment is illustrated schematically;
[0038] Figure 11 The operational phases of the alarm management system according to an embodiment are schematically illustrated;
[0039] Figure 12 The illustration shows an example scenario where prediction-based alerts according to the embodiments are beneficial;
[0040] Figure 13 The proxy model training workflow according to an embodiment is illustrated schematically;
[0041] Figure 14An example of using a proxy model for predicting alarms, according to an embodiment, is illustrated schematically;
[0042] Figure 15 The method of using online simulation according to an embodiment is illustrated schematically;
[0043] Figure 16 An example of using data from simulation and machine learning to perform root cause analysis, according to an embodiment, is illustrated schematically.
[0044] Figure 17 The method according to an embodiment is illustrated schematically; Detailed Implementation
[0045] Figure 1 The collaboration between an industrial process 50 and a machine learning model 10 according to an embodiment is schematically illustrated. The industrial process 50 can be viewed or observed via a first time series 21 of observable process values PV. The industrial process 50 also has a third time series 23 of at least one internal variable IV, which may be unobservable, i.e., not directly observable and / or "observable" through simulation, etc. The industrial process 50 is controlled and / or manipulated by a second time series 22 of at least one manipulated variable MV. MV may be input into the system by operators, service personnel, and / or one or more automated processes. The machine learning model 10 may be an artificial neural network (ANN). The machine learning model 10 has input data 20, which includes the first time series 21, the third time series 23, and the second time series 22. The third time series 23 of IV may be displayed in an IV table 68.
[0046] The machine learning model 10 also has score data 30, which includes a first critical value 32 and a fourth time series 34. The first critical value 32 may be based on and / or may be a function of the current observable process value PV and / or based on a first time series 21 of the observable process value PV, thus considering a longer time span of PV. The function may be constructed by a mapping device 62. The mapping device 62 may also output the first critical value 32 to an alarm display 64. One or more alarm displays 64 may be present in the system. The alarm displays 64 may also be fed by other components and / or modules of the system, such as sensor outputs, such as temperature, pressure, etc., depending on the industrial process 50. The fourth time series 34 includes at least one predicted observable process value PPV of the industrial process 50. The prediction of the predicted observable process value PPV may be based on historical data. The machine learning model 10 outputs an output value 40. The output value 40 includes a second critical value 42 and a fifth time series 44. The fifth time series 44 can be similar to the fourth time series 34, and / or simply “feed forward” the fourth time series 34, thus making the simulation results available as a component of the model data. The second critical value 42 is a function of at least one predicted observable process value PPV. Therefore, the second critical value 42 is a kind of “condensed knowledge” of process behavior, which may include aspects of the future development of PV.
[0047] Figure 2 An example lookup table 70 according to an embodiment is schematically shown, including input data 20 and a second threshold 42. Input data 20 may include multiple time series 31, 22, 23 of multiple PVs, MVs, and IVs. Score data 30 may include a first threshold and / or other values. Lookup table 70 may also include a scenario number 46 of the industrial process 50, indicating the current "state" of the industrial process 50. Example lookup table 70 may also include a predefined time distance T1, which may, for example, indicate the basis of the second threshold 42, i.e., the predicted value in time distance T1.
[0048] Figure 3The diagram schematically illustrates some elements, input data, and their interactions according to an embodiment. Solution elements and their interactions are shown. The simulator may include models of industrial processes and equipment, as well as models of control systems. The virtual operator is a software module that can apply predefined or pre-programmed operator actions (e.g., setpoint or MV changes) to the simulated control system. Data exchange can be accomplished via appropriate APIs (Application Programming Interfaces), OPC DA (Application Programming Interface Data Access), or via scripted user actions (e.g., performing cursor movement and keyboard input). The virtual operator performs certain operator actions during a simulation run of the first-principles model starting from a given initial state (e.g., simulating internal behavior). Data generated by the first-principles model of the process facility is stored in a database and used for machine learning by an ANN or machine learning model. Alternatively, the ANN can be trained using data streams from the simulator without having to store the data.
[0049] Using the control loop setpoint from time t0 to t n Process values (PV), manipulated variables (MV) within a time window, or within the same time window from t0 to t n Operator actions within and from t n to t end The future setpoints of the plan are used to train the artificial neural network (sometimes called "ML learning algorithm" or "ML algorithm"). The target (the process variable to be predicted) is from t n+1 to t end The process values. During the operation, the trained ML model will be fed data similar to a predictor: from t0 to t... n The process values and setpoints, and from t n+1 to t end The model calculates the planned future trajectory from the operator's setpoint. It can output the expected facility behavior based on the trajectory of future process variables.
[0050] Figure 4 The training process according to an embodiment is illustrated schematically. This can be about... Figure 3 The subsequent activities are as follows: Initially, the virtual operator loads the setpoint configuration file. Then, the initial state of the simulation is loaded and the simulation is run. During the simulation run, the virtual operator manipulates the setpoint. Optionally, data is stored, possibly for other uses, and machine learning models are trained and saved.
[0051] Figure 5A predictive process, typically operating during the runtime of an industrial process according to an embodiment, is illustrated schematically. Past process values and setpoints are collected from a facility automation system, including process facility history records. The operator provides a planned setpoint trajectory; in the simplest case, this may include only one setpoint. A trained model performs predictions within a specified timeframe. Optionally, predictive alarm logic is used to indicate to the operator whether an alarm will be triggered by the predicted process value trajectory. Finally, the process value trajectory and the corresponding alarm are shown to the operator.
[0052] Figure 6 The proxy model training workflow according to an embodiment is illustrated schematically. A first-principles dynamic facility model is used to create data, which is recorded for use as training data. Simulation may be a viable method because operators tend to avoid reaching critical thresholds during normal production processes, resulting in insufficient training data. However, in simulation, such training data can be generated without negatively impacting production processes. Furthermore, historical facility data can be used to enrich the training data. Training samples are generated based on the recorded simulation data and, optionally, historical facility data. The training samples consist of: process values, setpoints, and control outputs at n time points as predictor variables, and KPI values, along with the next m values of one or more KPIs as dependent variables. A common configuration is to use several PVs, setpoints, and control outputs, along with multiple KPIs and the next m values of a single KPI. Here, m must be chosen to be large enough that the predictions encompass the relevant time range.
[0053] Training samples are used to train machine learning algorithms, such as recurrent neural networks or ANNs. The trained algorithm, such as a recurrent neural network with trained weights, becomes a surrogate model or part of it.
[0054] Figure 7 An example of using a proxy model for predicting alarms, according to an embodiment, is illustrated schematically. The proxy model is fed, for example, the same setpoint from a DCS (Distributed Control System), the past n values of control outputs, and process values and KPI (Key Performance Indicator) values used during training, and generates an m-step trajectory. The KPI values may originate from the process and / or the DCS. Optionally, the m-step trajectory is analyzed by alarm logic to check for violations of relevant target values, and if so, an alarm is presented on the corresponding human-machine interface (e.g., in an alarm list). Alternatively, future KPI trajectories can be presented to the operator.
[0055] Figure 8The diagram schematically illustrates the intervals of an alarm management system corresponding to the time evolution of process variables according to an embodiment. An ANN system, sometimes referred to as Alarm Intelligent Deferment (AID), can proactively manage events before they cause alarms. AID can reduce the number of alarms issued to the Product Owner (PO) while controlling the timing of alarm arrivals at the PO. The goal of AID is to reduce operator workload and minimize human error. PVs typically operate in environments such as... Figure 8 The system is positioned within the safe zone shown. The Alarm Management System (AMS) can maintain a multi-zone alarm system to prevent hazardous operating conditions from being reached. If a process variable leaves the safe zone, an alarm (visual or audible warning) is triggered to the operator, who must intervene manually by manipulating the setpoint. If the process enters a hazardous zone, hard-coded instrumented safety systems, such as safety valves, are triggered. If the process variable (PV) cannot be controlled by these means, it enters a destructive zone, where a complete system shutdown may occur, resulting in significant economic costs and safety hazards. The purpose of the Alarm Management System (AID) is to keep the PV within the safe zone to the greatest extent possible and allow only a limited number of events to reach the facility operator (PO). Human operators can then comfortably apply their experience and knowledge without being overwhelmed by multiple alarms. In this way, few events (if any) will escape the alarm zone and lead to more dangerous situations. The AID will continuously grow its knowledge base as it monitors the actions of the PO, such as resolving unseen situations and running more simulations to expand its understanding of alarm events.
[0056] Figure 9 An interval or alarm limit according to another embodiment is illustrated schematically. These intervals include typical thresholds in alarm management located between critical upper and critical lower limits. Alarm thresholds can be defined backward from the critical upper limit, giving operators sufficient time to react. Defining thresholds can be a challenging task, especially in terms of giving operators sufficient time to react. Furthermore, statically defined thresholds do not reflect the current state and dynamics of the facility. If a response time is defined, a predictive model can be used to determine whether a process value will exceed any threshold within a time window, thus giving operators sufficient time to react to the alarm. If the predictive model is a dynamic model of the facility, the alarm will also be able to interpret the current state and dynamics of the facility.
[0057] Figure 10The training process of a machine learning model according to an embodiment is illustrated schematically. The training phase can be initiated offline and / or before using the machine learning model. Training data can be generated from multiple sources. The first source may include simulations of events and operator actions and the corresponding results of these actions. The second source may include monitored data from real processes. Similar data formats are collected, more specifically: "setpoint" plus "past PV time series" and the corresponding "future PV time series". The time series is defined as PV over time steps {PV...} t-τ ,PV t-τ+1 ,…,PV t The evolution of}, and the setpoint constitutes the manipulator accessible to the PO, {MV t+Τ The combination of PV and MV produces PV{PV} for the advance fmv(T). t+1 ,PV t+2 ,…,PV t+T The evolution of}.
[0058] Training data is collected in a knowledge base and used by an ML-training system that trains two ML modules: (1) ML Delay Module: This module is trained using the corresponding data if there is a combination of setpoints (MVs) that lead to a feasible solution to a problem at a given time head. The module outputs setpoint actions for a given time series that preserve the future evolution of the PV within a safe interval. Setpoint actions are ranked from most effective to least effective based on their KPIs. (2) ML Delay Module: This module is trained using the corresponding data if there is no combination of setpoints in the knowledge-based data that provides a feasible solution (preserved within a safe interval) for a given lead time. Different setpoint actions are ranked based on the time delay they can be attached to the PV before leaving the safe interval (compared to no action). There may not be any setpoint actions in the current knowledge base that can add a delay to the PV.
[0059] Human operators can be employed to guide ML training and facilitate its implementation. The primary input for the PO (Product Owner) is specifying which MV (Moment of Performance) is best suited for each event or PV (Performance Value). Furthermore, the PO can suggest approximate setpoint actions, such as quantifying the MV based on empirical knowledge, to aid further ML training. The ML training module prompts the PO through queries in a graphical user interface to facilitate interaction. Queries are prompted during periods of low or zero cognitive load for the operator. This allows the ML training system to significantly limit its exploration space and provides a good initial condition for the training of the ML module.
[0060] Figure 11The operational phases of an alarm management system according to an embodiment are schematically illustrated. Multiple PVs of interest can be continuously monitored. This data is fed into a prediction system. The prediction system assesses the probability that any PV will leave the safe zone and trigger an alarm within a predefined timeframe. If the probability exceeds a threshold, the ML module is invoked to address the impending event.
[0061] If a feasible solution exists for this event, the ML delay module is invoked, which initiates all corresponding actions to keep the PV within a safe range. If the actions taken are successful, the prediction system stops issuing alarm predictions.
[0062] If the ML Deferred module's knowledge base does not contain a feasible solution to successfully resolve the predicted event, the ML Delay module is invoked. The ML Delay module attempts to insert a time buffer before issuing the actual alarm. If the Product Owner (PO) is under heavy cognitive load (the PO is already processing multiple alarms for other events), the ML Delay module selects the maximum feasible delay. If the PO is under low or zero cognitive load, the module selects a small or zero delay accordingly.
[0063] If no feasible solution exists and no delay can be added to the evolution of the predictive alert, the Product Owner (PO) is appropriately notified. The PO's actions in resolving the alert are recorded and the knowledge base is expanded accordingly to enable autonomous resolution of events in the future, thereby enhancing AID's problem-solving capabilities.
[0064] Figure 12 The illustration schematically depicts an example scenario where prediction-based alerts, according to an embodiment, are beneficial. It illustrates four different process value trajectories:
[0065] (a) The hypothetical PV trajectory used to define the static threshold. The threshold selection method allows operators sufficient time to respond to alarms and prevent PV from exceeding the critical threshold.
[0066] (b) The PV increases faster than assumed. Alarms will activate too late, and operators will not have enough time to respond.
[0067] (c) Contrary to (b), the PV increases more slowly than assumed. The alarm is activated too early. This is also undesirable according to alarm management standards. Operators may ignore the alarm.
[0068] (d) The trajectory changes after exceeding the static threshold and never exceeds the critical threshold. No alarm is needed.
[0069] Figure 13The surrogate model training workflow according to an embodiment is illustrated schematically. A first-principles dynamic facility model is used to create data, which is recorded for use as training data. Simulation is used because operators would avoid reaching critical thresholds during production and training data would be insufficient. However, in simulation, such training data can be generated without negatively impacting production. However, historical facility data can be used to enrich the training data. Training samples are generated based on the recorded simulation data and optional historical facility data. The training samples consist of: process values, setpoints, and control outputs at n time points as predictor variables, and the next m values of one or more process values as dependent variables. A common configuration is to use several PVs, setpoints, and control outputs, and the next n values of a single PV. Here, m must be chosen large enough that the prediction includes the relevant time range, i.e., longer than the reaction time. The training samples are used to train a machine learning algorithm, such as a recurrent neural network. The trained algorithm, such as a recurrent neural network with trained weights, becomes the surrogate model.
[0070] Figure 14 An example of using a proxy model for predicting alarms, according to an embodiment, is illustrated. The proxy model is fed, for example, the same setpoint from the DCS, the past n values of the control output, and process values used during training, for example, from the process (technically also possibly from the DCS), and generates an m-step trajectory. The number of m steps corresponds to the time required to respond to the alarm. The alarm logic analyzes the m-step trajectory to check if a relevant threshold has been violated; if so, the alarm is presented on the corresponding human-machine interface, for example, an alarm list.
[0071] In the variant, the alarm logic can still be evaluated based on a static threshold, but the analysis considers whether (a) there is still sufficient time to respond if the alarm is issued at the static threshold, (b) there is more time than required for a response, or (c) the PV may not have reached the critical threshold. In case (a), the logic might activate the alarm earlier than usual and provide causal information to the HMI, such as the expected trajectory. In case (b), the logic might not suppress the alarm but instead add additional information to the HMI indicating that there is still time to respond, such as the expected trajectory. In case (c), the logic might not suppress the alarm but instead add additional information to the HMI indicating that a response may not be necessary at all, such as the expected trajectory.
[0072] Figure 15A method of using online simulation according to an embodiment is illustrated schematically. For this purpose, current facility data is fed into the system. (0) In the first step, the state of the facility can be estimated, and the estimated state is used to (1) initialize the simulation or the simulation initialized by feeding the facility data into the simulation until the simulation converges with the true facility state. This true facility state may be past, ideally at the time of the disturbance, or even earlier. Next, (3) a disturbance curve is selected. The selection may be based on readings from the facility, for example by excluding certain unlikely disturbances. It is also possible to select multiple disturbance curves for the same simulation run. (4) The simulation is run using the disturbance curves. Steps (3) and (4) may occur in concurrent execution. The data generated by the simulation, i.e., the process value PV and the setpoint value SP (or MV), will be matched with the actual facility data. Matching can occur through appropriate distance measurements, such as Euclidean distance, dynamic time warp (DTW), Jaccarq, Levenshtein, based on correlation, autocorrelation, etc. If the match shows a sufficiently small distance, the possible root cause is (5) presented to the user. Alternatively, the disturbance curve with the smallest metric will be presented. If, according to expert or machine learning definitions, countermeasures are known (methods) for certain types of interference, the system can (6) recommend these measures to the user or directly trigger the execution of actions. The simulation model may not be a first-principles model, but rather a proxy model used to meet the real-time requirements of online simulation.
[0073] Figure 16 This illustration schematically depicts an example of root cause analysis using data from simulation and machine learning, according to an embodiment. Variations may not directly use simulation and / or surrogate models and disturbances. Instead, a machine learning model can be trained to identify potential root cause disturbances. The process is broken down into two steps: training and root cause analysis (RCA).
[0074] During training, the simulation is performed using a large number of disturbance curves and combinations thereof. The simulation generates training data using predictors—such as process values, setpoints, alarms, and events—and the disturbance curves used during the simulation (either as continuous signals or simply as disturbance identifiers). In the second step, the disturbance information is used as labels to train a machine learning classifier, or a machine learning regression to reproduce the disturbance curves. The resulting model is then used for the RCA task.
[0075] During RCA, the RCA is requested by the operator or a monitoring system (e.g., an anomaly detection system). Data collected from the facility is fed into a machine learning model. The output can then be presented to the operator as a possible root cause. If, according to expert or machine learning definitions, countermeasures (methods) are known for certain types of interference, the system can (4) recommend these measures to the user or directly trigger the execution of actions.
[0076] Variants could include first trying and evaluating actions from the perturbation method on the surrogate model. For example, using Bayesian optimization, the process of the actions—such as timing, sequence, setpoints, etc.—can vary in the optimization loop, and the actions can be optimized based on objective time, such as minimizing execution time or maximizing throughput during execution.
[0077] Variants can be implemented in the deployment of digital twins of facility processes. Digital twins digitally replicate facility processes using model-based dynamics. However, these are only approximations of the actual processes and can slowly deviate from what is happening in the facility. Standard practice is to synchronize the digital twin with the physical facility processes using measurements from them. However, these measurements are not always sufficient to distinguish between different internal states that might produce the exact same measurements. This can have different impacts on the future evolution of the facility. The digital twin may not run at a single instant that matches the actual facility state, but rather multiple possible scenarios weighted according to some probability. These scenarios can be maintained, discarded, or reweighted using ML models. Whenever a disturbing feature is detected, some instances running in parallel will be discarded or reweighted. Furthermore, if no internal state matches the ML model, the ML may need to augment its training data using relevant scenarios from the digital twin.
[0078] Figure 17 A flowchart 80 of the method according to an embodiment is schematically shown. In step 81, the machine learning model 10 (see...) Figure 1The process is trained using input data 20 and score data 30. The machine learning model 10—also known as an ML model—is an artificial neural network (ANN). Input data 20 includes a first time series 21 of at least one observable process value PV of the industrial process 50, a second time series 22 of at least one manipulated variable MV affecting the industrial process 50, and a third time series 23 of at least one internal variable IV of the industrial process 50. Score data 30 includes a first critical value 32 for each of the at least one observable process value PV indicating anomalous behavior of the industrial process 50, and a fourth time series 34 of at least one predicted observable process value PPV of the industrial process 50. In step 82, the trained machine learning model 10 is run by applying the first time series 21. In step 83, an output value 40 is output by the trained machine learning model 10, which includes at least a second critical value 42 of at least one predicted observable process value PPV indicating anomalous behavior of the industrial process 50 at a predefined time interval T1.
Claims
1. A method for detecting anomalous behavior in an industrial process (50), the method comprising the steps of: A machine learning model (10) is trained using input data (20) and score data (30), wherein the machine learning model (10) is an artificial neural network (ANN), and the score data (30) is the reward and / or penalty of the artificial neural network (ANN). The input data (20) mentioned above includes: The first time series (21) of at least one observable process value of the industrial process (50) that can be observed from the outside. A second time series (22) for controlling at least one manipulated variable of the industrial process (50), the at least one manipulated variable being input by an operator, service personnel, and / or one or more automated processes, and The third time series (23) of at least one internal variable of the industrial process (50) that cannot be directly observed and / or can be observed through simulation. And the score data (30) mentioned therein includes: The first critical value (32) of each of the at least one observable process value indicating the abnormal behavior of the industrial process (50), and A fourth time series (34) of at least one predicted observable process value of the industrial process (50). The trained machine learning model (10) is run by applying the first time series (21), the second time series (22), and the third time series (23) to the trained machine learning model (10); and The trained machine learning model (10) outputs an output value (40), which includes at least a second critical value (42) of at least one predicted observable process value indicating anomalous behavior of the industrial process (50) within a predefined time interval.
2. The method according to claim 1, wherein the output value (40) further comprises: The scene number (46) of the industrial process (50) depends on at least one of the first time series (21), the second time series (21) and / or the third time series (23).
3. The method according to claim 1, wherein the output value (40) further comprises: A fifth time series (44) that depends on at least one of the first time series (21), the second time series (21) and / or the third time series (23).
4. The method according to claim 1, wherein the output value (40) further comprises: The first critical value (32) of the at least one observable process value.
5. The method according to claim 1, further comprising the following step: The output depends on the manipulated variable of at least one of the first time series (21) and / or the third time series (23).
6. The method according to claim 1, further comprising the following step: Determine the time distance to the second critical value (42) that exceeds the predefined critical value.
7. The method of claim 1, further comprising the following step: Determine the rate of increase of the second critical value (42); and An alarm is output when the rate of increase exceeds a predefined threshold.
8. A computer program product comprising instructions that, when executed by a computer and / or an artificial neural network (ANN), cause the computer and / or the ANN to perform the method according to any one of claims 1 to 7.
9. A computer-readable storage medium on which the computer program of claim 8 is stored.
10. A machine learning model (10) configured to perform any one of claims 1 to 7.
11. Use of machine learning models (10) in monitoring and / or controlling industrial processes (50).
12. An industrial facility comprising a computer and / or an ANN, wherein a computer program product comprising instructions is stored on the computer and / or the ANN, and when the program is executed by the computer and / or the ANN, the instructions cause the computer or the industrial facility to perform the method according to any one of claims 1 to 7.
Citation Information
Patent Citations
System and method for dynamic multi-objective optimization of machine selection, integration and utilization
US20090210081A1
Process and system for propagating cell cultures while preventing lactate accumulation
WO2019100040A1