Industrial Plant Monitoring
By employing time series residual data and data-centric algorithms for sensor grouping and predictive modeling, the method effectively reduces false positives and enhances anomaly detection in industrial plants, enabling proactive maintenance and reducing unplanned shutdowns.
Patent Information
- Application Number
- JP2022547299
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2020-02-04
- Filing Date
- 2021-01-29
- Publication Date
- 2025-05-20
- Estimated Expiration
- 2041-01-29
AI Technical Summary
Industrial plants face challenges in efficiently monitoring sensor outputs to detect anomalies due to sensor malfunctions or equipment deterioration, leading to unplanned shutdowns and increased maintenance costs, with existing methods often resulting in false positive events and inefficient maintenance scheduling.
A method involving the use of time series residual data from sensor objects, generated by a group of sensors, to monitor level and relevant signals, and generate anomalous event signals when these signals deviate from expected values, utilizing data-centric algorithms like self-organizing maps for automatic sensor grouping and predictive models to identify equipment anomalies.
This approach reduces false positive alerts, enhances anomaly detection sensitivity, and allows for proactive maintenance planning, thereby minimizing unplanned shutdowns and optimizing maintenance schedules in industrial plants.
Smart Images

Figure 0007680458000001 
Figure 0007680458000002 
Figure 0007680458000003
Abstract
Description
[Technical field]
[0001] Technical Field The present teachings relate generally to computer-based monitoring of industrial plants. [Background technology]
[0002] Background technology An industrial plant, such as a process plant, comprises equipment operated to produce one or more industrial products. The equipment may be, for example, machinery and / or heat exchangers requiring maintenance. The requirements for maintenance may vary depending on several factors, including the operating time and / or load of the equipment, the environmental conditions to which the equipment is exposed, etc. An improper or unplanned shutdown of equipment is generally undesirable, as it often leads to a stop in production, reducing the efficiency of the plant and causing waste. It may be difficult to plan the shutdown of equipment around the time that maintenance is actually required, as the period between two maintenances may vary. Thus, scheduled maintenance may be performed earlier than actually required, or the operation of the equipment may exceed the maintenance period. The latter may affect the life of the equipment and / or cause inefficient operation. As will be recognized, the latter may also increase the risk of unplanned shutdowns, which may cause waste of materials that could not be processed by the equipment due to the unplanned shutdown. Although the risk of unplanned shutdowns can be reduced using the former approach, the approach is not always desirable, as it may lead to more frequent maintenance and increased costs.
[0003] The plant also includes a number of sensors for measuring or detecting one or more parameters associated with the equipment. Some sensors may also require their own maintenance, such as preventative maintenance and / or calibration, to ensure reliability in measuring and / or detecting the parameters they are supposed to measure or detect.
[0004] In many cases, a change in the output of a sensor can also indicate the health of the equipment it is measuring. However, a change in the output can also be due to a malfunction of the sensor itself, rather than a deterioration in the health of the equipment. Plants often contain hundreds or thousands of sensors. Large industrial plants can contain tens of thousands of sensors, or even more. Thus, it can be difficult to obtain an indication of the health of the equipment by monitoring each sensor. Furthermore, even if the sensor output drifts, it can be difficult to determine the state of the equipment. As a result, false positive events can be triggered. Frequent false event signals and alarms can reduce the usability of such systems.
[0005] Therefore, there is a need for a method for monitoring a plant that more efficiently utilizes sensor outputs to detect anomalies in the plant. Summary of the Invention [Means for solving the problem]
[0006] overview It is indicated that at least some of the problems inherent in the prior art are solved by the subject matter of the accompanying independent claims.
[0007] From a first perspective, there can be provided a method for monitoring a plant including a plurality of sensors and one or more functionally connected processing units, the method comprising:
[0008] - providing, at any of the one or more processing units, time series residual data for a sensor object, the sensor object being a group of at least some sensors from a plurality of sensors, the residual data comprising, for each of the sensors of the sensor object, a residual signal that is a difference between a measured output of the sensor and a predicted output of the sensor; - monitoring a level signal via any of the one or more processing units, the level signal being indicative of collective time-based variations in the time series residual data; and - monitoring, via any of one or more processing units, a relevant signal, the relevant signal indicating a variance and / or a relevant structure of the time series residual data; - generating, via one or more processing units, an anomalous event signal when, at a particular time, the value of the level signal and / or the value of the associated signal change from an expected value of the respective signal at or around that time.
[0009] The last step can alternatively be expressed as: - generating, via either one or more processing units, a level event signal and / or a related event signal, where a level event signal is generated when at a particular time the value of the level signal changes from an expected value of the level signal at or around that time, and a related event signal is generated when at a particular time the value of the related signal changes from an expected value of the related signal at or around that time.
[0010] The occurrence of the level event signal and / or the associated event signal may be considered as an abnormal event by the processing unit. In other words, the level event signal and / or the associated event, or more generally the abnormal event signal, is indicative of an abnormality of at least one device in the plant. Either or both of the values may be time-dependent.
[0011] Those skilled in the art will appreciate that the term "time-dependent" in this disclosure refers to such values or parameters that may change over time. Such values may not be directly dependent on time; rather, due to the time-varying nature of the signals and process parameters associated with a plant that are time-series data, such values may be expressed or calculated as a series of discrete or continuous values along a time scale, or time-series values. It will also be appreciated, therefore, that it is not necessary that such values must always or periodically change over time.
[0012] In some cases, the final step may further be: - generating, via either one or more processing units, an anomalous event signal when the magnitude of the residual signal exceeds a residual threshold at a particular time, whereby the value of the level signal and / or the associated signal changes from an expected value of the respective signal at or around that time.
[0013] In some cases, the residual data may be provided as input data at one or more functionally connected processing units. In such a case, the first step may specifically be as follows:
[0014] - receiving, at any one of one or more processing units, time series residual data for a sensor object, the sensor object being a group of at least a portion of a plurality of sensors, the residual data including, for each sensor of the sensor object, a residual signal that is a difference between a measured output of the sensor and a predicted output of the sensor.
[0015] The residual data may be received directly at one or more inputs of the one or more processing units, or the residual data may be received at a computer memory operatively connected to any of the one or more processing units.
[0016] In some cases, the residual data may even be generated by one or more functionally connected processing units. In such a case, the first step may specifically be as follows:
[0017] - generating, at any one of one or more processing units, time series residual data for a sensor object, the sensor object being a group of at least some sensors from a plurality of sensors, the residual data including, for each sensor of the sensor object, a residual signal that is a difference between a measured output of the sensor and a predicted output of the sensor.
[0018] In some cases, the residual data may also be monitored, as described below. In such cases, the method may also include:
[0019] - monitoring time series residual data of the sensor object via one or more processing units;
[0020] The monitoring of the residual data can be done by monitoring each residual signal, i.e., for each sensor or any one or more residual signals. Thus, for any sensor whose magnitude of the residual signal exceeds a predefined value or threshold at any particular time, a residual event can be recorded at that time in a memory location operatively connected to any of the one or more processing units. If the level signal and associated signal are as expected, then no action is taken, e.g., generating an alarm or recording an anomaly. On the other hand, if the level event signal and / or associated event signal occurs at or around a particular time when a residual event from one or more sensors occurs. In such a case, the residual event signal from one or more sensors can be analyzed to find the cause or root cause of the anomaly. By at or around a particular time, it is meant that the residual event signal may have occurred simultaneously with, prior to, or after the occurrence of the level event signal and / or associated event signal in order to be considered for the analysis. It will be appreciated that rather than measuring the magnitude of the residual, the monitoring of the sensor signal can even be based on the magnitude of the measured signal of any sensor that exceeds a particular threshold from the expected value for that sensor. Thus, both cases are considered equivalent and may be used interchangeably in the present disclosure, e.g., with reference to the threshold.
[0021] A "sensor object" may refer to a specific group of at least some sensors from a plurality of sensors of a plant. In other words, a sensor object is a group of sensor signals whose residual signals are collectively monitored in the form of level signals and / or related signals. In large production plants, such as chemical and / or biological plants, the number of sensors monitored is large. Typically, the number of sensors in a chemical or biological plant can be well over 1000, and in many cases there can be tens of thousands of sensors, and in some cases hundreds of thousands or more sensors. In chemical and / or biological plants with complex production and value chains, such as multiple industrial plants or barband arrangements, the number of sensors can be enormous.
[0022] The proposed sensor object allows for a more effective application of the multivariate techniques disclosed herein. The applicant recognized that if the multivariate techniques are implemented on a group including all the sensors of the plant or on a suboptimal group of sensors, such a group may lack sensitivity with respect to the anomalies detected. Small but important deviations of some sensors may be over-controlled or dominated by large but less important deviations of other sensors in the group. The present teachings also allow for the definition of an appropriate group of sensors in the form of one or more sensor objects, each of which may allow for maintaining sensitivity to anomalies without being overcome by other sensor signals.
[0023] Compared to univariate methods, multivariate monitoring methods may have the advantage of reducing the false positive rate in generating alerts indicating anomalies in the sensor object. Thus, the sensor objects disclosed herein may be capable of maintaining low false positives while achieving high sensitivity for anomaly detection.
[0024] Clustering with some of the sensors from multiple sensors into one or more sensor objects can be done in many ways as mentioned above. However, suboptimal clustering or grouping of sensors in sensor objects can affect the sensitivity of detection. Furthermore, manual clustering to generate sensor objects can take a long time without guarantee of success. Even with experts, manual grouping of sensors with manual input can result in a huge amount of manual work, even when performed at a large scale. Since each plant can be unique in itself, and therefore many unknowns may be involved, it can be difficult to perform a proper grouping. Furthermore, often the information about the plant topology is not provided in a processable data format. For example, mathematical models of the plant that can be processed or evaluated are not available, at least not at the required level of detail.
[0025] According to a preferred embodiment, the sensor object is provided by at least partially automatically selecting at least some of the sensors from the plurality of sensors. The selection is based on the suitability of at least some of the sensors to be grouped in the sensor object. According to an embodiment, similarity detection using a data-centric algorithm is used to build the sensor object. Preferably, the selection is fully automatic. As explained, more specifically, the selection of at least some of the sensors, or the grouping of the sensors, is performed by one or more data-centric methods or algorithms. Preferably, the data-centric algorithm is configured to use at least one similarity measure to group the sensors from the plurality of sensors into the sensor object. For example, the data-centric algorithm uses at least one similarity measure to group at least some of the sensors to provide the sensor object. Here, by data-centric algorithm is meant an algorithm configured to leverage sensor data, such as historical time series data of the plurality of sensors, to at least partially automatically group or select at least some of the sensors. More specifically, the data-based algorithm may be one or more clustering algorithms. The clustering algorithm may be an unsupervised learning algorithm, such as a neural network model trained by unsupervised learning using historical data of multiple sensors, or any other suitable algorithm for clustering and dimensionality reduction. The unsupervised learning algorithm may be, for example, a self-organizing map ("SOM"). Applicants have found that SOMs are particularly useful for automatically grouping suitable sensors into sensor objects.
[0026] Thus, according to one aspect, using one or more self-organizing maps, any one or more processing units are configured to arrange or order sensors in a computer memory according to similar patterns or one or more similarity measures in their time-dependent signals. For this purpose, historical time series data from the corresponding sensors can be utilized. The sensors can be arranged in a matrix or computer-readable 2D map space. In a next step, this matrix or map is subdivided or fragmented to generate one or more sensor objects. For example, the map space can be fragmented or cut using a symmetric or asymmetric grid. Additionally or alternatively, the map space is cut using distance values around clusters of sensors on the map. For example, clusters can be automatically detected via a similarity measure that includes selecting sensors that are within a certain distance from each other. All sensors that are within the distance value can then be grouped into a sensor object. The distance value measurement can also include multiple distance values, for example to capture sensors that form asymmetric clusters. To better capture asymmetric clusters, occasionally sub-clusters within a cluster can be detected based on one or more distances. In some cases, entire multiple sensors are divided into sensor objects as described. A sensor object may have two or more sensors. In a preferred embodiment, a sensor object may have 20 or a few sensors, but not less than two sensors. The population of sensors in a sensor object may vary. For example, some sensor objects may include more than 20 sensors, for example about 100 sensors. According to a preferred embodiment, a sensor object includes signals from 20 or about 20 sensors. According to a more general preferred embodiment, a sensor object includes signals from 10, about 10, or tens of sensors. The term "tens of sensors" here means including any integer number of sensors equal to or less than 100, for example 4, 12, 25, or 30 sensors. The proposed automatic clustering can capture sensors suitable for inclusion in a sensor object, for example by enabling automatic detection of similarities between sensors.Anomaly detection can therefore be performed without expert users and with little or no specification of the sensors or plant topology – which can be a key advantage for complex and large plants such as chemical and biological plants.
[0027] According to one embodiment, the SOM is trained using unsupervised learning with sensor data from multiple sensors to generate a two-dimensional discretized representation of input vectors, in this case sensor data that needs to be clustered into sensor objects, i.e., data from multiple sensors. In accordance with the present teachings, those of input vectors or sensor signals that are similar in high-dimensional space are mapped to nearby nodes in two-dimensional ("2D") space. Similarity can be measured in terms of distance between sensor nodes mapped in the 2D space.
[0028] As a non-limiting example of automatic grouping using SOM, the 2D space can be predefined or specified, for example, by representing its geometry as a k*n grid. The sensor nodes can be initially randomly mapped into the 2D space, and then their positions are iteratively adjusted. In this way, adjacent points in the initial geometry of the input vector can be mapped to nearby points in the 2D space.
[0029] According to the present teachings, by selecting an appropriate distance / similarity measure for the sensor time series data, one or more SOMs can be used to cluster the time series into groups of matching shapes, where the group represents a sensor object, as should be understood by those skilled in the art.
[0030] It will be appreciated that the processing units need not be located at the same location or in the same physical location. For example, at least some of the processing units may be implemented as or in a cloud service. Because each of the processing functions for monitoring the level signals and related signals provides at least one technical advantage, such functions are also patentable in their own right.
[0031] Therefore, from another perspective, there can be provided a method for monitoring a plant including a plurality of sensors and one or more functionally connected processing units, the method comprising:
[0032] - providing, at any of the one or more processing units, time series residual data of a sensor object, the sensor object being a group of at least a portion of a plurality of sensors, the residual data comprising, for each sensor of the sensor object, a residual signal that is a difference between a measured output of the sensor and a predicted output of the sensor; - monitoring a level signal via any of the one or more processing units, the level signal being indicative of collective time-based variations in the time series residual data; and - generating, via any of the one or more processing units, a level event signal, the level event signal being generated when a value of the level signal at a particular time changes from an expected value of the level signal at or around that time, the level event signal indicating an anomaly in at least one piece of equipment within the plant.
[0033] Similarly, from yet another perspective, there may be provided a method for monitoring a plant including a plurality of sensors and one or more functionally connected processing units, the method including:
[0034] - providing, at any of the one or more processing units, time series residual data of a sensor object, the sensor object being a group of at least a portion of a plurality of sensors, the residual data comprising, for each sensor of the sensor object, a residual signal that is a difference between a measured output of the sensor and a predicted output of the sensor; - monitoring, via any of one or more processing units, a relevant signal, the relevant signal indicating a variance and / or a relevant structure of the time series residual data; - generating, via any of the one or more processing units, a related event signal, the related event signal being generated when a value of the level signal at a particular time changes from an expected value of the related signal at or around that time, the related event signal being indicative of an anomaly in at least one piece of equipment within the plant.
[0035] As previously discussed, preferably the sensor objects are provided at least partially automatically by grouping at least some of the sensors.
[0036] As will be appreciated, in any of the above cases, rather than generating an event indicating an anomaly based on the residual signal, i.e., when the measurement signal of the sensor exceeds a certain threshold from the predicted value of the sensor signal, the present teachings postpone the generation of the anomaly event signal based on monitoring of the level signal and / or the related signal. The anomaly event is generated when either or both of the level signal and the related signal change from their respective predicted values. Thus, in other words, the generation of the anomaly event signal is prevented when either of the residual signals exceeds their residual threshold while both the level signal and the related signal are within their respective predicted values, and thus the anomaly event is based on the result of monitoring the level signal and / or the related signal.
[0037] According to one embodiment, the respective predicted values of the level signal and the associated signal are each preferably provided as a respective range of values that these respective signals can effectively be. Thus, each of the level signal and the associated signal may be provided with at least one limit value at any particular time. If the signal value at that time is within the respective limit, the value is considered as predicted and no anomalous event is generated. In other words, any of the respective predicted values may be provided as a limit of the corresponding predicted value, specifying at a particular time a number of predicted values that the corresponding signal can effectively have without an anomalous event being generated. The limit values may also be time-dependent values. It will be understood that the multiple predicted values may be discrete values or may correspond to any values that the respective signals can take within the ranges specified by the corresponding limit values.
[0038] Applicants have discovered that the proposed event generation based on monitoring of levels and associated signals allows for a significant reduction in false positive events while focusing on detecting actual anomalies.
[0039] In response to the abnormal event signal, an alarm may be generated to notify an operator or user of the abnormality in the equipment with which the sensor object is associated. Alternatively, or in addition, any of the one or more processing units (hereinafter simply referred to as processing units) may backtrack the sensor data to determine the cause of the abnormality. Thus, according to one aspect, the method also includes:
[0040] - Determining the health of equipment associated with the sensor object in response to the abnormal event signal.
[0041] The health of the equipment may be determined, for example, by performing a root cause analysis via the processing unit.
[0042] As outlined above, in some cases, limits are provided for changes or deviations from expected values. If the signal value is within the relevant limit value, no event signal is generated. Thus, if any of the monitored signals, levels, or associations of the sensor object deviate beyond a predefined limit or control range, this deviation can be considered an anomaly and an alert can be generated in the form of an event signal (e.g., a level event signal, an associated event signal, or both). As proposed, the state of the sensor object is monitored or observed using the level signal and / or the associated signal, both of which are time-dependent or time-based signals. It will be appreciated that the specific limits can be defined according to the application of interest, for example, the acceptable range of deviations, the required sensitivity of the event, the criticality or importance of the sensor object, etc. According to one aspect, the limit values are derived using statistical limits within which the respective signals, levels and associated signals are expected to be under normal conditions. Thus, the limit values represent a range or probability space of values of the movement of the expected value or around which the monitored signal value can be validly placed. In some cases, there may be upper and lower limits, especially for the associated signal. Thus, a probability space that lies within the upper and lower bounds can be defined by a number of predicted values. In either case, the probability space can be determined, for example, using historical data from past monitoring. The bounds can thus be understood as defining a range of values within which an observation can be acceptably placed with a particular probability value. The bounds can be zero or non-zero values depending on the requirements of the application and / or the availability of historical sensor data. The bounds can also be time-dependent values. Thus, it is not necessary to give specific numbers to the bounds of the present disclosure.
[0043] In order to monitor the validity of the sensor data associated with the sensor object, the applicant has discovered that the above two indices, i.e., the level signal and the related signal, can be particularly effective in condensing the multivariate information from the sensor data into individual or integrated scores. The indices are applied to the residual signals of the sensor object. The residual signal generated for each sensor of the sensor object is the difference between the measured sensor output signal (or observed sensor signal) at a particular time and the predicted sensor signal at that time. The indices are then generated from the residual data, including the residual signals from the sensor object sensors. The applicant has recognized that the measured sensor output signal most often contains multiple pieces of information, most of which may be irrelevant for detecting anomalies or the health of the equipment. For example, the measured or observed sensor output signal may depend on the controller settings, the production mode, the operating conditions of the plant and or equipment, etc. Thus, information related to the health of the equipment may be drowned in extraneous information caused by various other parameters on which the sensor output depends. Applicant has realised that by not using the sensor output signal directly, but instead applying the proposed indices to the residual data, redundant information can be at least partially removed from the time-dependent sensor output, such that equipment health-related information becomes more detectable for further signal processing.
[0044] The first metric, the level signal, provides information or statistics related to the collective behavior of time-dependent residual signals in the time series residual data of the sensor object. The level signal values can be used to detect level changes or short-term trends in the time series residual data in the sensor object. The second metric, the association signal, provides information or statistics related to the fluctuations and / or association structure of the time series residual data. The association signal indicates changes in volatility and / or correlation structure between the time series residual signals in the sensor object.
[0045] An event signal for each signal, the level signal and / or the associated signal, can be generated when the magnitude of the signal value exceeds a predicted value or a certain limit of the predicted value for that signal at that time. An indicator is therefore compared at any particular time with respect to what can be called its predicted state at that time. The predicted state or the time-dependent predicted signal value can be provided by a model of the sensor object. The sensor object model can be at least partially a predictive model, such as a data-driven model, for example comprising a sensor object neural network trained using historical residual data. The results of the comparison, i.e. the deviation of the level signal from its predicted state and the deviation of the associated signal from its predicted state, are monitored via a processing unit over time. An abnormal event signal is generated when the magnitude of the level signal value and / or the associated signal value deviates beyond a limit from the predicted value of that signal at that time. The deviation is monitored via a processing unit over time.
[0046] The monitoring of the signal or even the deviation can be performed continuously or during discrete periods of equal or unequal length.
[0047] The sensor object may be provided in the processing unit, for example, via a memory operatively connected to the processing unit. According to one aspect, the generation of the residual data is performed via the same processing unit. Alternatively, the generation of the residual data may be performed by a separate processor and then provided in the processing unit.
[0048] An industrial plant, or simply a plant, includes infrastructure used for industrial purposes. The industrial purpose may be the manufacture or processing of one or more products, i.e., a manufacturing process or a process carried out by the plant. The product may be any physical product, such as, for example, chemical, biological, pharmaceutical, food, beverage, textile, metal, plastic, semiconductor, or the product may be a service product, such as electricity, heating, air conditioning, waste treatment, such as recycling, chemical processing, such as cracking or melting, or even incineration. Thus, the plant may be any or more of a chemical plant, a pharmaceutical plant, an oil and / or natural gas well, a fossil fuel processing facility, such as a refinery, a petrochemical plant, a cracking plant, and the like. The plant may further be any of a distillery, an incinerator, or a power plant. The plant may further be a combination of any of the above, for example, the plant may be a chemical plant including a cracking facility, such as a steam cracker, and / or a power plant. For the purposes of applying the present teachings, in some cases, a standby facility within a large plant may even be considered as a plant. The infrastructure may include equipment or process units, such as any one or more of the following: Heavy duty rotating equipment such as heat exchangers, columns such as fractionators, furnaces, reaction chambers, crackers, storage tanks, settling units, pipelines, stacks, filters, valves, actuators, transformers, circuit breakers, machinery, e.g. turbines, generators, crushers, compressors, fans, pumps, motors, etc.
[0049] A plant or industrial plant may further be part of a plurality of industrial plants. The term "multiple industrial plants" as used herein is a broad term and should be given its ordinary and customary meaning to those skilled in the art and should not be limited to a special or customized meaning. The term may specifically refer to, but is not limited to, a compound of at least two industrial plants having at least one common industrial purpose. Specifically, a plurality of industrial plants may include at least two, at least five, at least ten, or even more industrial plants that are physically and / or chemically bonded. A plurality of industrial plants may be bonded such that the industrial plants forming the plurality of industrial plants may share one or more of their value chains, extracts, and / or products. A plurality of industrial plants may also be referred to as a compound, a compound site, a bar band, or a bar band site. Furthermore, the value chain production of the plurality of industrial plants from various intermediate products to final products may be decentralized to various locations such as various industrial plants or integrated into a bar band site or chemical park. Such a barband or chemical park may be or may contain one or more industrial plants, and products manufactured in at least one industrial plant may serve as raw materials for another industrial plant.
[0050] Those skilled in the art will appreciate that plants also typically include instrumentation that may include several different types of sensors. Sensors are used to measure various process parameters or parameters related to equipment. For example, sensors can be used to measure process parameters such as flow rate in a pipeline, level in a tank, temperature in a furnace, chemical composition of gas, and some sensors can be used to measure turbine vibration, fan speed, valve opening, corrosion in a pipeline, voltage across a transformer, and the like. The differences between these sensors can be based not only on the parameters they sense, but also on the sensing principle they use. Examples of sensors based on the parameters they sense include temperature sensors, pressure sensors, radiation sensors such as optical sensors, flow sensors, vibration sensors, displacement sensors, chemical sensors such as those for detecting certain substances such as gas, and the like. Examples of sensors that employ different sensing principles include, for example, piezoelectric sensors, piezoresistive sensors, thermocouples, impedance sensors such as capacitive and resistive sensors, and the like. The exact parameters measured by the sensors or the principles used are not important to the generality of the present teachings.
[0051] Thus, plants are often equipped with sensors that continuously or periodically measure the value of certain quantities within the plant (e.g., temperature in a column, pressure in a pipe, mass flow rate, etc.). Each of these values is typically stored in a Plant Information System ("PIMS") along with additional information. The additional information can be one or more metadata such as the sensor name, a timestamp of the measurement, the measured value, units, measurement quality, etc.
[0052] In a plant, at least some of the sensor outputs are received directly or indirectly by a control system that controls at least some of the plant's operations. Some plants may further comprise multiple control systems that can be configured to operate in a hierarchy or in parallel. The exact architecture of the one or more control systems, the industrial control system ("ICS"), is not essential to the scope of generality of the present teachings. A plant also typically comprises a data acquisition system ("DAS") that receives data from multiple sensors in the plant. The sensor data is stored in a long-term computer memory or a database. The DAS may be the same system as the PIMS or may be a different system. The sensor data typically includes metadata that indicates the history or time information of the data collected from the sensor. Thus, the historical sensor data is typically stored in or can be restored from a database as time series data. Historical related and / or level signal data may also be stored in the same database or in a separate database operatively connected to the processing unit. As mentioned above, the metadata may also include units and / or labels of the sensor's time series data. Sensors are often also referred to as tags or sensor tags in industrial plants. Some further examples of plant control and / or monitoring systems include programmable logic controllers ("PLCs"), distributed control systems ("DCSs"), and supervisory control and data acquisition ("SCADA"). Furthermore, in some plants, the functions of any two or more of the above systems may be performed by a single control and / or monitoring system. Most control and monitoring systems today are digital systems; that is, they operate on digital signals, and thus sensor data received by such systems is also converted to digital signals, either by the sensor itself or at any stage before the analog signal is processed by the system's digital processor. Thus, the present teachings apply to any kind of plant monitoring and / or control system that includes one or more processing units.Some sensor outputs may depend on the state of a control system, for example, via a process parameter that is directly or indirectly influenced by the control system or controller. By integrating the residuals into the level and related signals and monitoring the residuals, such effects can be at least partially removed, thereby improving the detection of anomalies while reducing false positive detections.
[0053] The applicant has realized that the proposed method or its monitoring system allows for early detection of anomalies or the onset of abnormal operation of the plant before the problem is observable by an operator or a conventional monitoring system. Thus, it is possible to prevent unanticipated shutdowns of at least a portion of the plant equipment due to anomalies that appear over time. Such prevention can be achieved by using the present teachings to detect anomalies and plan maintenance so that interruptions to the industrial process can be at least reduced. It can also be provided to direct the operator's attention to specific plant areas or equipment where problems may become apparent later and corrective actions can be planned.
[0054] According to one aspect, in response to the occurrence of the abnormal event signal, the processing unit determines a state of health of at least one piece of equipment in the plant. According to a further aspect, in response to the abnormal event signal, the processing unit initiates further analysis including analyzing historical and / or real-time time series data of at least one individual sensor or a subgroup of sensors in the plurality of sensors. According to one aspect, the processing unit prioritizes further analysis of at least one sensor having one or more sensor outputs whose residual signals exceed a respective residual threshold or an associated residual event signal that occurred approximately simultaneously with the occurrence of the abnormal signal.
[0055] By directly responding to the abnormal event signal or by later analyzing data related to one or more sensors, the processing unit can more specifically determine which parts of the equipment or sensors may require maintenance. More preferably, by analyzing said time series signals over a certain period of time, the processing unit can predict a maintenance schedule for at least one piece of equipment. As will be appreciated, by doing so, the processing unit can provide future maintenance requirements related to the equipment and / or one or more sensors. Furthermore, by monitoring the proposed levels and related signals, the amount of sensor data that the processing unit needs to monitor and analyze in real time can be reduced. Especially under normal operation of the equipment or when no detectable anomalies are present, the resources used for monitoring can be significantly reduced compared to analyzing each sensor signal individually. As mentioned above, this can also have additional advantages in terms of false positives compared to univariate approaches, i.e., monitoring and analyzing each sensor individually.
[0056] According to one embodiment, in response to the level event signal, the processing unit analyzes the time series residual signal of each sensor in the sensor object to determine one or more major drivers or most dominant factors for the level signal value. The determination can be made using, for example, effect size calculations and / or value distribution analysis. An effect size is a measure of how far a residual signal is from a particular residual value within a particular time period. A particular residual value is typically the most likely value of that residual signal within a particular time period.
[0057] According to one aspect, in response to the associated event signal, the processing unit analyzes the time series residual signals of each sensor in the sensor object to determine one or more major drivers or most dominant contributors to the associated signal value. In a further aspect, the processing unit analyzes the covariance of the time series residual signals of each pairwise combination of sensor residual signals in the sensor object to determine one or more major drivers or most dominant contributors to the associated signal value.
[0058] The proposed teachings can leverage similarities and / or associations between residual data from sensors within a sensor object, focusing computational monitoring efforts on scenarios where data from one or more sensors deviate in a way that impacts the overall level and association of residual data within a sensor object. By grouping sensor residual data, it can be monitored according to the proposed multivariate approach. It will be appreciated that this frees up computational resources while still focusing on the overall behavior of related sensor data contained within a sensor object.
[0059] As described, according to one aspect, the grouping is performed at least partially automatically via the processing unit, for example, using a self-organizing map. The automatic grouping may be performed by the processing unit based on data characteristics, for example, based on similarity of sensor time series data of the sensor objects and / or interdependencies between output signals of the sensors, and / or based on the type of sensor, and / or similarity of sensor responses. Additionally or alternatively, according to one aspect, the grouping is performed based on input from a user. Thus, the grouping may be performed at least partially based on operator preferences or experience. In either case, automatic or user-guided, the proposed multivariate approach using level signals and associations has the advantage that interdependencies between residuals can be used to detect behavior that may not be detectable at individual signal levels. As previously described, false positives can be reduced compared to univariate monitoring approaches, i.e., monitoring data from each sensor individually while maintaining sensitivity for detecting anomalies in the sensor object.
[0060] Therefore, from another perspective, there can also be provided a method for monitoring a plant including a plurality of sensors and one or more functionally connected processing units, the method including:
[0061] - providing, at any of the one or more processing units, time series residual data of a sensor object, the sensor object being a group of at least a portion of the sensors from the plurality of sensors, the residual data comprising, for each sensor of the sensor object, a residual signal that is a difference between a measured output of the sensor and a predicted output of the sensor, the sensor object being provided by at least partially automatically grouping at least a portion of the sensors; - monitoring a level signal via any of the one or more processing units, the level signal being indicative of collective time-based variations in the time series residual data; and - monitoring, via any of one or more processing units, a relevant signal, the relevant signal indicating a variance and / or a relevant structure of the time series residual data; - generating an anomalous event signal when the value of the level signal and / or the value of the associated signals at a particular time, via either one or more processing units, changes from the predicted value of the respective signal at or around that time.
[0062] As already explained, the automatic grouping is performed using at least one data-centric algorithm, which may be a clustering algorithm, for example a SOM algorithm.
[0063] According to one aspect, a fully automatic grouping is performed first, and then parts of the grouping are rearranged according to user input, which can be useful for onboarding new plants and can reduce manual intervention.
[0064] According to one embodiment, a first group of sensor outputs from a first group of sensors in the plurality of sensors are configured or tagged as covariate signals. Additionally, a second group of sensor outputs, different from the first group of sensor outputs, are configured or tagged as monitored signals. The sensor object is realized using residuals from the second group of sensor outputs. The first group of outputs, or covariate signals, are signals representing parameters that may cause a change in the behavior of at least one of the monitored signals. Thus, the covariate signals preferably represent influential factors or parameters that may affect the monitored signals. The covariate signals may represent parameters such as ambient temperature, cooling water temperature, load, input flow, output flow, controlled measurements, or set screw position. Thus, the one or more monitored signals are preferably at least partially dependent on at least one of the covariate signals. The processing unit may automatically determine the covariate signals from the monitored signals by checking for interdependencies between the sensor outputs. Alternatively, at least some of the covariates and / or monitored signals may be defined based on user input. The advantages of integrated monitoring of sensor data can still be maintained.
[0065] According to one embodiment, data from the first group of output or covariate signals are also input to the sensor object model to generate predicted level signals and predicted related signals and / or their respective limits. Thus, the time-dependent values of the level signals and related signals are compared to the time-dependent values of the predicted level signals and predicted related signals. In this way, the applicant has realized that taking into account the covariates can make the predicted values and their limits more accurate, which can synergistically improve the detection of anomalies earlier while automatically taking into account factors that may affect the sensor data. Thus, the predicted values can also be adapted according to the covariate signals. Combining this with monitoring performed at the sensor object level can further result in a solution that requires reduced monitoring resources and reduced false positives. Thus, the reliability of the detection can be improved.
[0066] According to yet another aspect, the predicted output of the sensor is provided by a predicted state model, a machine learning ("ML") model. The predicted state model is a predictive model, e.g., a data-driven model, such as a predicted state neural network, that is trained using training data that includes, at least in part, the sensor's historical time-series output data.
[0067] Thus, the predictive state model may be or may include a predictive model that, when trained using training data including historical time series output data of the sensor as a machine learning ("ML") module, becomes a trained data-driven model. A "data-driven model" refers to a model derived at least in part from user training data, which may include historical data related to data, in this case the sensor. In contrast to a strict model derived purely using physicochemical laws, a data-driven model may describe relationships that cannot be modeled by physicochemical laws. A data-driven model may be used to describe relationships related to processes that take place, for example, within a respective manufacturing process, without solving equations from physicochemical laws. This may reduce computational power and / or increase speed.
[0068] The data-driven model may be a regression model. The data-driven model may be a mathematical model. The mathematical model may describe a relationship between the provided performance characteristic and the determined performance characteristic as a function.
[0069] Thus, in the present context, a data-driven model, preferably a data-driven machine learning ("ML") model or simply a data-driven model, refers to a trained mathematical model that is parameterized to reflect the reaction kinetics or physicochemical processes associated with the plant and / or one or more pieces of equipment according to a respective training data set, such as, for example, historical time series output data of a sensor. An untrained mathematical model refers to a model that does not reflect the reaction kinetics or physicochemical processes. For example, an untrained mathematical model is not derived from physical laws that provide scientific generalizations based on empirical observations. Thus, kinetic or physicochemical properties may not be inherent to an untrained mathematical model. An untrained model does not reflect such properties. Feature engineering and training using the respective training data set allows the parameterization of the untrained mathematical model. The result of such training is simply a data-driven model, preferably a data-driven ML model, that reflects the reaction kinetics or physicochemical processes associated with the respective plant and / or one or more pieces of equipment or assets of the plant, preferably only as a result of the training process.
[0070] The predicted state model may further be a hybrid model. A hybrid model may refer to a model that includes a first principles part (so-called white box) and a data-driven part (so-called black box) as previously described. The predicted state model may include a combination of white box and black box models and / or grey box models. The white box model may be based on physicochemical laws expressed, for example, as equations. The physicochemical laws may be derived from first principles. The physicochemical laws may include one or more of chemical kinetics, the law of conservation of mass, momentum and energy, particle populations of any dimension. The white box model may be selected according to the physicochemical laws governing the respective plant, its manufacturing process, or a part thereof. The black box model may be based on historical data, such as historical time series output data of sensors. The black box model may be built using one or more of machine learning, deep learning, neural networks, or other forms of artificial intelligence. The black box model may be any model that provides a good fit between the training data set and the test data. The grey box model is a model that combines partial theoretical structures and data to complete the model.
[0071] The trained model may include a serial or parallel architecture. In a serial architecture, the output of the white-box model may be used as the input of the black-box model, or the output of the black-box model may be used as the input of the white-box model. In a parallel architecture, a combined output of the white-box model and the black-box model may be determined, such as by superposition of the outputs. As a non-limiting example, the first sub-model may predict at least one of the performance parameters and / or at least some of the control settings based on a hybrid model with an analytical white-box model and a data-driven model acting as a black-box collector trained with the respective historical data. This first sub-model may have a serial architecture, where the output of the white-box model is input to the black-box model, or the first sub-model may have a parallel architecture. The predicted output of the white-box model may be compared to a test data set that includes a portion of the historical data. The error between the calculated white-box output and the test data may be learned by the data-driven model and applied to any prediction. The second sub-model may have a parallel architecture. Other examples are possible.
[0072] As used herein, the term "machine learning" or "ML" may refer to statistical methods that allow a machine to "learn" a task from data without explicit programming. Machine learning techniques may include "traditional machine learning," which is a workflow in which features are manually selected and then a model is trained. Examples of traditional machine learning techniques may include decision trees, support vector machines, and ensemble methods. In some examples, data-driven models may include data-driven deep learning models. Deep learning is a subset of machine learning loosely modeled on the neural pathways of the human brain. Deep refers to multiple layers between the input layer and the output layer. In deep learning, algorithms automatically learn which features are useful. Examples of deep learning techniques may include convolutional neural networks ("CNNs"), recurrent neural networks such as long short-term memory ("LSTM"), and deep Q-networks.
[0073] Alternatively, or in addition, similar to above, the sensor object model may be or may include a sensor object machine learning ("ML") module, a predictive model where training data including historical time series data becomes a trained sensor object data-driven model.
[0074] Thus, the predicted output of the sensor is generated by inputting data from at least one of the first group of outputs or covariate signals into the predicted state model. It will be appreciated that each sensor or tag may be provided with a respective predicted state model trained using historical data from that particular sensor. The covariate signals input into the predicted state model may be all covariate signals of the plant or may be a subset of the plant covariate signals. The plant may be considered as a closed system with expected and unexpected interdependencies between different equipment. For example, the external temperature of a furnace during operation may cause the local ambient temperature to be higher than other parts of the plant. This may therefore cause an increase in the temperature of a pipe close to the furnace, which in turn changes the density of the liquid flowing through the section of the pipe close to the furnace. Such interdependencies may manifest themselves in various ways in the operating parameters of the plant. In the strict sense, all covariate inputs may be required for the predicted state model and / or the sensor object model to get a complete picture of the interdependencies. However, this may not be practical due to processing power requirements, etc. Thus, preferably, the covariate input signals to the predictive state model and / or the sensor object model are a subset of all covariate signals of the plant. The latter can save the processing power used by the predictive state model. Thus, the predictive state model and / or the sensor object model can also be faster. To maintain the accuracy of the predictive state model and / or the sensor object model, the processing unit can determine the covariates that are dominant for the respective models. The dominant covariates are preferably a subset of all covariates and can be determined via the processing unit by analyzing the predictive power of each covariate signal of the respective model. Thus, if the variation of the covariate does not affect the output of the model, the covariate is prevented as an input to the model. The processing unit can use the historical time series data to analyze the predictive power of each covariate signal of the plant for each signal monitored.Thus, the processing unit can determine a predicted state model and / or a sensor object model using a selected subset of covariate inputs that have an observable effect on the model outputs.
[0075] Covariates can be analyzed not only by considering the respective sensor data that occurred at or near the same time, but also by analyzing them with an additional time lag. The time lag can be one or more periods between an occurrence in the covariate and the detection of the effect of that occurrence as the output of a particular sensor. This allows capturing interdependencies associated with delays or time constants. As an example, if the fuel input to the furnace is increased at time t, the temperature rise of the liquid heated by the furnace will be at time t+t d can be detected only at t d represents the time lag of the system. Applicant has recognized that this allows the most important covariate signals to be determined even if the most important covariate signals are inadvertently ignored and the effects of such covariates are not analyzed with lag.
[0076] According to one aspect, the processing unit selects the type of the predictive state model by analyzing which model type provides the smallest error between the predicted output and the actual output. For example, the error can be measured by calculating any one or more of an absolute error value, a mean squared error value, a weighted mean squared error value, or even a combination thereof, between the predicted output and the actual output. To calculate one or more of these values, for example, a predictive state model can be trained using a particular portion of the historical time series data, and the error can be calculated by applying the trained predictive state model to covariate signal data from another portion of the time series data. The output of the trained model responsive to this covariate signal data can be compared to the actual output of the historical data to calculate the error. The processing unit can evaluate multiple model types, each based on a different prediction method, and then select the model that provides the smallest error. In some cases, a model with a particular accuracy performance score can be selected. For example, the score can be a figure of merit ("FOM"), such as the lowest "(mean absolute error)*(processing resources)". The FOM can also be generated from other types of errors or other metrics. Processing resources may refer to the processing time, energy, or a combination thereof used by a processing unit to perform a predictive function in a model.
[0077] The weighted mean squared error can be calculated, for example, by assigning different weights to one or more different sections of the training data, which has the advantage that it can improve the accuracy of the model by centering the model's behavior around one or more sections of the time-series history data that more accurately reflect the sensor's behavior.
[0078] The determination of the model type can usually be performed as a first step. As explained, an error and / or FOM analysis can be performed based on historical sensor data. One or more predictive state models under evaluation can be trained by specifying a time window in the historical time series data. The trained predictive state models can then be compared by the processing unit for error and / or performance, for example, by using the remaining historical time series data of the sensor. For this, the processing unit can also determine the covariate signals to be used as input to the model. As explained, these can be either all covariates of the plant or a subset thereof. The subset can be at least partially specified by the user or, as explained before, the processing unit can select the dominant covariates at least partially based on the predictive power of each covariate on the model output. The processing unit can then select the best predictive state model to generate the predicted output of the sensor using the current time series data from the sensor.
[0079] As previously indicated, a similar approach to that described above for the sensor's predicted state model can also be used for the sensor object model.
[0080] Preferably, the training data for training the sensor predictive state model includes sensor data related to normal operating conditions. In this case, training based on undesirable operating conditions can be prevented or reduced. Undesirable deviations in the sensor data can be better captured. Similarly, the sensor object model is preferably trained using residual data related to normal operating conditions. This can therefore provide a synergistic effect when used together with the proposed sensor object, i.e., by condensing the number of monitored parameters while improving the visibility of unwanted changes in the sensor data caused by anomalies, monitoring of level and related signals. Also in the case of the predictive state model, this can improve the visibility of anomalous sensor outputs.
[0081] According to yet another aspect, the plurality of sensors are subdivided into categories such as sensors belonging to a plurality of plant areas. The sensors belonging to each plant area, or the plant area sensors, can be subdivided into sensors belonging to a plurality of process groups. The sensors belonging to each process group, or the process group sensors, can be subdivided into sensors belonging to a plurality of sensor objects. In the above context, it will be understood that each sensor object is realized by grouping the time series output data from the sensors belonging to that sensor object. An advantage of structuring the plurality into plant areas may be to enable easier navigation in the user interface ("UI") to the part of the plant that the user is interested in. To configure each process group, the sensors or tags belonging to it are configured as covariates or as monitored tags. The difference between covariate tags and monitored tags is that they are mutually exclusive within the same process group, but monitored tags in a process group may be covariates in another process group. As will be understood, the subdivision into sensor objects is based on the sensor data to be monitored simultaneously. Thus, the level signal and the related signal are the indicators monitored for each sensor object, as explained before. As an example, if the plant is a thermal power plant, multiple sensors in the thermal power plant can be subdivided or tagged according to plant areas such as reservoirs, generating units, switchyards, etc. The generating unit areas can be subdivided into process groups such as boilers, feedwater loops, turbines, condensers, generators, etc. Monitoring or sensor objects can be created from sensors that belong to the same process group or from sensors that belong to different process groups. Such subdivision of the plant can also help to reduce the number of covariates associated with forecasting using predictive state models and / or sensor object models. Thus, instead of using all covariate signals available in the plant for a particular observed sensor or sensor object, only covariates in the vicinity of the observed sensor or sensor object can be considered.While some covariate signals related to essentially all observed sensors or sensor objects, e.g., ambient temperature, may still be present, other covariates that have a more localized effect within a particular area of the plant may be ignored in other areas. The proposed subdivision can therefore further simplify the generation of predictive state models and / or sensor object models.
[0082] To generate the residual signal for each individual sensor, the proposed teaching provides two different states of the sensor, namely, an actual state, which represents the observed or measured output value of that individual sensor at any particular time t, and a predicted state, which represents the predicted value of that sensor at time t. The predicted state is preferably defined from a normal plant operating mode or operation. The actual state of the sensor is compared with the predicted state of the sensor to generate a residual signal for that sensor. Thus, the residual signal represents the deviation of the actual or observed state of the sensor from the predicted state of the sensor.
[0083] As explained before, a level event signal is generated when the level signal value changes beyond a predicted level signal value or a specific level signal limit. Similarly, an associated event signal is generated when the associated signal value changes beyond a predicted associated signal value or a specific associated signal limit. For both signals, the limit values can be absolute or relative to the respective signal values. The same can be true for the thresholds of the residual signals. As explained before, the thresholds and limits can be defined per application. According to another aspect, the thresholds of one or both of the residual signals are determined by the processing unit using a control chart. According to yet another aspect, the limits of one or both of the level signal and one or both of the associated signals are determined by the processing unit using a respective control chart. It is understood that the control chart is generated using the respective residual data or score data. According to one aspect, the limits are defined, for example, as quantiles of the predicted values of the respective scores. As a non-limiting example, the upper control can be at or around the 99.5% quantile. As a further non-limiting example, the lower control can be at or around the 0.5% quantile.
[0084] According to one embodiment, the predictive state model and / or the sensor object model are retrained between one or more predetermined time intervals. Doing so allows for a correct recalibration of the predictive state determination taking into account changes in the sensor's response that may be caused by natural factors such as aging and / or calibration drift. Retraining of the model may also be automatically triggered in response to the model's performance dropping below a minimum performance threshold for the model. Alternatively, or in addition, retraining of the sensor object model may be automatically triggered in response to changes in process parameters, such as user inputs to change the plant output, such as an increase or decrease in production rate. Thus, retraining can capture changes in the operation of the equipment and / or the plant.
[0085] According to one aspect, the level signal values are generated using a distance estimator, e.g., the T2-Hotelling statistic on the residual data of the sensor object. T2-Hotelling is a generalization of the t-statistic and indicates deviation from the multivariate mean of a group of variables. In general, the higher the value of the T2-Hotelling statistic, the further the observation from the mean. When used to calculate the level signal, it can indicate when and to what extent the residual data deviates from a normal, expected, or average state. Any suitable distance estimator can be used to calculate the level signal. According to a further aspect, the relevant signal values are generated using a measure of multivariate dependency, e.g., the generalized variance ("GV") statistic on the residual data of the sensor object. GV can be calculated as the determinant of the variance-covariance matrix of a sample of observations and is a multivariate generalization of variance. Thus, it can be used to measure the variance of the time series residual data in the sensor object at a particular time. Each signal value is calculated for the time series residual data from a respective time window, each of which is of a particular length. The time windows can be selected, for example, based on the quality of the training data.
[0086] According to one embodiment, the historical level signals and / or related signals are recorded as time series data on a database operatively connected to the processing unit.
[0087] According to an embodiment, any of the time series data (e.g., sensor time series data and / or residual data and / or level signal data and / or associated signal data) also includes annotation data. The annotation data may be provided via user input, but in some cases may further be provided at least partially automatically by the processing unit. The annotation data may be provided by type and / or level. The annotation type may classify features of the data, such as what a particular section of the time series data is relevant to. For example, maintenance activity, fault, data issue, etc. The level type may specify which level of the plant the annotation is relevant to. For example, the plant level may be used to refer to annotations associated with all sensor objects and their sensors, the area level may refer to annotations associated with all sensor objects and their sensors that belong to a particular plant area, similarly specifying annotations at the process group level, sensor object level, and even sensor level. According to an embodiment, the processing unit uses one or more annotations to automatically select an appropriate portion of the historical data for training the predictive state model and / or the sensor object model. For example, the processing unit may avoid historical time windows that include certain types of annotations in the time series. In doing so, training of data that is inconsistent with the model may be avoided. Thus, the accuracy of the model can be improved by selecting training data that provides appropriate information about normal plant or equipment operation. According to one embodiment, the processing unit automatically places annotations to define a historical time window according to desired limits and / or thresholds for the scores and / or residuals, respectively. This can thus be used to influence how closely the observed signal needs to track the predicted signal. The user-specified annotations can be received at a user interface ("UI") and stored, for example, in an analysis database. The annotations can also include timestamps to specify the start and end of the annotation. The annotations can also be used to retrieve desired sections of the time series data.
[0088] In some plants, creeping processes such as pipe fouling may cause sensor data trends to slow or drift. Such trends may cause sensor data values to slowly rise or fall over time. Trends may have small slopes or rates of change, so that it may take, for example, a week, weeks, or even months for any particular detectable change in the sensor output caused by such creeping processes to appear. Retraining the predictive state model and / or sensor object model by monitoring the levels and / or associated signals, as well as residuals, may fail to detect such slow trends. To prevent this, the processing unit may perform trend detection. Applicant has discovered that calculations of the strength, smoothness, and recency of the sensor's historical data may be particularly useful in detecting drift in the time series data.
[0089] Thus, the method may also include: - detecting, via either one or more processing units, a drift in an output signal of a sensor, the sensor being among at least some of the sensors, the drift being calculated from historical time series data of the sensor, the historical data of the sensor being for a period of at least one week, and the drift being detected by calculating the strength, smoothness, and recency of the historical data of the sensor.
[0090] Those skilled in the art should understand that strength in this context refers to the strength of the signal trend. Strength can therefore be expressed via a measurement of the slope of the trend. Strength represents the strength and recognizability of the drift, which can be measured, for example, using the Mann-Kendall test on the sensor's historical data. Such a test results in a value score that can indicate whether a trend can be detected and whether it is a positive or negative trend. Thus, a measurement of strength can indicate whether the trend or drift is weak or strong.
[0091] Smoothness indicates whether the drift is fairly smooth or whether it is due to more abrupt level shifts in the historical data. Smoothness thus represents a standardized measure of the degree to which the data is free of sudden features such as spikes, level shifts, etc. The presence of such features may increase the uncertainty of the trend. For example, the strength of the trend may be artificially inflated or deflated. Thus, the present teachings propose to use the smoothness value in the context of intensity and recency to detect or calculate a drift in the sensor's output signal that approximates the actual drift of the sensor output.
[0092] Some trends may have been active for long periods of time, such as weeks or months, and therefore may be recognizable if a time window of a few weeks or months is considered, but may have cooled or become inactive recently since then. Such characteristics can be detected using recency or reality tests. Recency, prevalence, or concurrency of a trend refers to a measure of whether an actual trend identified during observation is part of a long-term trend of change in the sensor output data. Thus, drift that is no longer active can be ignored.
[0093] The above criteria thus allow synergistic detection of actual drifts that may be related to anomalies or potential anomalies. This therefore allows maintaining such actual drifts while ignoring such variations in the sensor output that do not or are unlikely to represent anomalies. Thus, more reliable visibility can be maintained for slower moving effects that have not yet manifested as anomalies despite retraining of the predictive state model.
[0094] According to one embodiment, the time period for calculating drift is a period of one month or about one month in length. Additionally or alternatively, according to another embodiment, the time period is three months or about three months in length. Additionally or alternatively, according to yet another embodiment, the time period is six months or about six months in length. Preferably, the historical data is essentially a long portion of the time period of the time series data up to the time that trend detection is being performed.
[0095] The present teachings relate, for example, at least in part, to a black-box monitoring method, and may also be used to provide a monitoring system for a plant, as outlined above. Thus, a monitoring and / or control system for a plant including a plurality of sensors may also be provided, the system including one or more processing units configured to perform any of the method steps disclosed herein.
[0096] For example, a monitoring and / or control system for a plant including a plurality of sensors may be provided, the system including one or more operatively connected processing units, the system including: - generating, via one or more processing units, time series residual data for a sensor object, the sensor object being a group of at least a portion of a plurality of sensors, the residual data comprising, for each sensor of the sensor object, a residual signal that is a difference between a measured output of the sensor and a predicted output of the sensor; - monitoring, via any of one or more processing units, a level signal, the level signal being indicative of collective time-based variations of the time series residual data; - monitoring, via any of one or more processing units, a relevant signal, the relevant signal indicating a variance and / or a relevant structure of the time series residual data; - generating, via any of the one or more processing units, an anomalous event signal when the value of the level signal and / or the value of the associated signals at a particular time changes from an expected value of the respective signal at or around that time.
[0097] Similarly, in another aspect, a monitoring and / or control system for a plant including a plurality of sensors may be provided, the system including one or more functionally connected processing units, the system including: - providing, at any of the one or more processing units, time series residual data of a sensor object, the sensor object being a group of at least a portion of a plurality of sensors, the residual data comprising, for each sensor of the sensor object, a residual signal that is a difference between a measured output of the sensor and a predicted output of the sensor; - monitoring, via any of one or more processing units, a level signal, the level signal being indicative of collective time-based variations of the time series residual data; - generating, via any of the one or more processing units, a level event signal, the level event signal being generated when a value of the level signal at a particular time changes from an expected value of the level signal at or around that time, the level event signal indicating an anomaly in at least one piece of equipment in the plant.
[0098] Similarly, in yet another aspect, there may be provided a monitoring and / or control system for a plant including a plurality of sensors, the system including one or more operatively connected processing units, the system including: - providing, at any of the one or more processing units, time series residual data of a sensor object, the sensor object being a group of at least a portion of a plurality of sensors, the residual data comprising, for each sensor of the sensor object, a residual signal that is a difference between a measured output of the sensor and a predicted output of the sensor; - monitoring, via any of one or more processing units, a relevant signal, the relevant signal indicating a variance and / or a relevant structure of the time series residual data; - generating, via any of the one or more processing units, an associated event signal, the associated event signal being generated when, at a particular time, a value of the level signal changes from an expected value of the associated signal at or around that time, the associated event signal being indicative of an anomaly in at least one piece of equipment in the plant.
[0099] The method or system can be used to detect anomalous behavior and anomalous patterns in sensor data via the proposed automated data-driven techniques. The teachings can benefit plant operators, for example, by enabling more efficient tracking of large sensor data sets by directing the operator's attention at an early stage to areas of the plant where future action may be required. As previously outlined, the teachings can also be used to provide forecasts of upcoming maintenance requirements associated with the plant. According to one aspect, the teachings can also be used to realize automated systems for plant maintenance forecasting and control.
[0100] The user interface can be any suitable human machine interface ("HMI") that allows a user to interact with the monitoring system. The human machine interface ("HMI") can include any one or more of the following: a monitoring panel, a video display unit (e.g., LCD (Liquid Crystal Display), CRT (Cathode Ray Tube) display, touch screen), an alphanumeric input device (e.g., keyboard), a cursor control device (e.g., mouse), and / or a signal generating device (e.g., speaker). Thus, the HMI can be, for example, a visual interface such as a panel, a screen, and / or it can be an audio interface such as a speaker. Thus, the output can be displayed to the user and / or announced via the speaker.
[0101] The processing unit may be a computer, or even a general-purpose processing device such as a microprocessor, a microcontroller, a central processing unit ("CPU"), or the like. More specifically, the processing unit may be a CISC (complex instruction set computing) microprocessor, a RISC (reduced instruction set computing) microprocessor, a VLIW (very long instruction word) microprocessor, or a processor implementing other instruction sets, or a processor implementing a combination of instruction sets. The processing unit or processing means may also be one or more dedicated processing devices, such as an ASIC (application specific integrated circuit), an FPGA (field programmable gate array), a CPLD (complex programmable logic device), a DSP (digital signal processor), a network processor, or the like. The methods, systems, and devices described herein may be implemented as software in a DSP, a microcontroller, or any other side processor, or as hardware circuitry within an ASIC, a CPLD, or an FPGA. Also, as outlined above, it should be understood that the term "processing unit" or processor may also refer to one or more processing devices, such as a distributed system of processing devices located across multiple computer systems (e.g., cloud computing), and is not limited to a single device unless otherwise specified. Furthermore, any one or more of the processing units may be located in a different physical location than other processing units.
[0102] Viewed from another perspective, a computer program may be provided that includes instructions that, when executed by any one or more suitable processing units of a plant monitoring and / or control system operatively connected to a plurality of sensors, cause the system to perform the method steps disclosed herein.
[0103] For example, a computer program may be provided that includes instructions that, when executed by one or more operatively connected processing units of a plant monitoring and / or control system operatively connected to a plurality of sensors, cause the system to:
[0104] - generating, via one or more processing units, time series residual data for a sensor object, the sensor object being a group of at least a portion of a plurality of sensors, the residual data comprising, for each sensor of the sensor object, a residual signal that is a difference between a measured output of the sensor and a predicted output of the sensor; - monitoring, via any of one or more processing units, a level signal, the level signal being indicative of collective time-based variations of the time series residual data; - monitoring, via any of one or more processing units, a relevant signal, the relevant signal indicating a variance and / or a relevant structure of the time series residual data; - generating, via one or more processing units, an anomalous event signal when the value of the level signal and / or the value of the associated signals at a particular time change from an expected value of the respective signal at or around that time.
[0105] Similarly, from yet another perspective, a computer program product may be provided that includes instructions that, when executed by one or more operatively connected processing units of a plant monitoring and / or control system operatively connected to a plurality of sensors, cause the system to:
[0106] - providing, at any of the one or more processing units, time series residual data of a sensor object, the sensor object being a group of at least a portion of a plurality of sensors, the residual data comprising, for each sensor of the sensor object, a residual signal that is a difference between a measured output of the sensor and a predicted output of the sensor; - monitoring, via any of one or more processing units, a level signal, the level signal being indicative of collective time-based variations of the time series residual data; - generating, via any of the one or more processing units, a level event signal, the level event signal being generated when a value of the level signal at a particular time changes from an expected value of the level signal at or around that time, the level event signal indicating an anomaly in at least one piece of equipment within the plant.
[0107] Similarly, from yet another perspective, there can be provided a computer program comprising instructions that, when executed by one or more operatively connected processing units of a plant monitoring and / or control system operatively connected to a plurality of sensors, cause the system to:
[0108] - providing, at any of the one or more processing units, time series residual data of a sensor object, the sensor object being a group of at least a portion of a plurality of sensors, the residual data comprising, for each sensor of the sensor object, a residual signal that is a difference between a measured output of the sensor and a predicted output of the sensor; - monitoring, via any of one or more processing units, a relevant signal, the relevant signal indicating a variance and / or a relevant structure of the time series residual data; - generating, via any of the one or more processing units, a related event signal, the related event signal being generated when, at a particular time, a value of the level signal changes from an expected value of the related signal at or around that time, the related event signal being indicative of an anomaly in at least one piece of equipment within the plant.
[0109] Viewed from yet another point of view, a computer readable data carrier may be provided having stored thereon a computer program as disclosed herein.Thus, a non-transitory computer readable medium may be provided which stores a program for causing a suitable processing unit of a plant monitoring and / or control system to carry out any of the method steps disclosed herein.
[0110] The computer-readable data carrier includes any suitable data storage device on which one or more sets of instructions (e.g., software) embodying any one or more of the methodologies or functions described herein are stored. The instructions may also reside, completely or at least partially, in a main memory and / or in a processor during its execution by a processing unit or device that may constitute a computer system, a main memory, and a computer-readable storage readable medium. The instructions may further be transmitted or received over a network via a network interface device.
[0111] A computer program for implementing one or more of the embodiments described herein may be stored and / or distributed on a suitable medium, such as an optical storage medium or a solid-state medium supplied together with or as part of other hardware, but may also be distributed in other forms, such as the Internet or other wired or wireless communication systems, but the computer program may also be presented over a network, such as the World Wide Web, and may be downloaded from such a network into the working memory of a data processor.
[0112] Viewed from another point of view, a data carrier or data storage medium for making a computer program element downloadable can also be provided, which computer program element is arranged to perform a method according to one of the preceding embodiments.
[0113] The word "comprising" does not exclude other elements or steps, and the indefinite articles "a" or "an" do not exclude a plurality. A single processor or controller or other unit may fulfill the functions of several items recited in the claims. The mere fact that certain measures are recited in mutually different dependent claims does not indicate that a combination of these measures cannot be used to advantage. The term "functionally connected" may also be read as "operably connected" or connected in a direct or indirect manner. Reference signs in the claims are not to be interpreted as limiting the scope.
[0114] Hereinafter, embodiments will be described with reference to the accompanying drawings. [Brief description of the drawings]
[0115] [Figure 1] 1 illustrates an exemplary industrial plant in which certain aspects of the present teachings may be applied. [Diagram 2] 1 shows a block diagram including some of the signals generated in accordance with the present teachings. [Diagram 3] 1 shows a plot of a general variance statistical signal in accordance with the present teachings. [Figure 4] The effect of annotations on the control limits is shown. [Diagram 5] 1 illustrates a specific example of trend detection. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0116] Detailed Description 1 illustrates an example 100 of a plant 101 for purposes of illustrating how at least certain aspects of the present teachings may be applied. The layout 100 shows an industrial plant 101. The plant includes a number of instruments and sensors. The plant also includes a processing unit 110, which is shown distributed across three portions 110a, b, and c.
[0117] Shown is a first group 102 of equipment and sensors that are part of a plant, as well as a second group 103 of equipment and sensors. The plant 101 is used to manufacture one or more industrial products 150. The product 150 can be any physical product or service product, as previously outlined. For example, the product 150 can be a chemical product or a pharmaceutical product. The architecture or process shown in the example 100 is not important to the generality or scope of the present teachings. In this example, the first group 102 is located at a different location than the second group 103. An intermediate product is provided by the output of the first group 102 as an input product to the second group 103. The intermediate product is shown to be provided via a transportation medium, such as a pipeline 188, which can be, for example, a long pipeline. This is shown to demonstrate that most of the equipment in the first group 102 can be relatively isolated from the equipment in the second group 103. For example, there can be interdependencies between the first group 102 and the second group 103 due to parameters of the intermediate product being transferred via the pipeline 188. However, there may be certain factors, e.g., atmospheric pressure and temperature, that are common to the two groups 102 and 103. Such ambient parameters may affect the process parameters or sensor outputs on both sides. Thus, as they relate to the process, any such ambient parameters may be considered to be what was previously called a covariant signal.
[0118] Both the first group 102 and the second group 103 include multiple sensors, such as temperature sensors 132, 133, 142 and 148, pressure sensors 311, 135, 136 and 145, flow sensors 138, 143 and 147. The equipment in both groups includes a heat exchanger 130, a separation chamber 139, a reaction tank 120, a cooling unit 140, a filter 151, a fan 141, and pumps 134, 144, and 149.
[0119] The sensors from the first group 102 are monitored by the processing means 110, or more specifically by the first processing unit 110a. Signals from the sensors of the first group 102 are shown received via the first communication means 105a. The communication means 105a can be any means suitable for transmitting signals or data from the sensors, wired, wireless, or a combination thereof. For example, the first communication means 105a can be a bus as shown. The data received by the first processing means 110a can be processed by the first processing means 110a and / or by any other processing means 110b and c. At least some of the data can also be stored in a memory or database 111. The database 111 can be located in a single location or distributed as shown at 111a, b, and c. In addition to monitoring, the first processing unit 110a can also perform control functions, for example via the control bus 106a. The control bus 106a of the first processing unit 110a may be any communication means as previously described in the context of the bus 105a. In some cases, the bus 105a and the control bus 106a may furthermore be the same bus or communication means. The control functions may include, for example, the control of the pump 134. The processing unit 110 may even be provided by an HMI 112. The HMI 112 may be provided to each of the distributed processing units 110a, b, and c as shown, or to any one or more of them. The HMI may comprise a monitoring panel or a video screen and one or more input devices, such as a keyboard or mouse, for a user to interact with the processing means 110. The HMI may also include an audio device, such as a speaker. Events, such as alarms, may be communicated audibly and / or visually via the HMI.
[0120] Similarly, the sensors from the second group 103 are monitored by the processing means 110, or more specifically by the second processing means 110b. Signals from the sensors of the second group 103 are shown received via the second communication means 105b. The second communication means 105b can be any means, wired, wireless, or a combination thereof, suitable for transmitting signals or data from the sensors. For example, the second communication means 105b can be a bus as shown. The data received by the second processing means 110b can be processed by the second processing means 110b and / or by any other processing means 110a and c. Again, at least some of the data may also be stored in memory or database 111. In addition to monitoring, the second processing unit 110b can also perform control functions, for example via the second control bus 106b. The control functions may include, for example, control of pumps 144 and 149, fan 141, and valves 146. The second control bus 106b of the second processing unit 110b can be any communication means. In some cases, the bus 105b and the control bus 106b can even be the same bus or communication means.
[0121] The first processing unit 110a and the second processing unit 110b are operatively connected via a data link 190, which may be any suitable communication medium, wired, wireless, or a combination thereof. Thus, the processing units can exchange data, which may include any data or signals, such as sensor data, status data, and event signals. The data link may also be used to transfer data from one database or memory to another.
[0122] In some cases, a separate processing unit may be provided, for example, the third processing unit 110c. The third processing unit 110c may be at a higher hierarchical level and may be a plant-level monitoring and / or control system. The third processing unit 110c may be co-located with the plant or even at least partially located elsewhere, for example, a cloud-based platform. In some cases, the third processing unit 110c may be in the plant, but its database 111c may be implemented as a cloud storage or vice versa. The monitoring processing unit 110c may even be located in another plant, which is located differently from the plant 101. In some cases, the third processing unit 110c may be located between the first processing unit 110a and the second processing unit 110b, i.e., the data link 190 is divided into two sections, the first between the first unit 110a and the third unit 110c and the second between the third unit 110c and the second unit 110b. The particular architecture of the processing unit or plant is not essential to the scope or generality of the present teachings.
[0123] Another plant may be in another country. Therefore, even the first group 102 and the second group 103 may be located in different plants or countries. For example, a supplier plant and a consumer plant connected via a gas pipeline may be located in different countries.
[0124] Any of the processing units 110a, b, and c, and the databases 111a, b, and c, may be implemented, for example, as a cloud-based service provided by a third party. In some cases, the processing units 110a, b, and c, and / or the databases 111a, b, and c, may be co-located or they may even be the same unit.
[0125] To monitor equipment, conventional systems may individually monitor the state of one or more sensors, such as the output signal from the temperature sensor 148. An increase in temperature may be used to indicate overheating of the pump 149, for example due to a reduction in flow rate caused by a blockage of the filter 151. However, in reality, there may be cases where the temperature increase is due to an increase in the ambient temperature. Thus, such systems may lead to false positive events indicating an anomaly.
[0126] To solve this problem, the second processing unit 111b can compare the measured or observed output value of the temperature sensor 148 with its predicted value at that time. The predicted value can be generated by a predicted state model of the temperature sensor 148. To improve the predicted state forecast, a model, for example a neural network, can be trained using historical time series data of the sensor 148, preferably under the desired operating conditions. To further improve the forecast, the predicted state model can be input with covariate signals that influence the output of the sensor 148. For example, the ambient temperature can be one of the covariate signals. There can be other signals or parameters from the second group 103 that influence the output of the sensor 148. Such covariates are recognized during the model building phase of the predicted state model. The processing means 110 can use the entire covariate pool of the entire plant 101 to check which of the covariates have an influence or predictive power on the output of the sensor 148, using the historical data. Thus, the covariates that have a measurable influence on the output of the sensor 148 are selected as model inputs. Once the model is deployed on the processing unit 110, a residual signal is generated for the output of the sensor 148, which is the difference between the observed sensor output value and the output of the predicted state model at that particular time. If the sensor is operating normally, the residual signal will be mostly random noise.
[0127] To make the anomaly detection more resistant to noisy spikes and such imperfections in the residual signals, the present teachings propose to create a sensor object. A sensor object refers to a group of sensor residual signals that are integrated and monitored together. The group of sensor residual signals is a time series residual signal received from a preselected number of sensors. The preselected number of sensors may be selected manually or at least partially automatically via the processing unit 110. The processing unit may determine this, for example, based on similarity of sensor response, sensor type, covariate dependency, or a combination thereof. The sensor residual signal or group of residual data is then analyzed by the processing unit 110 to calculate a level signal. The level signal is indicative of the collective time-based variation of the time series residual data. The time-dependent level signal value is then compared to a predicted level signal value at that time. The processing unit 110 may generate a level event signal at any particular time when the value of the level signal changes beyond the predicted level signal value at or around that time. The level event signal is deemed to be indicative of an anomaly of at least one piece of equipment in the plant. In this example, if the level signal violates the expected level signal, the processing unit 110 can generate an alert. Additionally, the processing unit 110 can check which sensors in the sensor object had their sensor output violated the expected sensor output value at or around the time the level event signal was generated. This is used by the processing unit 110 to find the cause of the anomaly.
[0128] The predicted level signal values are preferably generated by the processing unit 110 using a sensor object model, which is a predictive model or neural network trained using historical residual data.
[0129] Preferably, the predicted level signal is provided as a range of values within which the level signal may lie. Thus, one or more limit values for the level signal may be provided. A level event signal is generated when the observed level signal value exceeds a predicted level signal limit. The predicted level signal limit may be an upper predicted level signal limit and / or a lower predicted level signal limit.
[0130] Preferably, the processing unit calculates another score to make the anomaly detection more resistant to noisy spikes in the residual signal. That is, a time-dependent related signal value is generated. Thus, the sensor residual signal or group of residual data is analyzed by the processing unit 110 to calculate a related signal. The related signal indicates the variation and / or related structure of the time series residual data. The time-dependent related signal value is then compared to the predicted related signal value at that time. The processing unit 110 can generate a related event signal at any particular time when the value of the related signal changes beyond the predicted related signal value at or around that time. The related event signal is considered to be indicative of an anomaly of at least one piece of equipment in the plant. Referring again to the example, the processing unit 110 can issue an alarm if the related signal violates the predicted related signal. Furthermore, the processing unit 110 can check which sensors in the sensor object have violated the predicted sensor output value at or around the time the level event signal was generated. This can also be used by the processing unit 110 to find the cause of the anomaly.
[0131] The predicted associated signal values are preferably generated by the processing unit 110 using a sensor object model.
[0132] Preferably, the predicted relevant signal is provided as a range of values that the relevant signal may be. Thus, one or more limit values of the level signal may be provided. The relevant event signal is generated when the observed relevant signal value exceeds a predicted relevant signal limit. The predicted relevant signal limit may be an upper predicted relevant signal limit and / or a lower predicted relevant signal limit. The limits may also be referred to as control limits.
[0133] Violation of either or both of the level and related signals may be considered to indicate an anomaly, either in the expected value or in the limit, respectively.
[0134] To catch anomalies that may appear slowly, the processing unit 110 can even perform trend detection. By retraining the model, slow moving drifts can be eliminated from the level and related monitoring observations. Values: Intensity, smoothness, recency, prevalence or concurrency of the sensor's historical data calculated by the processing unit to detect drift in the sensor's time series data.
[0135] In the above description, it may be referred to as a particular function being performed by "processing unit 110", but it will be understood that in some cases, it may be further implemented to be performed via any of one or more of processing units 110a, b, and c. It will also be understood that in some cases, additional processing units may be present. For example, some sensors may be further provided with a dedicated processor configured to calculate a residual signal of that sensor. In that case, the residual signal of such sensor may be provided directly to processing unit 110 as an input.
[0136] Similarly, for the first group 102, the processing unit 110, e.g., in some cases, the first processing unit 110a, can monitor another sensor object via the corresponding one or both levels and associated signals of that object. Each group may have multiple sensor objects.
[0137] In response to the event signal, the processing unit can backtrack the sensor data to find the cause of the anomaly, for example as outlined above. Additionally, the processing unit can predict the maintenance requirements of the anomaly, for example by providing an estimated date and time when maintenance should be performed to prevent a particular disruption. The disruption can be calculated as lost productivity or waste compared to a planned shutdown for maintenance.
[0138] Figure 2 shows a block diagram 200 representing the signals to be generated and monitored. The chart on the left shows the observations and predictions of five different sensors grouped into sensor objects. The grouping is preferably done automatically, for example using a self-organizing map, but may also be done at least in part based on manual feedback.
[0139] The first set of curves 201a relates to the measured output signals from the first sensor and their predicted outputs. Similarly, curves 201b-e relate to the measured output signals from the second through fifth sensors and their predicted outputs, respectively. By comparing the measured or observed output of each sensor with the respective predicted outputs, respective residual signals 202a-e are obtained. For example, the first residual signal 202a relates to the first sensor. As can be seen, each sensor output shown in curves 201a-e was quite different from the outputs from the other sensors, whereas the residual signals 202a-e are more uniform. As explained before, redundant information from the sensor outputs can be removed by generating residual signals.
[0140] It is clear that the signal is time-dependent or composed of time series values. By combining the residual signals 202a-e, a sensor object 204 is realized, which includes multi-dimensional residual data 203. From the residual data 203, a time-dependent level signal or score 205 is generated and shown. The level signal 205 is provided with a predicted level signal limit 207, which represents a probability space of predicted values in which the level signal can validly be. The predicted level signal limit 207 may also be a time-dependent value. As shown, shortly after time 209, the predicted level signal limit 207 is decreased by the sensor object model. It is also seen that 205p represents a peak in the level signal 205 when said signal changes beyond the predicted value of the level signal, or the predicted level signal limit 207 at that time. Thus, in such a case, a level event signal will be generated by the processing unit 110. The processing unit can then trace the root cause of the anomaly, for example, by analyzing one or more sensor signals 201a-e. The processing unit can use effect size calculations to find the sensor or sensors that contribute most to the change in the signal. An alarm can be displayed on the visual monitoring panel 210. For example, the relevant equipment on the panel can be highlighted.
[0141] Also from the residual data 203, a time-dependent associated signal or score 206 is generated and shown. The associated signal 206 is provided with a predicted associated signal limit 208, or more specifically, an upper associated signal limit 208a and a lower associated signal limit 208b. The distance between these limits represents the probability space of predicted values in which the associated signal can validly be. The predicted associated signal limits 208a and b may also be time-dependent values. It can be seen that 206p represents the peak portion of the associated signal 206 when said signal changes beyond the upper predicted value of the associated signal, or the upper predicted associated signal limit 208a for that time. Thus, in such a case, an associated event signal will be generated by the processing unit 110. The processing unit can then trace the cause of the anomaly, for example, by analyzing one or more sensor signals 201a-e. The processing unit can use effect size calculations to find the sensor or sensors that contribute most to the signal change. The associated score can also detect the rate of change of the residual signal movement. Similarly, an alarm can be displayed on the visual monitoring panel 210.
[0142] In some cases, event signals for either or both scores may be caused by plant activity that results in the residual data deviating from a predicted state. Such activity may be a repair or other event that changes the behavior of the observed state. In such cases, the user may recognize the cause of the event. The user may then annotate the event according to a particular classification or type. Thus, one or more annotations may be fed back into the model so that the model is trained to classify such events in the future.
[0143] FIG. 3 illustrates a plot 300 of a generalized variance ("GV") that can be used to calculate the variance and associated structure of the sensor object data. Thus, the GV can be used to generate the associated signal 206. The plot 300 illustrates the associated signal 306 with generalized variance values on the Y-axis 310 and time on the X-axis 302. Also illustrated are limit values 308, which may be referred to as an upper control limit ("UCL") 308a and a lower control limit ("LCL") 308b. The distance 310 between these values 308a and b represents a control range or probability space within which the value of the associated signal 306 may validly lie. Thus, a value that falls within the control range 310 at any particular time may be referred to as a predicted value at that time. The control range 310 may also be time dependent, although in this case it is shown as a constant.
[0144] It is found that within the first time period 304, the value of the associated signal 306 has changed from an expected value, or in this case has changed by more than the lower control limit 308b at or around that time. Thus, in this case an associated event signal is created or generated.
[0145] Similarly, within the second time period 305, the value of the associated signal 306 again changes from the expected value, or in this case, first beyond the lower control limit 308b and then beyond the upper control limit 308a. Thus, in this case, one or more associated event signals are also generated.
[0146] Such a chart 300 may also be referred to as a control chart used to monitor the numerical statistics of a signal over time. Similar control charts may also be generated for level signal value statistics.
[0147] FIG. 4 shows two charts 400 illustrating how annotations can be used to improve the focus or relevance of training data. The charts 400 are shown as control charts. The left control chart 430a shows control limits 410 and 420a of a score signal 490. The score signal in this case is the relevant signal. The Y-axis is therefore the relevant signal value, for example, the GV statistic. The X-axis 302 represents time. The chart has a peak region 455 that shows a suddenly high value of the signal 490. Such a high value may have occurred due to an abnormal event, such as a maintenance activity on a particular piece of equipment that contains the sensor in the sensor object. The chart also shows that the start time 401 and the end time 402 constitute a time window that can be used to train a sensor model. If the left chart 430a is used to train a sensor model, the control limits are determined as follows, namely the upper control limit 420a and the lower control limit 410 of the unannotated chart 430a. It will be appreciated that such a high upper control limit 420a may not be adequate to detect changes in the score signal from the expected value at that time, and therefore some anomalous events may not be flagged.
[0148] This can be addressed by placing annotations 460 within the annotation time window 440. Thus, the processing unit 110 ignores data from the peak region 455 for training the model. The effect of the annotations can be seen in terms of the upper control limit 420b of the annotated chart 430b. The latter upper control limit 420b is now more realistic. The lower control limit 410 is not affected, since the annotations are only relevant for high values of the score signal 490. Thus, the annotations can be used to modify the weighting of the training data in the normal operating region.
[0149] Similar to the annotations, the selection of the time window, i.e. the period enclosed between the start time 401 and the end time 402, is also selected by the processing means such that the selected window reflects the normal operation of the sensor object. Annotations can also be used to mark such windows that are desirable. Thus, the control limits can be adjusted to better detect anomalies.
[0150] As explained earlier, the control limits can be statistical quantile limits. FIG. 5 shows exemplary charts 501-500 to demonstrate how the processing means can perform trend detection. The proposed trend detection involves calculating three metric values on long-term data within a time window of at least one week length. The three metric values are intensity, smoothness, and recency or realism. Trend detection is also used in sensor objects. In the following example, a value between 0 and 1 is assigned to each metric, where 0 indicates false and 1 indicates true. Typically, the values are between 0 and 1 and indicate the probability or trend property of the metric. It is to be understood that the values shown below are examples and the limits specified above are not absolute. Scaling factors can be applied, for example, any of the values are between 0 and 100. Thus, the values mentioned herein should not be construed as specified in an absolute sense. A symbol can also be assigned to the intensity value to specify the direction of the trend.
[0151] A first chart 501 shows an upward trending signal and the metric values calculated for the first chart 501 are: strength:1, smoothness:0.9, realism:1.
[0152] The second chart 502 shows a sharply rising signal that is contained between the low and high noise floors. Thus, this trend has almost disappeared. The metric values calculated for the second chart 502 are: Strength: 0.69, Smoothness: 0.98, Realism: 0.25.
[0153] The third chart 503 shows a signal in the form of a triangular waveform. The metric values calculated for the third chart 503 are: Strength: 0.1, Smoothness: 0, Realism: 0.24.
[0154] The fourth chart 504 shows a signal with an initial noisy portion followed by an upward trend. The metric values calculated for the fourth chart 504 are: Strength: 0.73, Smoothness: 0.97, Realism: 1.
[0155] The fifth chart 505 shows a signal with an initial upward trend followed by an upper plateau. The metric values calculated for the fourth chart 504 are: Strength: 0.77, Smoothness: 0.97, Realism: 0.28.
[0156] Various examples have been disclosed above of methods for monitoring a plant, monitoring and / or control systems for a plant, and computer software products implementing any of the associated method steps disclosed herein. However, those skilled in the art will understand that changes and modifications can be made to the examples without departing from the spirit and scope of the appended claims and their equivalents. It will be further understood that aspects from the embodiments of the methods and products discussed herein can be freely combined.
[0157] Particular embodiments of the present teachings are summarized in the following clauses. Article 1. 1. A method for monitoring a plant including a plurality of sensors and one or more operatively connected processing units, the method comprising: - providing, at any of the one or more processing units, time series residual data of a sensor object, the sensor object being a group of at least a portion of a plurality of sensors, the residual data comprising, for each sensor of the sensor object, a residual signal that is a difference between a measured output of the sensor and a predicted output of the sensor; - monitoring a level signal via any of the one or more processing units, the level signal being indicative of collective time-based variations in the time series residual data; and - monitoring, via any of one or more processing units, a relevant signal, the relevant signal indicating a variance and / or a relevant structure of the time series residual data; - generating an anomalous event signal when the value of the level signal and / or the value of the associated signals at a particular time via any of the one or more processing units changes from an expected value of the respective signal at or around that time.
[0158] Article 2. 2. The method of claim 1, wherein each of the predicted values is provided as a corresponding predicted value limit specifying a number of predicted values at a particular time as ranges and / or discrete values that the corresponding signal can validly have without generating an anomalous event.
[0159] Article 3. 2. The method of claim 1, wherein either the predicted value or the prediction limit value is a time-dependent value.
[0160] Article 4. The method also - determining at least one root cause of the anomaly by performing one or more of the following in response to the anomalous event signal: checking which sensors in the sensor object have their measured outputs changed from their predicted outputs at or around the time of the anomalous event; analysing the time series residual signals of each sensor in the sensor object to determine one or more main drivers or most dominant contributors to the level signal values; analysing the time series residual signals of each sensor in the sensor object to determine one or more main drivers or most dominant contributors to the associated signal values; and analysing the covariance of the time series residual signals of each pairwise combination of sensor residual signals in the sensor object to determine one or more main drivers or most dominant contributors to the associated signal values.
[0161] Article 5. The method also - The method of any one of the preceding clauses, comprising determining a health state of at least one device associated with the sensor object in response to the abnormal event signal.
[0162] Article 6. 2. The method according to any one of the preceding clauses, wherein either the predicted value or the prediction limit is provided by a sensor object model, which is a predictive model trained using historical residual data of the sensor object.
[0163] Article 7. 7. The method of claim 6, wherein one or more covariate signals are provided as inputs to the sensor object model, each covariate signal being a signal representing a parameter on which at least one of the residual signals depends.
[0164] Article 8. 2. The method of any one of the preceding clauses, wherein the predicted output of at least one sensor is provided by a predicted state model that is at least partially a predictive model trained using historical time series output data of the respective sensor.
[0165] Article 9. 9. The method of claim 8, wherein one or more covariate signals are provided as inputs to the predictive state model, each covariate signal being a signal representing a parameter on which the output of a sensor depends.
[0166] Article 10. The method according to any one of the above clauses, wherein the sensor object is provided by at least partially automatically grouping at least some of the sensors, using at least one data-centric algorithm, such as a clustering algorithm, e.g. a self-organizing map algorithm, further, e.g. the sensor object is at least partially automatically generated by one of the one or more processing units using at least one self-organizing map.
[0167] Article 11. 11. The method of any one of clauses 8-10, wherein the predictive state model is automatically selected by the processing unit by analyzing a plurality of different predictive model types and selecting the model type as the predictive state model that provides the smallest error between the output of that model when trained on a particular training window of the historical time series data and the actual historical sensor output within the particular time window of the historical time series data.
[0168] Article 12. A method according to any one of the preceding clauses, wherein the level signal values are generated using a distance estimator which indicates the time and amount that the time series residual data deviates from its normal or expected or mean state.
[0169] Article 13. 4. The method according to any one of the preceding clauses, wherein the associated signal values are generated using a statistical measure of multivariate dependence of the residual data or a measure of the variance of the time series residual data at a particular time.
[0170] Article 14. The method also - detecting, via one or more processing units, a drift in the output signal of the sensor, at least among some of the sensors, the drift being calculated from historical time series data of the sensor, the historical data of the sensor being for a period of at least one week, and the drift being detected by calculating the strength, smoothness and recency of the historical data of the sensor.
[0171] Article 15. 1. A method for monitoring a plant including a plurality of sensors and one or more operatively connected processing units, the method comprising: - providing, at any of the one or more processing units, time series residual data of a sensor object, the sensor object being a group of at least a portion of a plurality of sensors, the residual data comprising, for each sensor of the sensor object, a residual signal that is a difference between a measured output of the sensor and a predicted output of the sensor; - monitoring a level signal via any of the one or more processing units, the level signal being indicative of collective time-based variations in the time series residual data; and - generating, via any of the one or more processing units, a level event signal, the level event signal being generated when a value of the level signal at a particular time changes from an expected value of the level signal at or around that time, the level event signal being indicative of an anomaly in at least one piece of equipment in the plant.
[0172] Article 16. 1. A method for monitoring a plant including a plurality of sensors and one or more operatively connected processing units, the method comprising: - providing, at any of the one or more processing units, time series residual data of a sensor object, the sensor object being a group of at least a portion of a plurality of sensors, the residual data comprising, for each sensor of the sensor object, a residual signal that is a difference between a measured output of the sensor and a predicted output of the sensor; - monitoring, via any of one or more processing units, a relevant signal, the relevant signal indicating a variance and / or a relevant structure of the time series residual data; - generating, via any of the one or more processing units, an associated event signal, the associated event signal being generated when, at a particular time, a value of the level signal changes from an expected value of the associated signal at or around that time, the associated event signal being indicative of an anomaly in at least one piece of equipment in the plant.
[0173] Article 17. 1. A method for monitoring a plant including a plurality of sensors and one or more operatively connected processing units, the method comprising: - providing, at any of the one or more processing units, time series residual data of a sensor object, the sensor object being a group of at least a portion of the sensors from the plurality of sensors, the residual data comprising, for each sensor of the sensor object, a residual signal that is a difference between a measured output of the sensor and a predicted output of the sensor, the sensor object being provided by at least partially automatically grouping at least a portion of the sensors; - monitoring a level signal via any of the one or more processing units, the level signal being indicative of collective time-based variations in the time series residual data; and - monitoring, via any of one or more processing units, a relevant signal, the relevant signal indicating a variance and / or a relevant structure of the time series residual data; - generating an anomalous event signal when the value of the level signal and / or the value of the associated signals at a particular time via any of the one or more processing units changes from an expected value of the respective signal at or around that time.
[0174] Article 18. A monitoring and / or control system for a plant including a plurality of sensors, the system including one or more processing units configured to perform the method steps according to any one of clauses 1 to 17.
[0175] Article 19. A computer program comprising instructions that, when executed by a processing unit of a plant monitoring and / or control system operatively connected to a plurality of sensors, cause the system to perform the method steps according to any one of clauses 1 to 17.
Claims
1. 1. A method for monitoring a plant including a plurality of sensors and one or more operatively connected processing units, comprising: providing, at any of said one or more processing units, time series residual data for a sensor object, said sensor object being a group of at least some of said sensors from said plurality of sensors, said residual data comprising, for each sensor of said sensor object, a residual signal that is the difference between a measured output of said sensor and a predicted output of said sensor; - monitoring, via any of said one or more processing units, a level signal, said level signal indicative of collective time-based variations of said time-series residual data; and - monitoring, via any of said one or more processing units, a relevant signal, said relevant signal indicating variability and / or relevant structure of said time series residual data; generating, via any of said one or more processing units, an anomalous event signal when the value of said level signal and / or the value of said related signal at a particular time changes from an expected value of said respective signal at or around that time; The method, wherein the sensor objects are provided by at least partially automatically grouping the at least some of the sensors using at least one data-centric algorithm.
2. 2. The method of claim 1 , wherein said any of said respective forecast values is provided as a corresponding forecast value limit specifying a number of forecast values at a particular time, as a range or discrete values that said corresponding signal can validly have without generating an anomalous event.
3. The method comprises: The method according to any one of claims 1 to 2, comprising determining at least one root cause of an anomaly by performing one or more of the following in response to the anomalous event signal: checking which of the sensors in the sensor object have their measured output changed from their predicted output at or around the time of the occurrence of the anomalous event; analysing the time series residual signals of each sensor in the sensor object to determine one or more main drivers or most dominant contributors to a level signal value; analysing the time series residual signals of each sensor in the sensor object to determine one or more main drivers or most dominant contributors to an associated signal value; analysing the covariance of the time series residual signals of each pairwise combination of the sensor residual signals in the sensor object to determine one or more main drivers or most dominant contributors to the associated signal value.
4. A method according to any one of claims 1 to 3, comprising determining a health state of at least one device associated with said sensor object in response to said abnormal event signal.
5. 5. The method of claim 2, wherein either the predicted value or the prediction limit is provided by a sensor object model that is at least partially a predictive model trained using historical residual data of the sensor object.
6. The method of claim 5 , wherein one or more covariate signals are provided as inputs to the sensor object model, each covariate signal being a signal representing a parameter on which at least one of the residual signals depends.
7. 7. The method of claim 1, wherein the predicted output of at least one sensor is provided by a predictive state model, the predictive state model being a predictive model trained using historical time series output data of the respective sensor.
8. The method according to any one of claims 1 to 7, wherein the sensor objects are generated at least partly automatically by any of the one or more processing units using at least one self-organizing map.
9. 9. The method of claim 7, wherein the predicted state model is automatically selected by the processing unit by analyzing a number of different predictive model types and selecting the model type as the predicted state model that provides the smallest error between the output of that model when trained on a particular training window of the historical time series data and actual historical sensor outputs within a particular time window of the historical time series data.
10. A method according to any one of claims 1 to 9, wherein the level signal values are generated using a distance estimator which is indicative of the time and amount that the time series residual data deviates from its normal or expected or average state.
11. 11. The method of any one of claims 1 to 10, wherein the associated signal values are generated using a statistical measure of multivariate dependence in the residual data or measuring the variance of the time series residual data at a particular time.
12. Detecting, via any of the one or more processing units, a drift in the output signal of a sensor that is among at least a portion of the sensors, the drift being calculated from historical time series data of the sensor, the historical data of the sensor being for a period of at least one week, the drift being detected by calculating the strength, smoothness and recency of the historical data of the sensor.
13. A monitoring and / or control system for a plant including a plurality of sensors, comprising one or more processing units configured to carry out the method steps of any one of the preceding claims.
14. A computer program comprising instructions which, when executed by a processing unit of a plant monitoring and / or control system operatively connected to a plurality of sensors, cause said system to perform the method steps of any one of claims 1 to 12.
Citation Information
Patent Citations
Diagnostic system and method for predictive condition monitoring
JP2004531815A
Application of abnormal event detection technology to delayed coking unit
JP2009534746A
Monitoring system
JP2011090382A
Abnormality analysis method, program, and system
WO2018104985A1