Predicting the industrial aging process using machine learning methods

A data-driven approach using LSTM and ESN models predicts chemical process equipment degradation, improving maintenance planning and reducing downtime by analyzing current and planned conditions.

CN114787837BActive Publication Date: 2025-07-15BASF SE +1
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202080081945.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-11-26
Filing Date
2020-11-25
Publication Date
2025-07-15
Estimated Expiration
2040-11-25

AI Technical Summary

Technical Problem

Existing mechanical models for predicting the degradation of chemical production facilities struggle to accurately forecast the deterioration of process equipment in real-world industrial environments due to their reliance on laboratory data and inability to account for complex, unpredictable factors.

Method used

A data-driven approach using machine learning models, such as LSTM and ESN, to predict the degradation of chemical process equipment by analyzing current operational data and planned conditions, incorporating historical data and process parameters to estimate future performance indicators.

Benefits of technology

Enhances the accuracy of predicting equipment degradation, allowing for better maintenance planning and reducing unexpected downtime in chemical production facilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114787837B_ABST
    Figure CN114787837B_ABST
Patent Text Reader

Abstract

By accurately predicting the industrial aging process (IAP), such as the slow deactivation of catalysts in chemical plants, maintenance events can be scheduled further in advance, ensuring the cost - effectiveness and reliable operation of the plant. So far, these degradation progressions have typically been described by mechanical models or simple empirical prediction models. To accurately predict the IAP, data - driven models are proposed, comparing some traditional stateless models (linear and kernel ridge regression, and feed - forward neural networks) with more complex state - recursive neural networks (echo state network and long - short - term memory network). In addition, variants of the stateful models are discussed. In particular, stateful models that use mechanical pre - knowledge about the degradation dynamics (hybrid models). Stateful models and their variants may be more suitable for generating near - perfect predictions when trained on a sufficiently large dataset, while hybrid models may be more suitable for better generalization in the case of smaller datasets under changing conditions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a computer-implemented method and apparatus for predicting the deterioration progress of a chemical production plant. The present invention further relates to a computer program unit and a computer-readable medium. Background Art

[0002] The aging of critical assets is a prevalent phenomenon in any production environment, leading to substantial maintenance expenditures or production losses. Therefore, in both discrete manufacturing and the process industry, understanding and predicting potential deterioration progress is crucial for reliable and economical plant operation.

[0003] Focusing on the chemical industry, well-known aging phenomena include: deactivation of heterogeneous catalysts due to coking, sintering, or poisoning; blockage of process equipment (such as heat exchangers or pipes) on the process side due to coke layer formation or polymerization; fouling of heat exchangers on the water side due to microbial or crystalline deposits; corrosion of equipment (such as nozzles or pipes) installed in fluidized bed reactors; and so on.

[0004] For almost any important aging phenomenon in chemical engineering, the corresponding scientific community has a detailed understanding of its microscopic and macroscopic driving forces. This understanding is usually condensed into complex mathematical models. Examples of such mechanical deterioration models involve coking of steam cracking furnaces, sintering or coking of heterogeneous catalysts, or crystalline fouling of heat exchangers.

[0005] While these models provide valuable insights into the dynamics of experimentally inaccessible quantities and may help to verify or falsify hypotheses about general deterioration mechanisms, they may not be transferable to the specific environment of real-world plants or only with a significant amount of modeling effort: Broadly speaking, these models can typically describe the "clean" observations of deterioration progress in a laboratory environment and may not reflect the "dirty" reality of production, where additional effects are difficult or impossible to mechanically model. To give just one example, even in a "clean" system of Wulff-shaped particles on a flat surface, the sintering dynamics of supported metal catalysts are difficult to quantitatively model - and in a real heterogeneous catalyst, the surface morphology and particle shape may deviate significantly from this assumption. Therefore, mechanical models are rarely used in production environments to predict the deterioration dynamics of critical assets. Summary of the Invention

[0006] There may be a need to provide a reasonable prediction of the expected progress of the industrial aging process (IAP) of a chemical production plant.

[0007] The object of the invention is solved by the subject matter of the independent claims, wherein further embodiments are incorporated into the dependent claims. It should be noted that the aspects described below for the invention also apply to computer-implemented methods, devices, computer program units, and computer-readable media.

[0008] A first aspect of the invention relates to a computer-implemented method for predicting the deterioration progress of a chemical production plant. The method comprises:

[0009] a) receiving, via an input channel, currently measured process data indicative of current process conditions of a current operation of at least one chemical process equipment for a chemical production plant, wherein the at least one chemical process equipment has one or more deterioration key performance indicators (KPIs) for quantifying the deterioration progress of the at least one chemical process equipment;

[0010] b) receiving, via the input channel, one or more expected operation parameters indicative of planned operation conditions of the at least one chemical process equipment within a prediction range;

[0011] c) applying, by a processor, a data-driven model to an input data set comprising the currently measured process data and the one or more expected operation parameters to estimate future values of one or more deterioration KPIs within the prediction range, wherein the data-driven model is parameterized or trained according to a training data set, and wherein the training data set is based on a historical data set comprising process data and one or more deterioration KPIs; and

[0012] d) providing, via an output channel, the future values of the one or more deterioration KPIs within the prediction range, the future values being usable for monitoring and / or control.

[0013] In other words, a method is provided for predicting the short-term and / or long-term deterioration progress of one or more equipments of a chemical production plant based on the current process conditions and planned operation conditions of the chemical production plant. On a shorter time scale, the selected parameters may exhibit fluctuations that are not driven by the deterioration progress itself but by changing process conditions or background variables such as ambient temperature. In other words, one or more deterioration KPIs are largely determined by the process conditions rather than uncontrolled external factors such as a defective pipeline bursting. On a time scale longer than the typical production time scale, for example, the batch time of a discontinuous process or the typical time between set-point changes of a continuous process, the selected parameters change substantially monotonically to higher or lower values, indicating the occurrence of irreversible deterioration phenomena.

[0014] In some examples, the method may further include the steps of comparing future values of one or more KPIs with a threshold and determining a future time that meets the threshold. This time information can then be provided via an output channel or used for predictive maintenance events.

[0015] The method uses a data-driven model, such as a data-driven machine learning (ML) model, which does not involve prior physicochemical processes of one or more chemical process equipment in a chemical production plant. The data-driven model is capable of using one or more key performance indicators (KPIs) to predict short-term and long-term deterioration progress in a chemical production plant, as a function of input parameters that include one or more expected operating parameters indicating planned operating conditions of at least one chemical process equipment, and process data derived from sensors available in the production plant. A software product for performing the method is also provided. As an application example, the method can be used to predict and forecast at least one of the following deterioration progress in a chemical production plant: deactivation of a heterogeneous catalyst due to coking, sintering, and / or poisoning; blockage of chemical process equipment on the process side due to coke layer formation and / or polymerization; fouling of a heat exchanger on the water side due to microorganisms and / or crystalline deposits; and corrosion of equipment installed in a fluidized bed reactor.

[0016] A data-driven model refers to a trained mathematical model that is parameterized based on a training dataset to reflect the dynamics of the true degradation progression in a chemical production plant. In some examples, the data-driven model may include a data-driven machine learning model. As used herein, the term "machine learning" may refer to statistical methods that enable a machine to "learn" a task from data without explicit programming. Machine learning techniques may include "traditional machine learning" - a workflow where features are manually selected and then a model is trained. Examples of traditional machine learning techniques may include decision trees, support vector machines, and ensemble methods. In some examples, the data-driven model may include a data-driven deep learning model. Deep learning is a subset of machine learning that is loosely modeled after the neural pathways of the human brain. Depth refers to multiple layers between the input layer and the output layer. In deep learning, the algorithm automatically learns which features are useful. Examples of deep learning techniques may include convolutional neural networks (CNNs), recurrent neural networks (such as long short-term memory or LSTM), and deep Q networks. A general introduction to machine learning and corresponding software frameworks is described in "Machine Learning and Deep Learning frameworks and libraries for large-scale data mining: a survey"; Artificial Intelligence Review; Giang Nguyen et al., June 2019, Volume 52, Issue 1, pp 77-124. As will be explained below and particularly with respect to Figures 5 to 9 the exemplary embodiments shown in, the data-driven model may include a stateful model, which is a machine learning model with a hidden state that is continuously updated with new time steps and contains information about the entire past time series. Alternatively, the data-driven model may include a stateless model, which is a machine learning model whose predictions are based only on the inputs within a fixed time window prior to the current operation. In other words, the stateless model also relies on the past values of the degradation KPIs and the operating parameters on the input side. Alternatively, the data-driven model may include a hybrid model, i.e., a combination of a stateful model and a stateless model.

[0017] The at least one chemical process equipment can be one of the key components of a chemical production plant, as the health status of the key components has a significant impact on the maintenance activities of the chemical production plant. Sources of information regarding the selection of key components can be bad actor analysis or general operating experience. Examples of the deterioration progress of such chemical process equipment can include, but are not limited to: deactivation of heterogeneous catalysts due to coking, sintering, and / or poisoning; blockage of chemical process equipment on the process side due to coke layer formation and / or polymerization; fouling of heat exchangers on the water side due to microorganisms and / or crystalline deposits; and corrosion of equipment installed in fluidized bed reactors.

[0018] The at least one chemical process equipment can have one or more KPIs for quantifying its deterioration progress. The one or more deterioration KPIs can be selected from parameters including: parameters included in a set of measured process data and / or derived parameters representing functions of one or more parameters included in a set of measured process data. In other words, the one or more deterioration KPIs can include parameters directly measured using sensors (e.g., temperature sensors or pressure sensors). The one or more deterioration KPIs can alternatively or additionally include parameters indirectly obtained through proxy variables. For example, while catalyst activity is not directly measured in the process data, it manifests itself as a reduced yield and / or conversion rate of the process. The one or more deterioration KPIs can be defined by: a user (e.g., a process operator) or by a statistical model (e.g., an outlier score measuring the distance to the "health" state of the equipment in the multivariate space of relevant process data, such as Hotelling T 2 score or the DModX distance derived from principal component analysis (PCA)). Here, the health state can refer to the majority of states typically observed during a period of historical process data, which are marked as "normal" / "problem-free" / "good" by experts in the production process.

[0019] Process data can refer to quantities indicating the operating state of a chemical production plant. For example, such quantities can be related to measurement data collected during the production run of a chemical production plant and can be directly or indirectly derived from such measurement data. For example, process data can include sensor data measured by sensors installed in a chemical production plant, quantities directly or indirectly derived from such sensor data. Sensor data can include measured quantities available in a chemical production plant by means of installed sensors (e.g., temperature sensors, pressure sensors, flow sensors, etc.).

[0020] This set of process data may include raw data, which refers to basic, unprocessed sensor data. Alternatively or additionally, this set of process data may include processed or derived parameters directly or indirectly derived from the raw data. For example, while catalyst activity is not directly measured in the process data, it manifests itself as a reduced yield and / or conversion rate of the process. Examples of derived data for catalyst activity may include, but are not limited to: the average inlet temperature of multiple catalytic reactors derived from the corresponding temperature sensors, the steam-oil ratio derived from the raw data of steam flow and reactant flow, and any type of normalized data, such as production values normalized by catalyst volume or catalyst mass.

[0021] In the context of the current production run, the process data may include information about the current operating conditions, as reflected by the set of operating parameters, such as the feed rate of the reactor, which can be selected and / or controlled by plant personnel. As used herein, the term "current" refers to the most recent measurement, as measurements of some equipment may not be performed in real time.

[0022] Useful prediction ranges for equipment deterioration may be between hours and months. The prediction range applied can be determined by two factors. First, the prediction must be accurate enough to be used as a basis for decision-making. To achieve accuracy, input data for future production plans must be available, which is only available for a limited prediction range. Additionally, the prediction model itself may lack accuracy due to the potential prediction model structure or due to ill-defined model parameters, which may be the result of the noisy and limited nature of the historical data set used for model identification. Second, the prediction range must be long enough to address relevant operating issues, such as taking maintenance actions and making planning decisions.

[0023] The planned operating conditions may refer to the operating conditions under which the chemical production plant may operate in the future within the prediction range. The planned operating conditions are reflected by one or more expected operating parameters, which may be known and / or controllable within the prediction range, rather than uncontrolled external factors. Examples of uncontrolled external factors may include catastrophic events, such as the rupture of a defective pipeline. Other examples of uncontrolled external factors may include less catastrophic but more frequent external disturbances, such as changing external temperature or changing raw material quality. In other words, one or more expected operating parameters can be planned or anticipated within the prediction range.

[0024] One or more prospective operating parameters are employed, which can be used to simulate "what-if" scenarios, such as changes in process conditions, such as reducing feed load, feed composition, and reactor temperature within a prediction range. It should be noted that the proposed data-driven model does not infer future operating states from past and / or current operating states, but rather requires user input for one or more prospective operating parameters in order to interpret the changing operating conditions of a future plant. The use of prospective operating parameters can explain future changes in plant operation. Key performance indicators are functions of input parameters, including one or more prospective operating parameters indicating planned operating conditions of at least one chemical process unit and process data derived from sensors available in a production plant. By using prospective operating parameters, future loads, for example, can be included in the system for prediction. Allowing the values of future operating parameters to vary based on planning in the plant can provide additional degrees of freedom, which can improve the quality of the prediction model and can make the prediction more robust.

[0025] The input data set for the data-driven model can include current operating parameters. Current operating parameters can include raw data, which refers to basic, unprocessed sensor data. For example, the temperature and / or pressure in a reactor, the feed rate into the reactor, which can be selected and / or controlled by plant personnel. Alternatively or additionally, the set of process data can include processed or derived parameters, which are directly or indirectly derived from the raw data, such as a steam-to-oil ratio derived from raw data of steam flow and reactant flow, and any type of normalized data.

[0026] According to an embodiment of the present invention, the at least one chemical process unit operates in a cyclic manner including multiple runs. Each run includes a production phase, followed by a regeneration phase. The input data set includes at least one process information from the last run.

[0027] In other words, in the case of asset cyclic operation, the input data set can further include at least one process information from the last run. The last run can be the run before the "current run", where the current operation is used. Exemplary process information from the last run can include, but is not limited to, time on stream since the last regeneration, time on stream since the last exchange, process conditions at the end of the last run, regeneration duration of the last run, and duration of the last run. For prediction purposes, the input data set can additionally include information on planned operating conditions within the prediction range.

[0028] According to an embodiment of the present invention, the one or more degradation KPIs are selected from parameters including: parameters included in a set of measured process data and / or derived parameters representing functions of one or more parameters included in a set of measured process data.

[0029] According to an embodiment of the present invention, the selected parameter has at least one of the following characteristics: it tends to a higher or lower value in a substantially monotonic manner on a time scale longer than the typical production time scale, thereby indicating the occurrence of an irreversible degradation phenomenon and returning to the baseline after the regeneration phase.

[0030] The regeneration phase is a very important specific part of the process because, even without replacing the process equipment, the regeneration process may cause the KPI to return to its baseline. The existence of the regeneration phase leads to complex degradation behavior. In this case, the process equipment or catalyst may experience degradation on different time scales. We have a degradation behavior within a cycle, a regeneration phase at the end of the cycle, and at the same time we observe the degradation over the entire life cycle of the process equipment or catalyst charge. This will be explained below and in particular with respect to Figure 2 the example shown in

[0031] The existence of the regeneration phase can have an impact on the definition of the input parameters of the data-driven model. In this case, additional input parameters may be beneficial for improving the accuracy of the prediction.

[0032] Although there is a wide variety of affected asset types in a chemical production plant and the physical or chemical degradation progress behind them is completely different, the selected parameter representing one or more degradation KPIs may have at least one of the following characteristics:

[0033] On a time scale longer than the typical production time scale, such as the batch time of a discontinuous process or the typical time between setpoint changes in a continuous process, the selected parameter changes substantially monotonically to a higher or lower value, thereby indicating the occurrence of an irreversible degradation phenomenon. The term "monotonic" or "monotonically" refers to the increase or decrease of the selected parameter representing the degradation KPI on a longer time scale (such as the time scale of the degradation cycle), while the fluctuations on a shorter time scale do not affect this trend. On a shorter time scale, the selected parameter may exhibit fluctuations that are not driven by the degradation progress itself but by changing process conditions or background variables (such as ambient temperature). In other words, one or more degradation KPIs are largely determined by process conditions rather than uncontrolled external factors such as a defective pipeline bursting, external temperature changes, or changes in the quality of raw materials.

[0034] After the regeneration phase, the selected parameters may return to their baseline. As used herein, the term "regeneration" may refer to any event / procedure that reverses degradation, including replacement of the process equipment or catalyst, cleaning of the process equipment, in-situ reactivation of the catalyst, burning off of the coke layer, etc.

[0035] In an example, the degradation includes at least one of the following: deactivation of a heterogeneous catalyst due to coking, sintering, and / or poisoning; blockage of chemical process equipment on the process side due to coke layer formation and / or polymerization; fouling of a heat exchanger on the water side due to microorganisms and / or crystalline deposits; and corrosion of equipment installed in a fluidized bed reactor.

[0036] According to an embodiment of the present invention, the data-driven model includes a stateful model, which is a machine learning model with a hidden state that is continuously updated with new time steps and contains information about the entire past time series. Alternatively or additionally, the data-driven model includes a stateless model, which is a machine learning model whose prediction is based only on inputs within a fixed time window prior to the current operation.

[0037] A stateless model is a machine learning model whose prediction is based only on inputs within a past fixed time window. Examples of stateless models can include, but are not limited to, linear ridge regression (LRR), kernel ridge regression (KRR), and feedforward / forward neural networks (FFNN). LRR is an ordinary linear regression model with a regularization term added, which can prevent the weights from taking extreme values due to outliers in the training set. KRR is a non-linear regression model that can be derived from LRR using the so-called "kernel trick". Similar to LRR, FFNN learns a direct mapping between some input parameters and some output values. Stateless models (such as LRR, KRR, and FRNN) can accurately capture the instantaneous changes in the degradation KPIs due to changing process conditions. In addition, training a stateless model requires only a small amount of training data.

[0038] Compared with stateless models, stateful models explicitly use only the input x(t), rather than the past inputs x(t-1),..., x(t-k), to predict the output y(t) at a certain time point t. Instead, they maintain a hidden state h(t) of the system, which is continuously updated with each new time step and thus contains information about the entire past time series. Then, both the current input conditions and the hidden state of the model can be used to predict the output. Stateful models can include recurrent neural networks (RNNs), such as echo state networks (ESNs) and long / short-term memory networks (LSTMs). Stateful models can be beneficial for correctly predicting long-term changes.

[0039] According to an embodiment of the present invention, the stateful model includes a recurrent neural network (RNN).

[0040] Recurrent Neural Networks (RNNs) have a hidden state or "memory" that allows them to remember important features of the input signal, which only affect the output at a later time. This can be seen as an improvement over "memoryless" machine learning methods, as degradation phenomena can exhibit significant memory effects.

[0041] According to an embodiment of the present invention, the RNN includes at least one of the following: an Echo State Network (ESN) and a Long Short-Term Memory (LSTM) network.

[0042] RNNs are a powerful method for time series modeling. However, they can be difficult to train because their depth increases with the length of the time series. This can lead to gradient divergence during the error backpropagation training process, which can result in very slow convergence if the optimization fully converges ("vanishing gradient problem").

[0043] ESN is an alternative RNN architecture that bypasses the training-related problems of the above RNNs by training without using error backpropagation at all. Instead, ESN uses a very large randomly initialized weight matrix that, combined with the recurrent mapping of past inputs, essentially acts as a random feature expansion of the input; collectively called the "reservoir". Since the only parameters learned are the weights of the linear model for the final prediction, ESN can be trained on smaller datasets without much risk of overfitting.

[0044] Another exemplary architecture for dealing with the vanishing gradient problem in RNNs is the Long Short-Term Memory (LSTM) architecture. LSTM is trained using error backpropagation as usual, but avoids the vanishing gradient problem by using an additional state vector called the "cell state" along with the usual hidden state. Since modeling the gates that regulate the cell state requires multiple layers, LSTM may require a large amount of training data to avoid overfitting. Despite its complexity, the stability of the LSTM gradients makes it very suitable for time series problems with long-term dependencies.

[0045] According to an embodiment of the present invention, the state model includes a feedback state model that includes information about the predicted output or the true output from the previous time step into the input dataset for the current time step. The predicted output is one or more predicted KPIs at the previous time step. The true output is one or more measured KPIs at the previous time step.

[0046] Although it is possible to predict the key performance indicators (KPIs) of a process using only operating parameters, incorporating past KPIs as inputs can be a powerful new source of information, especially given the high autocorrelation of KPIs over time within the same cycle. One way to incorporate past KPIs into a stateful model such as an LSTM might be to include the predicted output or the true output (if available) from the previous time step into the input vector for the current time step.

[0047] According to an embodiment of the present invention, the input data set further includes an indicator variable that indicates whether the output from the data-driven model from the previous time step is a predicted output or a true output.

[0048] Including the predicted output (or true output) can lead to large prediction errors. The reason is that the predicted output is only an approximation of the true output and is therefore less reliable than the true output. Since the previous predicted output will be used for the next prediction, any small error in the predicted output value will propagate into the prediction of the next output. Over a long period of time, these small errors will accumulate and can result in a prediction that is very different from the true output time series, leading to very large errors. Therefore, it is crucial to distinguish between the reliable true output and the unreliable predicted output of the network so that the network can independently estimate the reliability of these two variables.

[0049] One way to achieve this can be to include an indicator variable alongside each feedback output value, which will indicate whether that output value is a true output, i.e., the actual measured KPI from the process, or a predicted KPI, i.e., the output from the stateful model at the previous time step. In other words, therefore, the feedback state model can be implemented simply by appending two values to the input vector at each time step: the output value from the previous time step, and the indicator variable, which is 0 if the feedback value is the truly measured KPI, or 1 if the feedback value was predicted by the stateful model at the previous step.

[0050] According to an embodiment of the present invention, step a) further includes receiving previously measured process data that indicates the past process conditions of at least one chemical process equipment of a chemical production plant during a predetermined period before the current operation. Step b) further includes receiving one or more past operating parameters that indicate the past process conditions of at least one chemical process equipment during a predetermined period before the current operation. In step c), the input data set further includes the previously measured process data and one or more past operating parameters.

[0051] Previously measured process data can also be referred to as lag data. Thus, the stateless model is more robust. In contrast, a stateless model without lag variables represents a system that responds specifically to current events. The model developer can select, for example, a predefined period before the current operation of the lag data according to the type of equipment. For example, the predefined period can be 5%, 10%, or 15% of the typical time period between two maintenance actions of the equipment.

[0052] According to an embodiment of the present invention, the stateless model includes at least one of the following: linear ridge regression (LRR), kernel ridge regression (KRR), and feedforward neural network (FFNN).

[0053] According to an embodiment of the present invention, the data-driven model is a hybrid model that includes a stateful model for predicting the degradation trend of one or more degradation KPIs and a stateless model for predicting the additional instantaneous impact of operating parameters on one or more degradation KPIs. The degradation trend represents a monotonic change in the performance of a chemical process equipment over a time scale longer than the typical production time scale. The additional instantaneous impact of the operating parameters does not include the time delay of the effect of the model input on one or more degradation KPIs.

[0054] In this way, the stateful model (e.g., RNN) improves data efficiency by providing mechanistical pre-information about the process. To make the learning problem of the stateful model simpler, the problem is divided into predicting the short-term or instantaneous effect and the long-term behavior of the degradation KPI.

[0055] In the basic problem setting of predicting an industrial aging process (IAP), all considered processes are affected by some potential degradation progress, which reduces process efficiency over time. Since this degradation is long-term and occurs throughout the cycle, it is difficult to predict because it is affected by the conditions in the early stage of the cycle, but this correlation is largely unknown and difficult to learn due to the large time lag. However, since engineers generally understand the basic dynamics of the degradation progress, some parametric prototype functions can be used to parameterize the degradation of the KPI, and its parameters can be fitted to perfectly match the degradation curve of a given cycle. To make the learning problem of the stateful model simpler, the problem is divided into predicting the instantaneous effect and the long-term effect of the input on the KPI.

[0056] One way to isolate the transient effect can be to train a linear model without any time information. For example, when the effect of degradation is still small and the time variable is not used as an input, the LRR model can be trained only on the initial time period of the cycle, so that the model does not attempt to learn from the time context but only from the transient effect of the input on the KPI. For example, the initial time period of the cycle can be the initial 1%-10% of the entire cycle, preferably 1%-5%, where the degradation effect can be expected to be negligible. Although this method will only learn the linear transient effect, usually this is sufficient to remove most of the transient artifacts from the cycle, so that the residual reflects the degradation curve. As previously mentioned, the residual can then be modeled using a parametric prototype function, and the parameters of this function will fit each degradation curve. In this way, instead of predicting the individual values at each time point of the degradation trend, which is usually highly non-stationary, only a stateful model is used to predict a set of parameters for each cycle, and these parameters are used in the prototype function to model the entire degradation curve. This in turn makes the learning problem more constrained, because only functions in the form given by the prototype can be used to model the degradation.

[0057] According to an embodiment of the present invention, the stateful model includes a combination of mechanical pre-information about a process represented by a function with a predefined structure and a stateful model that estimates the parameters of the function.

[0058] Wherein the mechanical pre-information is represented by a physics-based model, including ordinary differential equations or partial differential equations (ODE / PDE) and linear or nonlinear algebraic equations, such as heat or mass balance equations.

[0059] According to an embodiment of the present invention, the stateless model includes a linear model.

[0060] The linear model can be used to capture transient linear correlations, while the stateful model can be used to capture long-term degradation trends. In an example, the linear model can include LRR.

[0061] In some examples, the hybrid model further includes a nonlinear model. Generally, the linear model only captures transient linear correlations, while the stateful model will ideally capture long-term degradation trends. However, since the prototype function may not always perfectly fit the degradation, and there will still be some artifacts that are nonlinear or transient and thus cannot be captured by the linear model, we need a nonlinear model, such as LSTM, which will attempt to model these additional short-term artifacts separately at each time point. In other words, a linear model and two stateful models can be combined in the hybrid model. For example, one LSTM is used for long-term degradation and one LSTM is used for short-term artifacts, and we name this model a two-speed model.

[0062] According to an embodiment of the present invention, the input data set further includes at least one transformed process data, which represents a function of one or more parameters of the currently measured process data and / or the previously measured process data.

[0063] In other words, engineering features constructed from the process data can be used as additional inputs. These engineering features may include the runtime since the last regeneration (e.g., catalyst or heat exchanger), the runtime since the last exchange (e.g., catalyst or heat exchanger), the process conditions at the end of the last run, the regeneration duration of the last run, the duration of the last run, and so on.

[0064] In some examples, the historical data may include one or more transformed process data, which encode information about the long-term effects of the degradation of at least one chemical process equipment. The method may further include estimating future values of at least one key performance indicator within the prediction range of multiple runs. In other words, these engineering features may be particularly relevant because they may encode information about long-term effects in the system, such as coke residues accumulated on time scales of months and years. By including these long-term effects in the historical data, a data-driven model can be trained to predict the degradation in the current operating cycle and the long-term effects of the degradation on multiple operating cycles.

[0065] A second aspect of the present invention relates to a device for predicting the degradation progress of a chemical production plant. The device includes an input unit and a processing unit. The input unit is configured to receive the currently measured process data, which indicates the current process conditions of the current operation of at least one chemical process equipment in the chemical production plant, wherein at least one chemical process equipment has one or more degradation key performance indicators (KPIs) for quantifying the degradation progress of at least one chemical process equipment. The input unit is further configured to receive one or more expected operation parameters, which indicate the planned process conditions of at least one chemical process equipment within the prediction range. The processing unit is configured to perform the method steps as described above and below.

[0066] A third aspect of the present invention relates to a computer program unit for instructing the device as described above and below. The computer program unit is adapted to perform the method steps as described above and below when executed by the processing unit.

[0067] A fourth aspect of the present invention relates to a computer-readable medium storing the program unit.

[0068] As used herein, the term "aging" may refer to the effect of a component undergoing some form of material degradation and damage (usually but not necessarily associated with the time of use), with an increased likelihood of failure within the service life. Aging equipment refers to equipment that has evidence or the potential for significant degradation and damage since it was new, or for which there is insufficient information and knowledge to understand the extent to which that potential exists. The importance of aging and damage is related to the potential effects on equipment functionality, availability, reliability, and safety. Just because a piece of equipment is old does not necessarily mean it has significantly degraded and been damaged. All types of equipment can be vulnerable to aging mechanisms. Generally, an aging plant is one that is considered or potentially no longer considered fully fit for purpose due to degradation or obsolescence of its integrity or functional performance. "Aging" is not directly related to actual age.

[0069] As used herein, the term "degradation" may refer to the potential degradation of plants and equipment due to aging-related mechanisms such as coking, sintering, poisoning, fouling, and corrosion.

[0070] As used herein, the term "algorithm" may refer to a set of rules or instructions that cause a trained model to perform the operation you want it to perform.

[0071] As used herein, the term "model" may refer to a training process that predicts an output given a set of inputs. BRIEF DESCRIPTION OF THE DRAWINGS

[0072] These and other aspects of the present invention will become apparent and be further elucidated from the following description by way of example embodiments described and with reference to the accompanying drawings, in which

[0073] Figure 1 shows a flowchart illustrating a computer-implemented method for predicting the progression of degradation in equipment of a chemical production plant.

[0074] Figure 2 shows an exemplary degradation behavior of process equipment in the presence of a regeneration stage.

[0075] Figure 3 shows an example of an industrial aging process (IAP) prediction problem.

[0076] Figure 4 shows an example of a one-month synthetic dataset showing the loss of catalytic activity in a fixed-bed reactor.

[0077] Figure 5 shows an example of one-month historical data of a real-world dataset showing the pressure loss Δp across a reactor.

[0078] Figure 6 shows a comparison of stateless and stateful models for time series prediction.

[0079] Figure 7 Shows an example of the ESN architecture.

[0080] Figure 8 Shows an example of the LSTM architecture.

[0081] Figure 9 Shows an example of the feedback state model.

[0082] Figure 10 Shows an example of the hybrid model.

[0083] Figure 11 Shows the mean squared error (MSE) of each of the five models (LRR, KRR, FFNN, ESN, and LSTM) on training and test sets of different training set sizes.

[0084] Figures 12A - 12D Shows graphs of the true and predicted conversion rates of the models LRR, KRR, FFNN, ESN, and LSTM for randomly selected periods from the training and test sets.

[0085] Figures 13A - 13B Shows graphs of the predicted and true KPIs of the feedback LSTM for randomly selected training and test samples from two data sets.

[0086] Figures 14A - 14B Shows graphs of the predicted and true KPIs of an example of the hybrid model for randomly selected training and test samples from two data sets.

[0087] Figure 15 Shows the mean squared error of some models on the training and test sets.

[0088] Figure 16 Schematically shows an apparatus for predicting the deterioration progress in equipment of a chemical production plant.

[0089] Figure 17 Schematically shows a system for predicting the deterioration progress in equipment of a chemical production plant.

[0090] It should be noted that these figures are purely schematic and not drawn to scale. In the drawings, elements corresponding to those already described may have the same reference numerals. Examples, embodiments, or alternative features, whether indicated as non - restrictive or not, should not be construed as limiting the claimed invention. Detailed Description

[0091] Figure 1 Shows a flowchart illustrating a computer - implemented method 100 for predicting the deterioration progress of a chemical production plant.

[0092] In step 110, which is step a), the currently measured process data is received via an input channel. The currently measured process data indicates the current process conditions of the current operation of at least one chemical process equipment in a chemical production plant.

[0093] In some examples, at least one chemical process equipment can be operated in a cyclic manner including multiple runs. Each run includes a production phase followed by a regeneration phase.

[0094] Figure 2 Illustrates an exemplary degradation behavior of process equipment in the presence of a regeneration phase. The solid line 10 represents the degradation behavior within one cycle, while the dashed line 12 represents the degradation of the process equipment over the entire life cycle of the process equipment (e.g., within one catalyst charge). In Figure 2 the example, the process equipment is operated in a cyclic manner, including 11 cycles over the entire life cycle of the process equipment. Each cycle has a production phase 14 followed by a regeneration phase 16. The regeneration phase 16 is a very important specific part of the process because this may cause the KPI to return to its baseline (indicated by the dashed line 12) after the regeneration process, even without replacing the process equipment. The presence of the regeneration phase can lead to complex degradation behavior. In this case, the process equipment or catalyst can experience degradation on different time scales. As Figure 2 shown, degradation behavior is observed within one cycle, with a regeneration phase at the end of the cycle, and at the same time degradation is observed over the entire life cycle of the process equipment or catalyst charge.

[0095] The presence of the regeneration phase also has an impact on the definition of the input parameters of the data-driven model. In this case, additional input parameters may be required to improve the accuracy of the prediction. For example, process information from the last run can be provided as an additional input parameter. The process information from the last run can further include at least one of the following: run time since the last regeneration (e.g., catalyst or heat exchanger), run time since the last exchange (e.g., catalyst or heat exchanger), process conditions at the end of the last run, regeneration duration of the last run, duration of the last run, etc.

[0096] In an example, the process data can include sensor data that can be obtained from a chemical production plant. Examples of sensor data can include but are not limited to temperature, pressure, flow rate, level, and composition. For the equipment, appropriate sensors can be selected to provide information about the health state of the equipment under consideration. Alternatively or additionally, the process data can include quantities directly or indirectly derived from such sensor data, i.e., one or more derived parameters that are functions of one or more parameters included in a set of measured process data.

[0097] Back to Figure 1, in step 120, i.e., step b), one or more expected operating parameters indicating the planned operating conditions of at least one chemical process equipment within the prediction range are received via the input channel. The one or more expected parameters may be known and / or controllable within the prediction range. In other words, one or more expected operating parameters can be planned or anticipated within the prediction range. Step 110 and step 120 can be executed sequentially or in parallel.

[0098] In step 130, i.e., step c), a data-driven model is applied by the processor to an input data set including current measured process data and one or more expected operating parameters to estimate future values of one or more degradation KPIs within the prediction range. The data-driven model is parameterized or trained according to a training data set. The training data set is based on a historical data set including one or more degradation KPIs and process data of one or more chemical process equipments, wherein the one or more chemical process equipments are operated in a cyclic manner including multiple runs, and each run includes a production phase followed by a regeneration phase. The historical data set may include data from multiple runs and / or multiple plants.

[0099] The one or more degradation KPIs may be selected from parameters including parameters contained in a set of measured process data. Alternatively or additionally, the one or more degradation KPIs are selected from parameters including derived parameters that represent functions of one or more parameters contained in a set of measured process data.

[0100] Although the types of affected asset types are diverse and the physical or chemical degradation progress behind them is completely different, all these phenomena may share some of the following basic characteristics:

[0101] 1. The key assets under consideration have one or more key performance indicators (KPIs) for quantifying the degradation progress.

[0102] 2. On a time scale much longer than the typical production time scale (i.e., the batch time of a discontinuous process; the typical time between setpoint changes of a continuous process), the KPI drifts more or less monotonically to higher or lower values, indicating the occurrence of an irreversible degradation phenomenon. (On a shorter time scale, the KPI may exhibit fluctuations that are not driven by the degradation progress itself but by changing process conditions or background variables (such as environmental temperature).) For example, as indicated by the solid line 10 Figure 2 The degradation KPI shown in drifts monotonically to lower values, indicating the occurrence of an irreversible degradation phenomenon.

[0103] 3. The KPI returns to the baseline after a maintenance event (such as cleaning of a fouled heat exchanger, replacement or regeneration of an inactive catalyst, etc.). For example, Figure 2The degradation KPIs shown in [Figure 0] return to their baseline (indicated by the dashed line 12) after the regeneration process, even without replacing the process equipment.

[0104] 4. The degradation is not a "bolt from the blue" - such as, for example, a defective pipe bursting - but is driven by creep, inevitable wear and tear of the process equipment.

[0105] The present disclosure addresses any aging phenomenon having these general characteristics. The asset can be operated in a cyclic manner including multiple runs, where each run includes a production phase followed by a regeneration phase.

[0106] Characteristic (4) indicates that the evolution of the degradation KPIs is largely determined by the process conditions rather than uncontrolled external factors. This defines the central problem addressed by the present disclosure: given the planned process conditions within that time frame, develop an accurate model to predict the evolution of the degradation KPIs over a specific time frame.

[0107] Use a pre-trained data-driven model to determine the expected degradation behavior of chemical process components (i.e., individual assets such as heat exchangers or reactors) under expected operating conditions. Based on predefined end-of-run criteria, predict the end of the run (e.g., switching from the production phase to the regeneration phase, catalyst replacement).

[0108] In step 140, i.e., step d), future values of one or more degradation KPIs within the prediction range are provided via an output channel, which can be used for monitoring and / or control.

[0109] Based on this information, necessary control measures can be implemented to prevent unplanned production losses due to degradation or failure of the process equipment. For example, the future values of one or more KPIs can be compared with a threshold to determine the future time when the threshold is met. This time information can then be provided via the output channel or used to predict maintenance events. In this way, the planning and coordination of downtime between different chemical process equipment can be improved, for example, by avoiding parallel downtime of two or more chemical equipment. Typically, the data used for the prediction model in this context is created by sensors in the factory near the production process.

[0110] Below, we disclose examples of data-driven models for the IAP prediction task, comparing some traditional stateless models (including LRR, KRR, and FFNN) with more complex stateful recurrent neural networks ESN and LSTM. In addition, we also evaluate feedback state models such as feedback LSTM and hybrid models. To examine how much historical data each model requires in the training model, we first examine their performance on synthetic datasets with known dynamics. Then, in a second step, the models are tested on real data from a large chemical plant.

[0111] 1. Problem Definition

[0112] The general industrial aging process (IAP) prediction problem is as Figure 3 shown: The aim is to model the upcoming evolution of one or more deteriorations within the upcoming time window t ∈ [0, T (referred to as the i-th deterioration cycle) as a function of the planned process conditions during this cycle i :

[0113]

[0114] where ∈ i (t) represents the random noise of the deterministic relationship between the disturbance x i and y i .

[0115] Figure 3 shows the industrial aging process (IAP) prediction problem. The deterioration KPI (e.g., the pressure drop Δp in a fixed bed) increases over time (e.g., due to coking) and is affected by the (manually controlled) process conditions I and II (e.g., the reaction temperature T and the flow rate F). Although this example shows two process parameters, the claimed method is also applicable to one process parameter or more than two process parameters. The KPI recovers after a maintenance event, which divides the time axis into different deterioration cycles. The IAP prediction task is to predict the evolution of the KPI, i.e., the target (dependent) variable y i (t), given the upcoming process conditions, i.e., the input (independent) variable x i (t), in the current cycle i.

[0116] The deterioration phenomenon can exhibit significant memory effects, meaning that a certain input pattern x(t) may only affect the output y(t′) at a later time t′ > t. In addition, these memory effects can also occur on multiple time scales, making these processes very difficult to model. As an example, consider a heat exchanger with coking on the inner tube wall. The observed heat transfer coefficient is used as the KPI yi (t), and process conditions x i (t) includes the mass flow rate, chemical composition, and temperature of the fluid being processed. The time range is one cycle between two cleaning procedures (e.g., burning off). If, at an early time t1 in the cycle, an adverse combination of low mass flow rate, high coke precursor content, and high temperature occurs, a first coke spot may form at the wall, which is not yet sufficient to significantly affect heat transfer. However, they act as nuclei for further coke formation in the mid- to late-cycle, such that compared to when the process conditions were not adverse near t1 but had very similar process conditions throughout the rest of the cycle, y i (t) decreases more rapidly at t > t1.

[0117] Another complication can be that in real-world application cases, the distinction between the deteriorating KPI y, the process conditions x, and the uncontrolled influencing factors is not always clear. For example, consider the case of heterogeneous catalyst deactivation, where the loss of catalytic activity results in a reduced conversion rate. In this case, the conversion rate can be the target deteriorating KPI y, and the process conditions (such as temperature) that are manually controlled by the plant operator would be considered the input variables x of the model. However, the plant operator may try to keep the conversion rate at a certain setpoint, which can be achieved by increasing the temperature to counteract the effect of catalyst deterioration. This introduces a feedback loop between the conversion rate and the temperature, meaning that the temperature is no longer considered an independent variable because its actual value may depend on or be partially dependent on the target. Therefore, care may have to be taken because including such dependent variables as input x in the model can lead to reporting overly optimistic prediction errors, which will not hold up when the model is later used in reality.

[0118] 2. Data Set

[0119] To gain insight into and evaluate different machine learning models for the IAP prediction problem, we considered two datasets: one is a synthetic dataset that we generated ourselves using a mechanistic model, and one containing real data from a large BASF plant. These two datasets will be described in more detail below.

[0120] The reason for using synthetic data is that it enables us to control two important aspects of the problem: the amount of data and the quality of data. The amount of data, measured, for example, by the number of catalyst life cycles in the dataset, can be made arbitrarily large for synthetic data to test machine learning methods that require the most data. The quality of data refers to the level of noise in the dataset, or in other words, the extent to which the deteriorating KPI y(t) is uniquely determined by the process conditions x(t) provided in the dataset. In a synthetic dataset based on a deterministic degradation model, we know that there is a functional mapping between x and y, i.e., there is no fundamental reason that can prevent the machine learning model from learning this relationship and eliminating prediction errors. In contrast, for real data, bad prediction errors may be due to problems with the method and / or the dataset, which may not contain enough information on the input side x to accurately predict the output quantity y.

[0121] 2.1 Synthetic Data Set

[0122] In the following example, a synthetic dataset is used to simulate process data from a reactor experiencing catalyst deactivation and periodic regeneration. For the synthetic dataset, we model the widespread phenomenon of slow but steady loss of catalytic activity in a continuously operating fixed-bed reactor. Eventually, catalyst deactivation leads to unacceptable conversion or selectivity rates in the process, requiring catalyst regeneration or replacement, which marks the end of a cycle.

[0123] The chemical process in the reactor under consideration is the gas-phase oxidation of olefins. To generate time series of all variables, we used a mechanistic process model with the following components:

[0124] ■ Mass balance equations for all five relevant chemical species (olefin reactant, oxygen, oxidation product, CO2, water) in the reactor, which are modeled as an isothermal plug-flow reactor under the assumption of the ideal gas law for simplicity. The reaction network consists of a main reaction (olefin + O2 → product) and a side reaction (olefin combustion to form CO2).

[0125] ■ A highly non-linear deactivation law for the catalyst activity, which depends on the reaction temperature, flow rate, and incoming oxygen, as well as the activity itself.

[0126] ■ Kinetic laws for the reaction rates.

[0127] ■ A stochastic process that determines the process conditions (temperature, flow rate, etc.).

[0128] Based on the current process conditions and hidden state of the system, the mechanical model generates a multivariate time series [x(t), y(t)] of approximately 2000 degradation cycles. The final dataset includes five operating parameters (mass flow rate, reactor pressure, temperature, and mass fractions of two reactants, olefin and O2) as input x(t) for each time point t, and two degradation KPIs y(t) (conversion and selectivity).

[0129] To give an impression of the simulated time series, Figure 4 shows data for one month, which shows a synthetic dataset for one month, showing the loss of catalytic activity in a fixed-bed reactor. At each time point t, the process condition vector x(t) includes the reactor temperature T, mass flow rate F, reactor pressure p, and mass fraction μ of the reactants at the reactor inlet z . The degradation KPI y(t) is the conversion rate and selectivity of the process. The duration of the deactivation period is approximately 8 - 10 days. The catalyst activity A(t) is the hidden state and is therefore not part of the dataset, but is only used to illustrate the dynamics of the problem: the system output y(t) (selectivity and conversion rate) is affected not only by the current operating parameters x(t), but also by the current catalyst activity A(t), which decreases non-linearly in each cycle.

[0130] In addition to the operating parameters, the cumulative feed of olefin in the current cycle is also added to the dataset as a potential input variable. This variable is generally regarded as a rough predictor of catalyst activity. Therefore, it is usually calculated and monitored in the plant. In the language of machine learning, this variable represents an engineered feature of the original input time series. In this way, some basic domain knowledge about catalyst deactivation is added to the dataset.

[0131] 2.2 Real - World Data Set

[0132] The second dataset contains process data for the production of organic substances in a BASF world-class continuous production plant. The process is a gas-phase oxidation in a multitubular fixed-bed reactor.

[0133] The catalyst particles in the reactor suffer degradation, which is coking in this example, i.e., surface deposition of elemental carbon in the form of graphite. This results in reduced catalytic activity and increased fluid resistance. The latter is a more serious consequence and leads to an increased pressure drop across the reactor, as measured by the difference Δp in gas pressure before and after the reactor. In this example, the KPI is the pressure drop.

[0134] When Δp exceeds a predefined threshold, the so-called end-of-run (EOR) criterion is reached. Then, the coke layer is burned off during a dedicated regeneration process by injecting air and additional nitrogen into the reactor at an elevated temperature for different numbers of hours. Operational reasons may cause a delay in burning off when Δp exceeds the EOR threshold, and vice versa, premature burning off when Δp has not yet reached the EOR threshold. Figure 5 Some exemplary cycles of Δp are shown, which Figure 5 shows one month of historical data of a real-world dataset, showing the pressure loss Δp on the reactor, which is the degradation KPI y(t) in this IAP prediction problem. When Δp reaches the order of magnitude of the EOR threshold of 70 mbar, the coke deposits are burned off, which marks the end of the cycle.

[0135] Since the coke is not completely removed by this burning-off process, coke residues accumulate continuously during the regeneration process, making the pressure drop problem more severe. Therefore, the entire catalyst bed must be replaced every 6 - 24 months.

[0136] As an option, the historical data may include one or more transformed process data that encodes information about the long-term effects of degradation on at least one chemical process equipment. The method may further include estimating future values of at least one key performance indicator within the prediction range of multiple runs. Therefore, these engineering features may be particularly relevant as they can encode information about long-term effects in the system, such as coke residues that accumulate on time scales of months and years. By including these long-term effects in the historical data, data-driven models can be trained to predict degradation in the current operating cycle, as well as the long-term effects of degradation on multiple operating cycles.

[0137] Factors that may affect the coking rate are:

[0138] 1. The mass flow rate F through the reactor ("feed load")

[0139] 2. The ratio of organic reactants to oxygen in the feed

[0140] 3. The intensity of the previous regeneration process

[0141] 4. The length of the previous degradation cycle

[0142] The dataset contains 7-year process data from four most relevant sensors extracted from the plant's plant information management system (PIMS), as listed in Table 1. Assuming a time scale of 4 to 7 days between two burning-off processes, this corresponds to 375 degradation cycles belonging to three different catalyst batches. The sampling rate of all variables is 1 / hour, and linear interpolation is performed on this time grid.

[0143] Variable Name Unit Description Type PD mbar Pressure difference Δp on the reactor y T ℃ Reaction temperature x F_R kg / h Inflow of organic reactant into the reactor x F_AIR kg / h Mass flow rate of air into the reactor x

[0144] Table 1

[0145] The task is to predict the pressure drop Δp caused by coking during the entire remaining duration of the prediction period at the mid - point t of the degradation cycle. Particular attention is paid to the prediction of the time point t k at which the EOR threshold Δp max = 70 mbar is reached. EOR

[0146] As mentioned above, several relevant operating parameters can be used as input variables x(t) of the model (see Table 1). In addition, engineering features constructed based on these operating parameters or the degradation KPIs Δp in previous cycles can be used as additional inputs. Examples of these additional inputs are listed in Table 2 below:

[0147]

[0148] Table 2

[0149] 3. Input Quantities

[0150] For an asset, key performance indicators directly or indirectly related to the degradation state are required. For each prediction, process data measured for chemical process elements are needed. Such process data can include current process conditions. The at least one chemical process equipment can be operated in a cyclic manner including multiple runs. Each run includes a production phase followed by a regeneration phase. The input data set of the data - driven model can further include at least one process information from the last run, such as the run time since the last regeneration (e.g., catalyst or heat exchanger), the run time since the last exchange (e.g., catalyst or heat exchanger), the process conditions at the end of the last run, the regeneration duration of the last run, the duration of the last run, etc. The key performance indicators are parameters provided as process data or derived from the provided process data. It is necessary to predict the expected operating conditions (such as flow rate, controlled reaction temperature) of the current production run of the chemical process element.

[0151] 4. Model Architecture

[0152] We will now formulate the IAP prediction problem in a machine - learning environment. To this end, the mapping defined in Equation (1) is expressed as a specific function f that is based on the process conditions x at this time and possibly up to k hours before t i : and returns i.e., the KPI estimate at the time point t in the i - th degradation cycle,

[0153]

[0154] The task is to predict the entire cycle (i.e., until T​i )'s y i (t), typically starting approximately 24 hours after the last maintenance event that ended the previous cycle.

[0155] In equation (2), the prediction function f is defined as a function of the current and past input variables x i Since the value of the typically deteriorating KPI y i is known at least for the first 24 hours of each cycle, in principle the set of input variables for f can be extended to also include y i (t′), where t′ < t. However, while this may improve the prediction at the start of the cycle, since our goal is to predict the full cycle starting from the first 24 hours, for the prediction at most time points, instead of the actual value y i (t′) can be used as input, but their predicted values must be used Since these predicted values typically contain at least one small error, the prediction for future time points will be based on input data with increasing noise, as the prediction errors in the input variables will accumulate rapidly. Thus, the only explicit input to the model is the predefined process conditions x i . However, the model variant discussed in Section 4.3 (“ Feedback State Model ”) overcomes this limitation.

[0156] Thus, the exact form of the function f depends on the type of machine learning method chosen for the prediction task. However, while the chosen machine learning model determines the form of the function, its exact parameters need to be adjusted to fit the dataset at hand in order to produce accurate predictions. To this end, the available data is first split into so-called “training” and “test” sets, where each of the two sets contains the entire multivariate time series of several mutually exclusive deterioration cycles from the original dataset, i.e., multiple input-output pairs consisting of the planned conditions x of the given process and the deteriorating KPI y Then, using the data in the training set, the machine learning algorithm learns the optimal parameters of f by minimizing the expected error between the predicted KPI and the true KPI y i (t). After the machine learning model has been trained, i.e., when f predicts y i (t) as accurately as possible on the training set, the model should be evaluated on new data to indicate its performance when used in actual applications later. To this end, the test set is used. If the performance on the training set is much better than that on the test set, the model does not generalize well to new data and is considered to be “overfitting” on the training data.

[0157] In addition to the regular parameters of f, many machine learning models also require setting some hyperparameters, such as determining the degree of regularization (i.e., the impact of possible outliers in the training set on the model parameters). To find adequate hyperparameters, cross-validation can be used: here, in multiple iterations, the training set is further split into validation and training parts, and a model with a specific hyperparameter setting is trained on the training part and evaluated on the validation part. Then, when training the final model on the entire training set, those hyperparameter settings that produced the optimal results on the validation split are used, and it is then evaluated on the reserved test set as described above.

[0158] Machine learning models for time series prediction can be divided into two main subgroups: stateless models and stateful models.

[0159] Figure 6 Comparison of stateless and stateful models showing time series casting. Figure 6 (a) Shows a stateless model that is based on predictions of information contained in a fixed past time window, while Figure 6 (b) Shows a stateful model where hidden states are used to maintain and propagate information about the past.

[0160] Stateless models directly predict the output given the current input, independent of the predictions at the previous time point. On the other hand, stateful models maintain an internal hidden state of the system that encodes information about the past and can also be used in addition to the current process conditions when making predictions.

[0161] Stateless models include the most typical machine learning regression models ranging from linear regression models to various types of neural networks. The stateless regression models that we will explore in this article are linear ridge regression (LRR), kernel ridge regression (KRR), and feedforward neural networks (FFNN), i.e., one linear and two non-linear prediction models. The most commonly used stateful models for sequence data modeling are recurrent neural networks (RNNs). Although RNNs are some of the most powerful neural networks capable of approximating any function or algorithm, they are also more involved in training. Therefore, in this article, we choose to use two different RNN architectures to model IAP that are specifically designed to handle the problems that arise when training conventional RNNs: echo state networks (ESNs) and long short-term memory (LSTM) networks.

[0162] In addition, two main variants of the basic stateful model are introduced to improve the performance on real-world datasets: including feedback loops that incorporate past prediction outputs as additional inputs; and splitting the model into two or more different models that separate different aspects of the prediction output (e.g., instantaneous effects versus long-term trends).

[0163] The following paragraphs introduce seven machine learning models. For simplicity, in many cases we write only x and y, omitting the reference to the current cycle i and time point t under discussion, where x may include process conditions at multiple time points in a past fixed time window (i.e., up to t - k).

[0164] 4.1 Stateless Model

[0165] A stateless model is a machine learning model whose prediction is based only on the input within a past fixed time window, i.e., exactly as described in Equation (2).

[0166] Linear Ridge Regression (LRR)

[0167] LRR is an ordinary linear regression model with a regularization term added, which prevents the weights from taking extreme values due to outliers in the training set. The target variable y is predicted as a linear combination of the input variables x, i.e.,

[0168]

[0169] where is the weight matrix, i.e., the model parameters of f learned from the training data. The simple model architecture, global optimal solution, and regularization of LRR all contribute to reducing the overfitting of the model. In addition, training and evaluating the model are not computationally expensive, which also makes it a feasible model for processing large amounts of data. Although they are relatively simple, linear models are widely used in many application scenarios and can usually be used to approximate real-world processes with a fairly high degree of accuracy, especially if additional (non-linear) hand-designed features are available. In addition, considering the limited amount of training data usually available for real-world IAP problems, it is necessary to be very careful to reliably estimate the parameters of more complex non-linear prediction models (such as deep neural networks), while linear models provide a more robust solution because they provide a global optimal solution and are less likely to overfit considering their linear nature. For a detailed discussion of the LRR model, please refer to the following publications: Draper NR, Smith H. Applied regression analysis, vol. 326. John Wiley & Sons; 2014, and Bishop CM, Nasrabadi NM. Pattern Recognition and Machine Learning. Journal of Electronic Imaging 2007; 16(4).

[0170] Kernel Ridge Regression (KRR)

[0171] KRR is a non - linear regression model that can be derived from LRR using the so - called "kernel trick". Instead of using the regular input features x, it uses a feature map φ corresponding to a certain kernel function k to map the features into a high (and possibly infinite) - dimensional space such that φ(x) T φ(x′)=k(x,x′). By computing the non - linear similarity k between a new data point x and the training samples x j where j = 1,...,N, the target y can be predicted as

[0172]

[0173] where α j are the learned model parameters.

[0174] Compared with LRR, the non - linear KRR model can accommodate more complex data, and the fact that the global optimal solution can be obtained analytically makes KRR one of the most commonly used non - linear regression algorithms. However, the performance of the model is also relatively sensitive to the choice of hyperparameters, so the hyperparameters need to be carefully selected and optimized. In addition, the fact that computing the kernel matrix scales quadratically with the number of training examples N makes it difficult to apply KRR to large training sets. For a detailed discussion of the KRR model, please refer to the following publications: Draper NR, Smith H. Applied regression analysis, vol. 326. John Wiley & Sons; 2014, Bishop CM, Nasrabadi NM. Pattern Recognition and Machine Learning. Journal of Electronic Imaging 2007; 16(4), and Scholkopf B, Smola AJ. Learning with kernels: support vector machines, regularization, optimization, and beyond. MIT press; 2001.

[0175] Feed - Forward Neural Network (FFNN)

[0176] FFNN is the first and most straightforward type of neural network, but due to their flexibility, they are still successfully applied to many different types of machine learning problems ranging from classification and regression tasks to data generation, unsupervised learning, etc. Similar to LRR, FFNN learns a direct mapping f between some input parameters x and some output values y. However, unlike linear models, FFNN can also approximate highly non-linear correlations between the input and output. This is achieved by transforming the input using a series of "layers", where each layer typically consists of a linear transformation followed by a non-linear operation σ:

[0177]

[0178] In some cases, FFNN may be difficult to train because the error function is highly non-convex, and compared to the global optimal solutions found by LRR and KRR, the optimization process usually only finds local minima. However, the loss in these local minima is usually similar to the global optimal value, so this property does not significantly affect the performance of a properly trained neural network. Additionally, due to the large number of parameters (W1,..., W l ) and high flexibility of FFNN, if not properly trained, it may overfit, especially when using a small training set. For a detailed discussion of the KRR model, please refer to the following publications: Draper NR, Smith H. Applied regression analysis, vol. 326. John Wiley & Sons; 2014, Bishop CM, Nasrabadi NM. Pattern Recognition and Machine Learning. Journal of Electronic Imaging 2007; 16(4), and Jaeger H. The “echo state” approach to analysing and training recurrent neural networks - with an erratum note. Bonn, Germany: German National Research Center for Information Technology GMD Technical Report 2001; 148(34):13.

[0179] 4.2 Stateful Model

[0180] Compared with stateless models, stateful models explicitly use only the input x(t), rather than past inputs x(t-1),..., x(t-k), to predict the output y(t) at a certain time point t. Instead, they maintain the hidden state h(t) of the system, which is continuously updated with each new time step and thus contains information about the entire past time series. The current input conditions and the hidden state of the model can then be used to predict the output:

[0181] Both of these stateful models belong to the class of Recurrent Neural Networks (RNNs). RNNs are a powerful method for time series modeling; however, they can be difficult to train because their depth increases with the length of the time series. If the training is not carefully executed, this can lead to gradient divergence during the error backpropagation training process, and if the optimization converges completely, this may result in very slow convergence ("vanishing gradient problem").

[0182] Echo State Network (ESN)

[0183] Figure 7 An exemplary structure of an ESN is shown. An ESN is an alternative RNN architecture that can alleviate some of the training-related problems of the above-mentioned RNNs by training without using error backpropagation at all. Instead, the ESN uses a very large randomly initialized weight matrix that, combined with the recursive mapping of past inputs, essentially acts as a random feature expansion of the input (similar to the implicit feature mapping φ used in KRR); collectively referred to as the "reservoir". In this way, the ESN can track the hidden state h(t) ∈ R of the system by updating h(t) at each time step to include a weighted sum of the previous hidden state h(t-1) and a combination of the randomly expanded input feature x(t) and the randomly recurrently mapped h(t-1) m where m >> d x ). Then the LRR is used to compute the final prediction of the output for the input and the hidden state, i.e.,

[0184] where

[0185] Typically, an Echo State Network is a very powerful type of RNN, and its performance in dynamic system prediction is usually comparable to or even better than that of other more popular and more complex RNN models (LSTM, GRU, etc.). Since the only parameter to be learned is the weight W of the linear model for the final prediction out , the ESN can also be trained on smaller datasets without much risk of overfitting.

[0186] LSTM Network

[0187] Another very popular architecture for dealing with the vanishing gradient problem in RNNs is the Long Short-Term Memory (LSTM) architecture, which was developed specifically for this purpose. Figure 8 An exemplary structure of an LSTM network is shown. LSTM is trained using error backpropagation as usual, but avoids the vanishing gradient problem by using an additional state vector called the "cell state" along with the usual hidden state. The cell state is the core component of the LSTM and runs through the entire recursive chain, while being slowly updated at each time step using only linear updates, enabling it to maintain long-term correlations in the data and keep stable gradients over long sequences. The inclusion or deletion of new information in the cell state is carefully regulated by special neural network layers called gates. Although the update of the hidden state h(t) of the LSTM network is much more complex compared to an ESN, the final prediction is still just a linear transformation of the internal hidden state of the network:

[0188] where

[0189] However, in this case, the parameter values of W o are optimized together with the other parameters of the LSTM network, rather than using a separate LRR model.

[0190] Since modeling the gates that regulate the cell state requires multiple layers, LSTM typically requires a large amount of training data to avoid overfitting. Despite its complexity, the stability of the LSTM gradients makes it very suitable for time series problems with long-term correlations.

[0191] 4.3 Variants of Stateful Model

[0192] Feedback State Model

[0193] So far, we have only used operating parameters to predict the key performance indicators (KPIs) of the process, but using past KPIs as inputs can serve as a powerful new source of information, especially because of the high autocorrelation of KPIs over time within the same cycle.

[0194] The main challenge here is that the KPIs from the previous time step are not easily obtainable. In fact, in real-world scenarios, we can expect that at most only a few KPI values are available at the start of a cycle, and we need to predict the KPIs for the remaining duration of the cycle. Since autocorrelation rapidly weakens over time, using only these KPI values at the start of the cycle is not very beneficial for any long-term prediction. However, assuming our prediction is accurate enough, we can use the predicted KPI from the previous time step as a reasonable approximation of the true KPI. This will enable us to utilize the high temporal autocorrelation between the outputs to improve our prediction accuracy.

[0195] One way to incorporate this into a stateful model is to include the predicted output (or true output, if available) from the previous time step into the input vector for the current time step. For example, Figure 9 shows an example of a feedback state model, showing the concatenation of the output from the previous time step and the input for the next time step. In Figure 9 the exemplary embodiment shown:

[0196]

[0197] However, such an implementation can easily lead to large prediction errors. The reason for this is that the predicted output is only an approximation of the true output and is thus less reliable than the true output. Since the previous predicted output will be used for the next prediction, any small error in the predicted output value will propagate into the prediction of the next output. Over a long period of time, these small errors accumulate and can cause the prediction to be very different from the true output time series, resulting in very large errors. Therefore, it is crucial to distinguish between the reliable true output and the unreliable predicted output of the network, so that the network can independently estimate the reliability of these two variables.

[0198] One way to achieve this is to include an indicator variable alongside each feedback output value, which will indicate whether that output value is a true output, i.e., an actual measured KPI from the process, or a predicted KPI, i.e., the output from the stateful model at the previous time step. Thus, the exemplary feedback state model can be implemented simply by appending two values to the input vector at each time step: the output value from the previous time step and the indicator variable, which is 0 if the feedback value is a truly measured KPI, or 1 if the feedback value is predicted by the stateful model at the previous step. Figure 6 A diagram of this model is given in , which shows an example of a feedback LSTM, as an example of a feedback state model, which shows the concatenation of the output from the previous time step and the input for the next time step. Preferably, the network will learn the connection between these two variables and thus learn to distinguish between reliable true feedback values and less reliable past LSTM predictions.

[0199] Hybrid Model

[0200] In the basic problem setting of predicting the Industrial Aging Process (IAP), all considered processes are affected by some underlying deterioration progress that reduces process efficiency over time. Since this deterioration is long-term and occurs throughout the cycle, it is difficult to predict because it is affected by the conditions early in the cycle, the correlation is largely unknown, and it is difficult to learn due to large time lags. However, since engineers are often aware of the basic dynamics behind the deterioration progress, some parametric prototype functions can be used to parameterize the deterioration of KPIs, and the parameters of this prototype function can be perfectly fitted to match the deterioration curve of a given cycle. We attempt to utilize this knowledge to simplify the learning problem of using LSTM as an example of a stateful model by separating the problem into predicting the instantaneous and long-term effects of the input on the KPI.

[0201] One way to isolate the instantaneous effect is to train a linear model without any time information. In our experiments, when the effect of the deterioration is still small and no time variable is used as an input, we train the LRR model as an example of a linear model only in the initial time period of the cycle (e.g., the first 1% - 10% of all observations in the cycle, preferably 1% - 5%), so the model does not attempt to learn from the time context but only attempts the instantaneous effect of the input on the KPI. Although this method will only learn the linear instantaneous effect, usually this is sufficient to remove most of the instantaneous artifacts from the cycle, making the residuals reflect the deterioration curve.

[0202] As mentioned before, the residuals can then be modeled using a parametric prototype function, and the parameters of this prototype function will be fitted to each deterioration curve. In this way, instead of predicting the individual values at each time point of the deterioration trend, which is usually highly non-stationary, only a set of parameters for each cycle are predicted using LSTM, and these parameters are used by the prototype function to model the entire deterioration curve. This in turn makes the learning problem more constrained because only functions in the form given by the prototype can be used to model the deterioration. We expect this feature to be particularly useful for real-world datasets, where the constraints enforced by the prototype function should reduce overfitting to smaller training sets.

[0203] As a final step, since LRR only captures the instantaneous linear correlation, while LSTM will ideally capture the long-term deterioration trend.

[0204] Dual-Speed Hybrid Model

[0205] In some cases, since the prototype function may not always perfectly adapt to degradation and there will still be some artifacts that are not linear or instantaneous and thus not captured by a linear model (such as LRR), we need another stateful model, such as LSTM, which will attempt to model these additional short-term artifacts separately at each time point. Due to this combination of two stateful models, one for long-term degradation and the other for short-term artifacts, we name this model the two-speed model, whose complete scheme is as Figure 10 shown above, which shows an overview of the two-speed hybrid model, showing three different model components (e.g., one LRR and two LSTMs) and the period decomposition they learn.

[0206] 5. Training Process

[0207] The data-driven model is parameterized according to a training data set, where the training data set is based on a historical data set including operating data, catalyst aging metrics, and at least one target operating parameter.

[0208] For example, for the ESN model, the parameters of the reservoir matrix are not trained but randomly generated, and training occurs after the hidden state features at each time point in the training data set are generated. After that, the final output matrix is parameterized / trained using linear ridge regression, resulting in a globally optimal linear mapping that minimizes the difference between the target and the prediction.

[0209] For the LSTM-based model, training is performed using stochastic gradient descent, where the model parameters are slowly updated using the gradients of a random subset of the training samples in order to minimize some error function (in this case, the difference between the prediction and the target). This process is repeated iteratively many times until the optimization converges to some (most likely) local minimum of the error function.

[0210] The machine learning model also has a set of hyperparameters that cannot be trained. To select a good set of hyperparameters, we use a validation set that is disjoint from the training set. The model is instantiated with different sets of hyperparameters and trained on the training set, and then the performance is measured on the validation set. Subsequently, for each model type, we select the hyperparameters that bring the optimal performance for that specific model on the validation set.

[0211] Finally, to evaluate the generalization performance of the model on new unknown samples, we use a test set that is different from the training set and the validation set.

[0212] The loss is calculated as the average of the root mean square error (RMSE) for all test periods. The predictions of the ESN and LSTM models are independent in different periods because the hidden state is reinitialized before the prediction for each new period.

[0213] 6. Results

[0214] In this section, we report our evaluation of the seven different machine learning models introduced in Section 3 using the synthetic and real-world datasets described in Section 2. To measure the prediction error of the machine learning models, we use the mean squared error (MSE), which we define somewhat differently than usual since our dataset is subdivided into periods: Let the dataset D consist of N periods, and let y i (t) denote the KPI at time point t ∈ 0,..., T i in the i-th period, where T i is the length of the i-th period. Then, given the corresponding model predictions the MSE of the model over the entire dataset is computed as

[0215]

[0216] Since the synthetic dataset and the real-world dataset are very different, they are used to examine different aspects of the models. The synthetic dataset is used to examine how the models perform in an almost ideal scenario, where data is freely available and the noise is very low or even non-existent. On the other hand, the real-world dataset is used to test the robustness of the models since it contains only a limited number of training samples and a relatively high noise level.

[0217] 6.1 Synthetic Data Set

[0218] To systematically evaluate the performance of different methods in a controlled environment, synthetic datasets were generated as described in Section 2. A total of 50 years of historical data were generated, consisting of 2153 periods with a total of 435917 time points. Approximately 10% of the periods in the dataset were randomly selected as the out-of-sample test set, resulting in a training set consisting of 1938 periods (391876 time points) and a test set consisting of 215 periods (44041 time points). Only the results for the transformation as the degradation KPI are discussed; the results for the selectivity are similar.

[0219] The hyperparameters of the LRR, KRR, and ESN models were selected using 10-fold cross-validation on the training set. The FFNN and LSTM models were trained using stochastic gradient descent with Nesterov momentum for parameter updates. The hyperparameters of the neural network models were determined based on the performance of the validation set, which consisted of 15% of the periods randomly selected from the training set. Early stopping was used to select the number of training epochs, and training was stopped if the validation set error did not improve in the last 6 epochs.

[0220] For stateless models such as LRR, KRR, and FFNN, the input vector at time point t consists of the operating parameters over the past 24 hours, providing the model with a past time window, i.e., x 24h (t) = [x(t); x(t - 1);...; x(t - 24)]. Further increasing this time window did not result in any significant improvement in the performance of any of the models. Since stateful models are able to encode the past in their hidden states, the input to ESN and LSTM at any time point t only contains the operating parameters at the current time point, i.e., x(t). Feedback state models, such as feedback LSTM, append two values to the input vector at each time step: the output value at the previous time step and the metric variable, which is 0 if the feedback value is a truly measured KPI or 1 if the feedback value was predicted by the LSTM at the previous step. The input to the hybrid model can be a combination of the inputs of the stateless and stateful models.

[0221] LRR, KRR, FFNN, ESN and LSTM

[0222] Figure 11Shows the mean squared error (MSE) of each of the five models on the training and test sets for different training set sizes. For most models, the error converges relatively early, meaning that even with only a small fraction of the complete dataset, as long as the corresponding model complexity allows, the model can manage to learn an accurate approximation of the synthetic dataset dynamics. This also indicates that the error present in the models is largely due to limitations in the flexibility of the models themselves, rather than due to the training set not being large enough. This is evident in LRR, which essentially uses 5% of the total dataset size to achieve its maximum performance. Since LRR is a linear model, it can only learn linear relationships between the input and output. Although this high bias prevents the model from learning most of the non - linear dynamics regardless of the training set size, it also means that the model has low variance, i.e., it does not tend to overfit the training data. For FFNN, the error decreases slowly as the number of samples increases, although at an increasingly slower rate, and the error using the complete training dataset is significantly lower than that of LRR. As for ESN and LSTM, judging from the difference between the training and test errors, both methods seem to be somewhat overfitting for smaller training set sizes. However, even so, the test errors are much lower compared to the three stateless models. The error convergence of the two models is around 50% of the entire dataset, after which there is little overfitting and there is also no significant improvement in performance for larger dataset sizes. The general lack of overfitting can be explained by the fact that the training and test sets are generated using exactly the same model, i.e., they are drawn from the same distribution, which is the optimal setting for any machine - learning problem. Additionally, the lack of noise in the synthetic dataset also helps to explain the lack of overfitting, as overfitting typically involves the model fitting the noise rather than the actual signal / pattern. Among all dataset sizes, the LSTM model clearly performs the best, with an error 5 times smaller than that of the ESN model when using the complete dataset.

[0223] Given the excellent performance of the ESN and especially the LSTM models, these experiments clearly show that it is in principle possible to predict the entire degradation cycle with very high accuracy even with a small amount of high - quality data.

[0224] Figures 12A - 12D Plots showing the true and predicted transition rates of different models for some randomly selected cycles from the training and test sets. These show that all models are able to accurately predict the instantaneous effect of the input parameters on the output, as the relationship is largely linear rather than time - dependent. However, the greatest differences between the models lie in the non - linear long - term degradation, where the stateless models only predict a roughly linear trend, FFNN is slightly closer to the actual degradation trend due to its non - linearity, while the ESN model better predicts the degradation but fails to capture the rapid decline at the end of each cycle. On the other hand, the LSTM model almost perfectly captures both the short - term and long - term effects, with only a small error at the end of cycles with less data due to the different lengths of the cycles.

[0225] Feedback State Model

[0226] The test scenario for the feedback model is that the output values for the first 12 hours of each period are known and can thus be used as true feedback. After that, the feedback must be obtained from the predictions of the LSTM at the previous time step. Therefore, all mean squared errors reported by the feedback model are obtained by evaluating the test set, where the first 12 hours of each period are given as true feedback.

[0227] Figures 13A - 13B A graph showing the predicted and true KPIs of an example of a feedback state model for training and test samples randomly selected from two data sets, namely the synthetic data set and the real-world data set in Factory C.

[0228] For the synthetic data set, after using a stage-by-stage training process, the error of the feedback model is significantly higher than that of the conventional LSTM. More precisely, the conventional LSTM has an MSE of 0.08, while the MSE of the feedback model is nearly 4 times larger, at 0.31 (0.32 training error).

[0229] Although it is not currently clear why the performance is affected in this case, our hypothesis is that the overall high accuracy of the predictions leads the network to learn that the feedback values are also reliable during prediction, causing the model to start strongly relying on the predicted feedback values for future predictions. As mentioned earlier, this leads to the accumulation of small errors in the feedback values, which can be the reason for the degraded performance of the feedback LSTM compared to the conventional LSTM.

[0230] In addition, these two additional input parameters can make the learning problem more complex, so they do not allow the feedback LSTM to converge quickly to a very low local minimum, which is actually useful in this case because it reduces overfitting, which can again lead to better performance on the test set.

[0231] Hybrid Model

[0232] For both the synthetic data set and the real data set, we used an exponential function of the following form

[0233] f deg (t) = g(p1(t), p2(t),..., p n (t))

[0234] where the parameter p1(t) is predicted by an LSTM, as an example of a stateful model, because of short-term artifacts and the parameters p2(t),..., p n (t) are predicted by a long-term LSTM.

[0235] Figures 14A - 14BA graph showing the predicted and true KPIs of a hybrid model example, for randomly selecting training and test samples from two datasets, namely the synthetic dataset and the real-world dataset in Factory C.

[0236] For the synthetic dataset, the MSE of the two-speed model is slightly higher than the error of the regular LSTM. The two-speed LSTM has an MSE of 0.13 (training error of 0.137), while the MSE of the LSTM is 0.08. This is somewhat expected, as the constraints caused by the prototype function make the LSTM slightly less flexible, which is disadvantageous for the synthetic dataset with a large amount of data. Therefore, overfitting is not a problem.

[0237] 6.2 Real - World Data Set

[0238] The real-world dataset is much smaller than the synthetic one, containing a total of 375 cycles. After removing some abnormal cycles (less than 50 hours), the final size of the dataset is 327 cycles of 36058 time points in total, that is, it is more than 10 times smaller than the complete synthetic dataset. Since the real-world dataset uses different catalyst loadings in the reactor over 3 time periods, we test the performance in a realistic way by selecting the third catalyst loading as the test set, which allows us to see to what extent the model can infer the different conditions caused by catalyst exchange. This results in a training set consisting of 256 cycles (28503 time points), while the test set consists of 71 cycles (7555 time points).

[0239] The hyperparameters of the real-world dataset are selected in a similar way to the synthetic dataset, except that due to the smaller size of the dataset and thus shorter epochs, early stopping is triggered when the validation error does not improve in the past 30 epochs.

[0240] For this dataset, the input of both the stateful and stateless models at time point t only contains the process conditions of that time point x(t). Extending the time window to the past few hours only reduces the performance, as it reduces the size of the training set (if we take the past k hours, the input for each cycle must start k hours later, resulting in a loss of k samples per cycle) and increases the number of input features, making it easier for all models to overfit.

[0241] LRR, KRR, FFNN, ESN, LSTM

[0242] Figure 15Shows the mean squared error of each of the five models on the training and test sets. Due to the large noise and small amount of data, the results here are different compared to the results of the synthetic dataset: the more complex models show overfitting as the test error is significantly greater than the corresponding training error, especially for KRR, which also has the largest test error among all models. On the other hand, LRR shows little overfitting and its performance on the test set is closer to the performance of the other models. Again, the performance of ESN and LSTM is better than that of the stateless models, but this time, the gap is much smaller and both models show very similar performance. Considering the greater noise level and fewer sample numbers, this may be due to the greater likelihood of overfitting of the LSTM model here.

[0243] Feedback State Model

[0244] For the real-world dataset, namely Figure 10 (c) and the Factory C dataset in 9(d), the error of the feedback model is 25.58 (18.67 training error), which is significantly lower than that of the conventional LSTM with an MSE of 33.35. Compared with the synthetic dataset, we obtained the opposite result, which we hypothesized to be the case because the prediction of the Factory C data has a higher noise level and overall worse accuracy. This means that the correlation between the previous prediction output and the next true output is not that high, so the feedback model does not rely too much on these values when predicting the next output, but due to the step-by-step training process, it still learns to rely on the true feedback, which may be the reason for the performance improvement.

[0245] Hybrid Model

[0246] For the real-world dataset, namely Figure 10 (c) and the Factory C dataset in 9(d), i.e., the real-world dataset, the MSE of the two-speed model is 21.9 (26.22 training error), i.e., significantly lower than the MSE of the conventional LSTM, which is 33.35. As expected, the reason is the constraint of the prototype function, which fits well the shape of the degradation progress and reduces overfitting, which is particularly useful because the Factory C dataset has a much smaller available training set.

[0247] 7. Use Case

[0248] Well-known degradation phenomena in chemical plants can be predicted by the above methods, including but not limited to:

[0249] ■ Deactivation of heterogeneous catalysts due to coking, sintering, or poisoning;

[0250] ■ Blockage of process equipment (such as heat exchangers or pipes) on the process side due to coke layer formation or polymerization;

[0251] ■ Fouling of heat exchangers on the water side due to microorganisms or crystalline deposits;

[0252] ■ Corrosion of equipment (such as nozzles or pipes) installed in fluidized bed reactors.

[0253] 8. Summary

[0254] Developing an accurate mathematical model of the industrial aging process (IAP) is crucial for predicting when critical assets need to be replaced or restored. In world-scale chemical plants, such predictions have great economic value as they improve plant reliability and efficiency. While mechanical models are useful for elucidating the factors influencing the progression of degradation under laboratory conditions, it is well known that it is very difficult to adapt them to the specific environment of individual plants. On the other hand, data-driven machine learning methods can learn models and make predictions based on the historical data of a specific plant, and thus can easily adapt to a variety of conditions as long as sufficient data is available. Although simpler, especially linear prediction models, have been studied in the context of predictive maintenance before, more recent and complex machine learning models (such as recurrent neural networks) have not been examined in detail so far.

[0255] In this disclosure, we address the task of predicting KPIs that indicate the slow degradation of critical equipment over the time span of an entire degradation cycle, based only on initial process conditions and predictions of how the process will operate during that period. To this end, we compared a total of seven different prediction models: three stateless models, namely linear ridge regression (LRR), non-linear kernel ridge regression (KRR), and feed-forward neural network (FFNN), two state models based on recurrent neural networks (RNN), echo state network (ESN) and LSTM, and variants of the state model, namely the feedback state model and the hybrid model. To evaluate the importance of the amount of available historical data for model predictions, we first tested them on a synthetic dataset that essentially contains an infinite number of noise-free data points. In a second step, we examined how these results translate to real data from a large chemical plant of BASF.

[0256] While the stateless models (LRR, KRR, and FFNN) accurately capture the instantaneous changes in KPIs caused by changing process conditions, they may not accurately capture the underlying trends caused by slower degradation effects. On the other hand, ESN and LSTM are able to additionally correctly predict long-term changes, however at the cost of requiring large amounts of training data to do so. Due to more parameters to be adjusted, non-linear models usually overfit to the specific patterns observed in the training data and thus make relatively more mistakes on new test samples. Additionally, two main variants of the basic LSTM model are expected to improve the performance on real-world datasets: including feedback loops which take past prediction outputs as additional inputs and splitting the model into two or more different models which predict different aspects of the output dynamics (e.g., instantaneous effects vs. long-term trends).

[0257] In general, all models can produce very promising predictions that are accurate enough to improve the scheduling decisions for maintenance events in a production plant. The choice of the optimal model in a particular case depends on the amount of available data. For very large datasets, we found that LSTM can produce almost perfect predictions over long time horizons. However, if only a few cycles are available for training or the data is very noisy, applying a hybrid model may be advantageous, which can significantly improve the performance of the LSTM model by reducing overfitting especially on small datasets.

[0258] While accurate prediction of IAP will improve the production process by allowing longer planning horizons and ensuring the plant operates economically and reliably, the ultimate goal is of course to better understand and subsequently minimize the degradation effects themselves. While mechanical and linear models are fairly straightforward to interpret, neural network models have long been shunned due to their opaque predictions. However, due to novel interpretation techniques such as layer-wise relevance propagation (LRP), this is changing, which makes it possible to visualize the contribution of individual input dimensions to the final prediction. By such methods, the predictions of RNNs such as LSTM can be made more transparent, thus revealing the influencing factors and production conditions leading to the aging process under study, which can then be used to help improve the underlying process engineering.

[0259] Figure 16 Apparatus 200 for predicting the degradation progress of a chemical production plant is schematically shown. Apparatus 200 includes an input unit 210 and a processing unit 220.

[0260] The input unit 210 is configured to receive currently measured process data indicative of current process conditions of a current operation of at least one chemical process equipment for a chemical production plant. The at least one chemical process equipment operates in a cyclic manner including multiple runs. The at least one chemical process equipment has one or more degradation key performance indicators (KPIs) for quantifying the progress of degradation of the at least one chemical process equipment. The input unit 210 is further configured to receive one or more expected operating parameters indicative of planned process conditions of the at least one chemical process equipment within a prediction range.

[0261] Thus, in an example, the input unit 210 can be implemented as an Ethernet interface, a USB(TM) interface, a wireless interface such as WiFi(TM) or Bluetooth(TM), or any comparable data transfer interface capable of data transfer between an input peripheral device and the processing unit 220.

[0262] The processing unit 220 is configured to perform any of the above method steps.

[0263] Thus, the processing unit 220 can execute computer program instructions to perform various processes and methods. The processing unit 220 can refer to, be part of, or include the following: an application specific integrated circuit (ASIC), an electronic circuit, a processor (shared, dedicated, or group), and / or a memory (shared, dedicated, or group) that executes one or more software or firmware programs, combinational logic circuits, and / or other suitable components providing the functions. Additionally, such a processing unit 220 can be connected to a volatile or non-volatile memory, a display interface, a communication interface, etc. known to those skilled in the art.

[0264] The apparatus 200 includes an output unit 230 for providing future values of one or more degradation KPIs within a prediction range, which can be used for monitoring and / or control.

[0265] Thus, in an example, the output unit 230 can be implemented as an Ethernet interface, a USB(TM) interface, a wireless interface such as WiFi(TM) or Bluetooth(TM), or any comparable data transfer interface capable of data transfer between an output peripheral device and the processing unit 230.

[0266] Figure 17An example of a system 300 for predicting the deterioration progress of a chemical production plant is schematically shown. The system 300 of the illustrated example includes a sensor system 310, which includes one or more sensors (not shown) installed in the chemical production plant, a data repository 320, a web server 330 including means 200 for predicting the deterioration progress of the chemical production plant as described above and below, a plurality of electronic communication devices 340a, 340b, and a network 350.

[0267] The sensor system 310 may include one or more sensors installed in the chemical production plant (e.g., installed in one or more chemical process equipments) for sensing temperature, pressure, flow rate, etc. Examples of sensors may include, but are not limited to, temperature sensors, pressure sensors, flow sensors, etc.

[0268] The data repository 320 may be a database that receives data generated by one or more sensors of the sensor system 310 in a production environment and operation parameters indicating process conditions. For example, the data repository 320 may collect sensor data and operation parameters from different chemical process equipments or from different chemical production plants. These chemical production plants may be located in the same physical location or in different cities, states, and / or countries, and they are interconnected through a network. In another example, the data repository may collect sensor data and operation parameters from different production sites, which are either in the same physical location or scattered in different physical locations. The data repository 320 of the illustrated example may be any type of database, including a server, a database, a file, etc.

[0269] The network server 330 of the illustrated example can be a server that provides network services to facilitate the management of sensor data and operating parameters in multiple data repositories. The network server 330 can include means 200 for predicting the deterioration progress of a chemical production plant as described above and below. In some embodiments, the network server 330 can interact with a user, for example, via a web page, a desktop application, a mobile application, to facilitate the management of sensor data, operating parameters, and the use of the means to predict the deterioration progress of a chemical production plant. Alternatively, the network server 330 of the illustrated example can be replaced by another device (such as another electronic communication device) that provides any type of interface (such as a command-line interface, a graphical user interface). These interfaces (such as web pages, desktop applications, mobile applications) can allow users to manage data via the network 350 using the electronic communication devices 340a, 340b. The network server 330 can also include an interface through which a user can authenticate (by providing a username and password). For example, a user account can be used to authenticate a system user of a specific chemical production plant to access some data repositories using the network server 330 to obtain the sensor data and operating parameters of the specific chemical plant to allow the means 200 to predict the deterioration progress of the specific chemical plant.

[0270] The illustrated example of the electronic communication devices 340a, 340b can be a desktop computer, a notebook computer, a laptop computer, a mobile phone, a smart phone, and / or a PDA. In some embodiments, the electronic communication devices 340a, 340b can also be referred to as clients. Each of the electronic communication devices 340a, 340b can include a user interface that is configured to facilitate one or more users to submit access to the network server. The user interface 12 can be an interactive interface, including but not limited to a GUI, a role-based user interface, and a touch-screen interface. Optionally, the illustrated example of the electronic communication devices 340a, 340b can include a memory for storing, for example, sensor data and operating parameters.

[0271] The illustrated example of the network 350 communicatively couples the sensor system 310, the data repository 320, the network server 330, and the multiple electronic communication devices 340a, 340b. In some embodiments, the network can be the Internet. Alternatively, the network 350 can be any other type and number of networks. For example, the network 350 can be implemented by several local area networks connected to a wide area network. Of course, any other configuration and topology can be used to implement the network 350, including any combination of wired networks, wireless networks, wide area networks, local area networks, etc.

[0272] This exemplary embodiment of the invention covers computer programs that use the invention from the beginning and computer programs that turn existing programs into programs using the invention by means of updates.

[0273] Furthermore, a computer program unit may be able to provide all the necessary steps to carry out the process of an exemplary embodiment of the above method.

[0274] According to another exemplary embodiment of the present invention, there is provided a computer-readable medium, such as a CD-ROM, on which a computer program unit is stored, the computer program unit being described in the foregoing part.

[0275] The computer program may be stored and / or distributed on a suitable medium, such as an optical storage medium or a solid-state medium provided together with or as part of other hardware, but may also be distributed in other forms, such as via the Internet or other wired or wireless telecommunication systems.

[0276] However, the computer program may also be presented via a network such as the World Wide Web and downloaded into the working memory of a data processor from such a network. According to another exemplary embodiment of the present invention, there is provided a medium for making a computer program unit available for download, the computer program unit being arranged to execute a method according to one of the foregoing embodiments of the present invention.

[0277] According to an embodiment of the present invention, the present application further provides the following embodiments:

[0278] Embodiment 1: A computer-implemented method for predicting the deterioration progress of a chemical production plant, comprising:

[0279] a) receiving, via an input channel, currently measured process data, the process data indicating current process conditions of a current operation of at least one chemical process equipment for a chemical production plant, wherein the at least one chemical process equipment operates in a cyclic manner including multiple runs, wherein each run includes a production phase followed by a regeneration phase, and wherein the at least one chemical process equipment has one or more deterioration key performance indicators (KPIs) for quantifying the deterioration progress of the at least one chemical process equipment;

[0280] b) receiving, via the input channel, one or more expected operating parameters, the expected operating parameters indicating planned operating conditions of at least one chemical process equipment within a prediction range;

[0281] c) applying, by a processor, a data-driven model to an input data set including the currently measured process data and the one or more expected operating parameters to estimate future values of one or more deterioration KPIs within the prediction range, wherein the data-driven model is parameterized or trained according to a training data set, and wherein the training data set is based on a historical data set including process data and one or more deterioration KPIs; and

[0282] d) Provide future values of one or more degradation KPIs within the prediction range via an output channel, and the future values can be used for monitoring and / or control.

[0283] Example 2: The method according to Example 1,

[0284] wherein the one or more degradation KPIs are selected from parameters including the following:

[0285] - Parameters included in a set of measured process data; and / or

[0286] - Derived parameters that represent functions of one or more parameters included in a set of measured process data.

[0287] Example 3: The method according to Example 2, wherein the selected parameters have at least one of the following characteristics:

[0288] - Tend towards higher or lower values in a substantially monotonic manner on a time scale longer than the typical production time scale, thereby indicating the occurrence of an irreversible degradation phenomenon; and

[0289] - Return to the baseline after the regeneration stage.

[0290] Example 4: The method according to any one of the foregoing examples, wherein the degradation includes at least one of the following:

[0291] - Deactivation of a multiphase catalyst due to coking, sintering, and / or poisoning;

[0292] - Blockage of chemical process equipment on the process side due to coke layer formation and / or polymerization;

[0293] - Fouling of a heat exchanger on the water side due to microorganisms and / or crystalline deposits; and

[0294] - Corrosion of equipment installed in a fluidized bed reactor.

[0295] Example 5: The method according to any one of the foregoing examples, wherein the data-driven model includes:

[0296] - A stateful model, which is a machine learning model with a hidden state, and the hidden state is continuously updated with new time steps and contains information about the entire past time series; and / or

[0297] - A stateless model, which is a machine learning model, and the prediction of the machine learning model is only based on the input within a fixed time window before the current operation.

[0298] Example 6: The method according to Example 5, wherein the stateful model includes a recurrent neural network RNN.

[0299] Example 7: The method according to Example 6, wherein the RNN comprises at least one of the following:

[0300] - Echo State Network (ESN); and

[0301] - Long Short-Term Memory (LSTM) network.

[0302] Example 8: The method according to any one of Examples 5 to 7,

[0303] wherein the stateful model comprises a feedback state model, and the feedback state model comprises information regarding a predicted output or a true output in an input data set from a previous time step to a current time step;

[0304] wherein the predicted output is one or more predicted KPIs at a previous time step; and

[0305] wherein the true output is one or more measured KPIs at a previous time step.

[0306] Example 9: The method according to Example 8,

[0307] wherein the input data set further comprises an indicator variable indicating whether an output of a data-driven model from a previous time step is a predicted output or a true output.

[0308] Example 10: The method according to any one of Examples 5 to 9,

[0309] wherein step a) further comprises receiving previously measured process data, and the previously measured process data indicates past process conditions of at least one chemical process equipment in a chemical production plant during a predetermined period before the current operation;

[0310] wherein step b) further comprises receiving one or more past operation parameters, and the past operation parameters indicate past process conditions of at least one chemical process equipment during a predefined period before the current operation; and

[0311] wherein in step c), the input data set further comprises the previously measured process data and the one or more past operation parameters.

[0312] Example 11: The method according to Example 5, wherein the stateless model comprises at least one of the following:

[0313] - Linear Ridge Regression (LRR);

[0314] - Kernel Ridge Regression (KRR); and

[0315] - Feed-Forward Neural Network (FFNN).

[0316] Example 12: A method according to any one of Examples 5 to 11,

[0317] wherein the data-driven model is a hybrid model, and the hybrid model includes a stateful model for predicting the degradation trend of one or more degradation KPIs and a stateless model for predicting the additional instantaneous impact of operating parameters on one or more degradation KPIs;

[0318] wherein the degradation trend represents a monotonic change in the performance of the chemical process equipment on a time scale longer than the typical production time scale; and

[0319] wherein the additional instantaneous impact of the operating parameters does not include the time delay of the effect of the model input on one or more degradation KPIs.

[0320] Example 13: A method according to Example 12, wherein the stateful model includes a combination of mechanical pre-information about the process represented by a function with a predefined structure and a stateful model for estimating the parameters of the function.

[0321] Example 14: A method according to Example 12 or 13, wherein the stateless model includes a linear model.

[0322] Example 15: A method according to any one of the foregoing examples, wherein the input data set further includes at least one transformed process data, and the at least one transformed process data represents a function of one or more parameters of the currently measured process data and / or the previously measured process data.

[0323] Example 16: A device for predicting the degradation progress of a chemical production plant, comprising:

[0324] - an input unit;

[0325] - a processing unit; and

[0326] - an output unit;

[0327] wherein the input unit is configured to:

[0328] - receive the currently measured process data, the currently measured process data indicating the current process conditions of the current operation of at least one chemical process equipment of the chemical production plant, wherein the at least one chemical process equipment operates in a cyclic manner including multiple runs, wherein each run includes a production stage followed by a regeneration stage, and wherein the at least one chemical process equipment has one or more degradation key performance indicators KPIs for quantifying the degradation progress of the at least one chemical process equipment;

[0329] - receive one or more expected operating parameters, the expected operating parameters indicating the planned process conditions of at least one chemical process equipment within the prediction range;

[0330] wherein, the processing unit is configured to perform the method steps according to any one of claims 1 to 15; and

[0331] wherein, the output unit is configured to provide future values of one or more degradation KPIs within a prediction range, and the future values can be used for monitoring and / or control.

[0332] Embodiment 17: A computer program unit for indicating the device according to Embodiment 16, which is adapted to perform the method steps according to any one of Embodiments 1 to 15 when executed by the processing unit.

[0333] Embodiment 18: A computer-readable medium storing the program unit of Embodiment 17.

[0334] It must be noted that the embodiments of the present invention are described with reference to different subjects. In particular, some embodiments are described with reference to method-type claims, while other embodiments are described with reference to device-type claims. However, those skilled in the art will learn from the above and the following descriptions that, unless otherwise stated, any combination of features related to different subjects is also considered to be disclosed together with this application in addition to any combination of features belonging to one subject. However, all features can be combined to provide a synergistic effect that is more than just a simple sum of the features.

[0335] Although the present invention has been described in detail in the drawings and the foregoing description, such description and description should be considered illustrative or exemplary rather than restrictive. The present invention is not limited to the disclosed embodiments. By studying the drawings, the disclosure and the dependent claims, those skilled in the art can understand and implement other variations of the disclosed embodiments when practicing the claimed invention. In the claims, the word "comprising" does not exclude other elements or steps, and the indefinite article "a" or "an" does not exclude a plurality. A single processor or other unit can implement the functions recited in several items of the claims. The fact that certain measures are recited in mutually different dependent claims does not indicate that a combination of these measures cannot be used advantageously. Any reference signs in the claims should not be construed as limiting the scope.

Claims

1. A computer-implemented method (100) for predicting the degradation progress of a chemical production plant, comprising: a) receiving (110) currently measured process data via an input channel, the process data indicating current process conditions of a current operation of at least one chemical process equipment for the chemical production plant, wherein the at least one chemical process equipment has one or more degradation key performance indicators (KPIs) for quantifying the degradation progress of the at least one chemical process equipment; b) receiving (120) one or more expected operating parameters via the input channel, the expected operating parameters indicating planned operating conditions of the at least one chemical process equipment within a prediction range; c) applying (130), by a processor, a data-driven model to an input data set comprising the currently measured process data and the one or more expected operating parameters to estimate future values of the one or more degradation KPIs within the prediction range, wherein the data-driven model is parameterized or trained according to a training data set, and wherein the training data set is based on a historical data set comprising process data and the one or more degradation KPIs; and d) providing (140), via an output channel, the future values of the one or more degradation KPIs within the prediction range, the future values being usable for monitoring and / or control.

2. The method according to claim 1, Among them, wherein the at least one chemical process equipment operates in a cyclic manner comprising multiple runs, wherein each run comprises a production phase followed by a regeneration phase; and wherein the input data set comprises at least one process information from the last run.

3. The method according to claim 1 or 2, Among them, wherein the one or more degradation KPIs are selected from parameters comprising: - parameters included in a set of measured process data; and / or - derived parameters that represent a function of one or more parameters included in a set of measured process data.

4. The method according to claim 3, Among them, wherein the selected parameters have at least one of the following characteristics: - tending to higher or lower values in a substantially monotonic manner on a time scale longer than a typical production time scale, thereby indicating the occurrence of an irreversible degradation phenomenon; and - returning to a baseline after a regeneration phase.

5. The method according to claim 1, Among them, wherein the degradation comprises at least one of the following: - deactivation of a multiphase catalyst due to coking, sintering, and / or poisoning; - blockage of a chemical process equipment on the process side due to coke layer formation and / or polymerization; - fouling of a heat exchanger on the water side due to microorganisms and / or crystalline deposits; and - corrosion of equipment installed in a fluidized bed reactor.

6. The method according to claim 1, Among them, wherein the data-driven model comprises: - a stateful model, which is a machine learning model having a hidden state, the hidden state being continuously updated with new time steps and containing information about the entire past time series; and / or - A stateless model, which is a machine learning model, and the prediction of the machine learning model is only based on the input within a fixed time window before the current operation.

7. The method according to claim 6, Among them, The stateful model includes a Recurrent Neural Network (RNN).

8. The method according to claim 7, Among them, The RNN includes at least one of the following: - An Echo State Network (ESN); and - A Long Short-Term Memory (LSTM) network.

9. The method according to any one of claims 6 to 8, Among them, The stateful model includes a feedback state model, and the feedback state model includes information about the predicted output or the true output in the input data set for the current time step from the previous time step; wherein, the predicted output is one or more predicted KPIs at the previous time step; and wherein, the true output is one or more measured KPIs at the previous time step.

10. The method according to claim 9, Among them, The input data set further includes an indicator variable, and the indicator variable indicates whether the output of the data-driven model from the previous time step is a predicted output or a true output.

11. The method according to any one of claims 6 to 8, Among them, Step a) further includes receiving previously measured process data, and the previously measured process data indicates the past process conditions of the past operations of the at least one chemical process equipment in the chemical production plant within a predetermined period before the current operation; wherein, step b) further includes receiving one or more past operation parameters, and the past operation parameters indicate the past process conditions of the at least one chemical process equipment within the predefined period before the current operation; and wherein, in step c), the input data set further includes the previously measured process data and the one or more past operation parameters.

12. The method according to claim 6, Among them, The stateless model includes at least one of the following: - Linear Ridge Regression (LRR); - Kernel Ridge Regression (KRR); and - Feed-Forward Neural Network (FFNN).

13. The method according to any one of claims 6 to 8, Among them, The data-driven model is a hybrid model, and the hybrid model includes a stateful model for predicting the degradation trend of the one or more degradation KPIs and a stateless model for predicting the additional instantaneous impact of operation parameters on the one or more degradation KPIs; wherein, the degradation trend represents a monotonic change in the performance of the chemical process equipment on a time scale longer than the typical production time scale; and wherein, the additional instantaneous impact of the operation parameters does not include the time delay of the effect of the model input on the one or more degradation KPIs.

14. The method according to claim 13, Among them, The stateful model includes a combination of mechanical pre-information about the process represented by a function with a predefined structure and a stateful model for estimating the parameters of the function.

15. The method according to claim 14, Among them, The stateless model includes a linear model.

16. The method according to claim 1, Among them, The input data set further includes at least one transformed process data, which represents a function of one or more parameters of the currently measured process data and / or the previously measured process data.

17. An apparatus (200) for predicting the deterioration progress of a chemical production plant, comprising: - an input unit (210); - a processing unit (220); and - an output unit (230); wherein the input unit is configured to: - receive currently measured process data, which indicates current process conditions of a current operation of at least one chemical process equipment of the chemical production plant, wherein the at least one chemical process equipment operates in a cyclic manner including multiple runs, and wherein each run includes a production stage followed by a regeneration stage, and wherein the at least one chemical process equipment has one or more deterioration key performance indicators KPIs for quantifying the deterioration progress of the at least one chemical process equipment; - receive one or more expected operation parameters, which indicate planned process conditions of the at least one chemical process equipment within a prediction range; wherein the processing unit is configured to perform the method steps according to any one of claims 1 to 16; and wherein the output unit is configured to provide future values of the one or more deterioration KPIs within the prediction range, and the future values can be used for monitoring and / or control.

18. A computer program unit for indicating the apparatus according to claim 17, which is adapted to perform the method steps according to any one of claims 1 to 16 when executed by a processing unit.

19. A computer-readable medium storing the program unit of claim 18.

Citation Information

Patent Citations

  • Computer System And Method For Building And Deploying Models Predicting Plant Asset Failure

    US20190188584A1