A method of controlling a chemical process

By using lower dimensional spaces to filter outliers and combining mechanistic and statistical models, the method addresses the challenges of controlling chemical processes with unpredictable inputs, enhancing accuracy and optimizing production outcomes.

GB2631340BActive Publication Date: 2026-04-20JOHNSON MATTHEY DAVY TECHNOLOGIES LTD
View PDF 4 Cites 0 Cited by

Patent Information

Authority / Receiving Office
GB · GB
Patent Type
Patents
Current Assignee / Owner
JOHNSON MATTHEY DAVY TECHNOLOGIES LTD
Filing Date
2024-05-15
Publication Date
2026-04-20

AI Technical Summary

Technical Problem

Existing control systems for chemical processes, particularly those with unpredictable inputs like renewable energy and variable feedstock, struggle with non-linear dynamics and data noise, leading to inaccurate control of variables such as heat input and flow rates.

Method used

A method involving data preprocessing through lower dimensional spaces to identify and discard outliers, combined with a mechanistic and statistical model training approach, enhances the detection of faulty data and optimizes control variables for chemical processes.

Benefits of technology

This method improves the accuracy of controlling chemical processes by effectively filtering out noise and anomalies, allowing for better management of variable inputs and non-linear dynamics, thus optimizing production outcomes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000001_0000
    Figure 00000001_0000
  • Figure 00000001_0001
    Figure 00000001_0001
  • Figure 00000002_0000
    Figure 00000002_0000
Patent Text Reader

Abstract

Control of a chemical process is performed by obtaining a set of parameters of the chemical process using sensors and then operating the chemical plant with control variables based on the sets of para
Need to check novelty before this filing date? Find Prior Art

Description

The present invention relates to the control of chemical processes by computational means. Particularly, but not exclusively, the methods that follow have particular application to processes with unpredictable inputs in the form of feedstock flow rates and compositions or sources of power. A preferred example in which the disclosed methods may be employed is the process of producing emethanol. The process of producing e-methanol involves providing a feedstock of a renewable source such as biomass or biogas to a chemical plant. The process is typically powered, at least in part, by a renewable source of energy (e.g., wind power). As such, the supply of feedstock and electricity is significantly variable and unpredictable, since it depends on the composition of the feedstock and availability of power. Although this problem is present in numerous types of chemical plants, it is particularly exacerbated in e-methanol production, which involves many non-linear processes are present because of, for example, the presence of recycle loops. In scenarios such as this, the processing of the feedstock in a chemical plant provides a more difficult control problem than previously considered. Prior art systems of control use mechanistic models to identify set points for control variables (such as heat input, pump pressure, etc.) based on sensed sets of parameters (such as temperature, pH, flowrate, etc.). That is, mechanistic models have been formed from systems of deterministic physical equations, derived to represent the connection between the parameters and variables of the process. Such models must be carefully tuned by process engineers using chemical plant data. An example of such a prior art method is described in “Optimal Control of Methanol Synthesis Fixed-Bed Reactor” by Flavio Manenti and Giulia Bozzano, Industrial &Engineering Chemistry Research 2013 52 (36), 13079-13091. There is therefore a need for a better approach to the problem of controlling chemical processes by computational means. According to the invention, there is provided a method of controlling a chemical process, comprising: providing a feedstock to a chemical plant; processing the feedstock in the chemical plant to implement a chemical process to produce a product; repeatedly obtaining a set of parameters of the chemical process using sensors; and operating the chemical plant in accordance with control variables based on the sets of parameters by: discarding outliers from the sets of parameters; and simulating the chemical process using the retained sets of parameters to estimate a set of control variables for the chemical plant that optimises an objective function, wherein the step of discarding outliers from the sets of parameters comprises: projecting the sets of parameters into a first lower dimensional space to provide projected parameters; identifying outliers using the projections of the parameters in the first 07 05 25 lower dimensional space; and discarding at least some of the sets of parameters corresponding to the outliers, wherein discarding at least some of the sets of parameters corresponding to the outliers comprises: projecting the sets of parameters into a second lower dimensional space to provide projected parameters, wherein the second lower dimensional space is different from the first lower dimensional space; identifying outliers from the projected parameters in the second lower dimensional space; and discarding the sets of parameters corresponding to the outliers in both the first and the second lower dimensional spaces. As will be explained in more detail below, the use of a lower dimensional space for the pre-processing 5 of the sets of parameters can enable a more accurate detection of faulty data. Faulty data may result, for example, from an error in a sensed parameter. This is particularly effective in the context of a highly variable, non-linear, chemical process, in which traditional methods of pre-processing (such as low pass filtering) are unable to respond correctly to the variation in the data. In particular, the method proposed can avoid unnecessary removal of data resulting from transitions between operating modes 10 of a chemical plant, while data from such transitions would potentially be considered an outlier by conventional approaches. The method of controlling a chemical process, may comprise: providing a feedstock to a chemical plant; processing the feedstock in the chemical plant to implement a chemical process to produce a 15 product; repeatedly obtaining a set of parameters of the chemical process using sensors; and operating the chemical plant in accordance with control variables based on the sets of parameters using a simulation of the chemical process to estimate a set of control variables for the chemical plant that optimises an objective function, wherein the simulation is obtained by: providing a mechanistic model of the chemical process; generating first training data using the mechanistic model; obtaining 20 second training data by sensing a set of parameters of a test process implemented by a chemical plant; and training a statistical model of the chemical process using both the first and the second training data to provide a trained statistical model. As will be explained in more detail below, the use of a mechanistic model of the chemical process can 25 generate training data that provide a broad representation of the state space in which a model of the process is to be applied. This is important when the real world data available for a chemical process is limited to data surrounding the preferred operating points of a chemical plant. Such real world data is more accurate than the generated data, but only over the limited range related to the operating modes of the chemical plant, which do not enable a model to extrapolate to novel scenarios. Additionally, 30 real-world data suffers from process and instrument noise, which can negatively impact model predictions. The combination of the two data sets enables a model to have broad applicability, without loss of the accuracy obtainable using the real world data. For a better understanding of the invention, and to show how the same may be put into effect, reference will now be made, by way of example only, to the accompanying drawings in which: Figure 1 shows a schematic representation of a chemical plant for producing e-methanol; 5 Figure 2 shows a flow chart of a method of controlling a chemical plant such as that shown in Figure 1; 07 05 25 Figure 3A shows a flow chart of a method of training a model for use in the method of Figure 2; Figure 3B shows a two-dimensional representation of training data for use in the method of Figure 3A; Figure 4 shows a flow chart of a method of detecting faulty data for use in the method of Figure 2; Figure 5 shows a flow chart of another method of detecting faulty data for use in the method of Figure 2; and Figure 6 shows a flow chart of a further method of detecting faulty data for use in the method of Figure 2. Figure 1 shows a chemical plant 100. An exemplary chemical plant 100 comprises: chemical processing apparatus 101; a source of power 105; a source of feedstock 110; a catalyst 115; one or more sensors 120; one or more controllable components 125; and a controller 130. The chemical processing apparatus 101 is configured for processing the feedstock provided by the source of feedstock 110 to produce a product 135. Preferably, the chemical process is the production of methanol using electrolytic hydrogen. The precise apparatus shown in the figure is merely exemplary. The source of power 105 may be a source of electricity. In particular embodiments, the source of power 105 is a renewable source. For example, the source of power 105 may include at least one of: a solar panel; a wind turbine; and / or a hydroelectric generator. The power may be used for some or all of the components of the chemical plant 100. The source of feedstock 110 may be a supply of raw materials, or may be apparatus that processes raw materials to provide a refined feedstock 110. In the example shown in Figure 1, the source of feedstock 110 may be an electrolysis system, powered by the source of power 105 (in this case renewable). Carbon dioxide (CO2) may be passed to the feedstock 110 from any source, such as direct air capture, CO2 recovery from combustion gases, or from sequestered sources of CO2. For example, the electrolysis system may use electricity supplied by a renewable source 105 to produce hydrogen as a feedstock 110 for the chemical plant 100. It is also possible to consider the electrolysis system 110 to be a component of the chemical processing apparatus 101, in which case, the source of fluid for electrolysis would be considered the source of feedstock 110. In either case, the intermittent nature of the power from the power source 105 can produce a significant, and unpredictable, variation in the production of hydrogen by the electrolysis system 110. The catalyst 115 can be of any type, but for a methanol plant may be a component of, or a coating of, pellets forming at least part of a reactor bed. The catalyst 115 may be of a type that fouls or degrades over time. The performance of a catalyst 115 can vary based on temperature, pressure and the presence of contaminants in the feedstock. It is typically not possible to sense the state of the catalyst directly. Accordingly, this must be inferred from the set of parameters obtained from the sensor data using a model as discussed below. Catalysts are often temperature-sensitive and over-heating can result in permanent damage that reduces the catalyst activity and selectivity, both of which are used in the design of the chemical plant. Therefore, maintaining the catalyst activity and selectivity are critical for optimal and safe production. The sensors 120 (there may be more than those depicted) may be conventional sensors for sensing parameters of the chemical process. The parameters of the chemical process may include one or more of: the parameters of the feedstock, the chemical plant, and / or the product produced. The sensors 120 may be located in, mounted on, or integrated within components of the chemical processing apparatus 101. The sensors 120 communicate the sensor data to the controller 130. Collectively, the sensors 120 provide a set of parameters representing a state of the chemical plant 100 and chemical process. In what follows, the phrase “set of parameters” refers to a plurality of values representing the quantity measure by each sensor. For example, these may be presented as a vector. For example, a set of parameters may be a measurement of each of temperature, input pressure, output pressure, and flow rate. The sensors 120 can simultaneously provide sensor data or provide sensor data at different times. However, the controller 130 can form a set of parameters from the sensor data representing a common time. This may be done, for example, by grouping sensed data from similar times, or by interpolating the time series of sensed values obtained from each sensor 120 to obtain a set of parameters for a particular time. Irrespective of the particular method, the controller 130 is arranged to use the sensors 120 to repeatedly obtaining a set of parameters of the chemical process. The sensors 120 may include one or more of: a temperature sensor; a flow rate sensors; a valve sensor for sensing the opening state of a valve; a pH sensor; a pressure sensor; a concentration sensor; etc. The controllable components 125 (there may be more than those depicted) may include one or more of: an actuator; a valve; a pump; a heater; a cooler; and / or an agitator. The controllable components 125 may be controlled by the controller 130 to influence the chemical process. The controller 130 communicates with the sensors 120 and controllable components 125 to monitor and influence the chemical process. The controller 130 may include a local processor or a remote processor. For example the controller 130 may include a computer. The controller 130 may include storage, such as a local historian, and / or store data on a remote server. The controller 130 may include or implement one or more control systems arranged to maintain certain parts of the chemical process at set points. In this context a set point is a desired parameter of the chemical process (whether directly measurable by the sensors 120 or not). For example, one set point may define a particular flow rate along a conduit of the chemical processing apparatus 101, while another set point may define a particular temperature within a chamber or vessel of the chemical processing apparatus 101. The controllable components 125 may be more or less direct in their influence on the set points. For example, a heater jacket may be a controllable component 125 that directly heats a vessel with the temperature of the vessel being desired to equal a set point. Alternatively, the influence may be indirect, such as a downstream flow rate of the product being the result of a chemical reaction in the vessel, influenced by that heater jacket, with the flow rate being desired to equal a set point. Moreover, the set points might not be directly measurable, but might be inferred. The controller 130 is arranged to control the controllable components 130 to influence the chemical process. For example, one controllable component 125 may be a heater and the controller 130 may increase the heat output of the heater. As another example, one controllable component 125 may be a continuously variable valve and the controller 130 may open or close the valve to modify a flow rate. The controller 130 may include a processor and be able to access a model of the chemical process. The model can be used to simulate the chemical process by using the set of parameters as inputs. It may be preferable to pre-process the set of parameters to remove noise. This could be done in the conventional way, for example, by using a low pass filter on each parameter individually to remove short term effects so that the data more closely represents a general trend. Preferably, however, as will be described below, entire sets of parameters may be classified as outliers using a faulty data detection method and discarded so that they are not used by the controller 130 in the simulation of the chemical process. That is, the entire set of parameters for a given time may be discarded. The controller 130 uses the remaining sets of parameters to estimate a set of control variables for the chemical plant. It estimates the control variables by using the model and the sets of parameters to simulate the chemical process and selecting the control variables that optimise the chemical process with respect to some objective function. The objective function may be any suitable function, but could be a function of one or more of: product yield (e.g. methanol yield); energy efficiency of the chemical plant; catalyst performance; catalyst lifespan; financial profitability of the chemical plant; feedstock use; etc. The model 130A may be trained. The model 130A may include a set of model coefficients. For example, the model 130A may include a set of model coefficients of the functions that enable the controller 130 to simulate the chemical process and thereby establish the effect of the control variables on the objective function for a given set of parameters. The model coefficients may be derived from training data in a known way or by the method set out below. The model 130A may be a mechanistic model. Such a model may comprise systems of deterministic equations describing underlying physical processes, with the coefficients of those equations being the model coefficients. The model 130A may be a statistical model. Such a model may comprise a representation of the data used to train the model, with the model coefficients defining the representation. A preferred statistical model is a Gaussian process ora multi-objective Gaussian process, which may be derived from training data by Gaussian process regression. This technique is known in mathematics. An explanation of how Gaussian process regression can be used for modelling chemical processes may be found in “A Bayesian data modelling framework for chemical processes using adaptive sequential design with Gaussian process regression" by L. Fleming et al, Applied Stochastic Models in Business and Industry, vol. 38, no. 5, pp. 787-805, 2022. A preferred kernel function for this application is Matern. An alternative statistical model is a neural network. This technique is known in mathematics. An explanation of how neural networks can be used for modelling chemical processes may be found in “The Rise of Neural Networks for Materials and Chemical Dynamics" by M Kulichenko et al, J. Phys. Chern. Lett., 12(26), 6227-6243, 2021. Figure 2 is a flow chart of a method of controlling a chemical plant 100. The method comprises the steps of: providing a feedstock 210; processing the feedstock to produce a product 220; repeatedly obtaining a set of parameters using sensors 230; derive a set of control variables 240; and operating the chemical plant based on the sets of parameters 250. The step of providing a feedstock 210 may involve supplying feedstock to a chemical plant 100 at an unpredictable rate. For example, the step of providing a feedstock 210 may involve supplying hydrogen to a chemical plant 100. This may comprise using an electrolysis system 110 to obtain the hydrogen. The electrolysis system 110 may be powered by a renewable source of power 105. In this way, the rate of supply of hydrogen is dependent upon unpredictable wind conditions. The step of processing the feedstock 220 may involve processing the feedstock in the chemical plant 100 to implement a chemical process to produce methanol (for example, using hydrogen as a feedstock). The step of repeatedly obtaining a set of parameters of the chemical process using sensors 230 preferably involves repeatedly measuring at least one parameter of the chemical process using at least one sensor 120 and providing this to the controller 130. More preferably, this involves repeatedly measuring at least one parameter of the chemical process using each of a plurality of sensors 120. Optionally, this additionally or alternatively involves repeatedly indirectly estimating (i.e., not measuring) at least one parameter of the chemical process using the plurality of sensors 120. It is not essential for the sensors 120 to obtain sensor data simultaneously. Some parameters may vary more slowly than others and so it may be unnecessary to measure or estimate that parameter as frequently as a more fast-changing parameter. Even so, the controller 130 is able to obtain a set of parameters from the sensor data. A set of parameters may be, for example, a vector representing the value of each parameters at a given time (e.g., whether measured, interpolated, or estimated). The step of deriving a set of control variables 240 may include estimating the set of control variables for the chemical plant that optimises a particular objective function. To achieve this, the controller 130 simulates the chemical process using the sets of parameters to estimate the set of control variables that optimise the objective function. The simulation is achieved by using a model, such as a mechanistic or statistical model. Preferably, the model used in step 240 is a statistical model, optionally trained using the method of Figure 3A. The set of parameters is an input for the statistical model. The objective function may include one or more terms (for example, a weighted sum) related to different desired outcomes. An example objective function may include a term related to a rate of production of methanol. In step 240, the controller 130 may simulate the chemical process using the set of parameters (and, optionally, the time series of previously obtained parameters and / or chemical process states derived from them) to derive the control variables that would produce the best value of the objective function (for example, those which maximise the rate of production of methanol). The step of operating the chemical plant 100 based on the sets of parameters 250 may therefore involve operating the chemical plant 100 in accordance with control variables derived by the simulation of the chemical process using the sets of parameters. The chemical plant 100 may be configured to operate in one of a plurality of operating modes. Each operating mode may be defined by a plurality of the control variables being within respective predefined ranges. The controller 130 may operate the chemical plant 100 in accordance with an operating mode by automatically adjusting the control variables within their respective predefined ranges forthat operating mode. In general, the controller 130 may be programmed to operate the chemical plant 100 by adjusting the control variables within respective predefined ranges of the one or more operating modes. Figure 3A shows a flow chart of a method of training a statistical model 300 for use in the method of Figure 2, in which the chemical process is simulated using a statistical model in step 240 in order to derive the control variables. In this embodiment, the method of training the statistical model is a two-stage process (the process uses what is known in mathematics as transfer learning). Two training sets of data are obtained, a first training set and a second training set (despite the terms “first” and “second”, the order of obtaining the data is unimportant). The statistical model is then trained (i.e., its model coefficients determined) on both the first set of training data and the second set of training data. Whilst it is possible to simply pool the first and second training data, it is preferred to use the approach of transfer learning as is well known in the field. In this context, the first training data would be considered the source task data, and the second training data would be considered the target task data. The method comprises the steps of: providing a mechanistic model 310; generating first training data from the mechanistic model 320A; obtaining second training data from a real process 320B; and training a statistical model of the chemical process via transfer learning using both the first training and second training data 340. Importantly, the order in which the first and second training data is unimportant and non-limiting (for example, step 320B may be first). Preferably, the methods of detecting faulty data 400, 500, 600 discussed below are applied to the second training data in step 330. The step of providing a mechanistic model of the chemical process 310 may involve providing a set of deterministic equations describing underlying physical processes within the chemical plant 100. As is known in the field, the coefficients of those equations may be derived from real-world historic data, or may be derived analytically from first principles and / or estimates. The step of generating first training data using the mechanistic model 320A may include simulating the use of the chemical plant 100 across a first range of the control variables and parameter values. The first range is not a one-dimensional range, but the range for all of the possible control variables and parameter values. For example, the first training data may be generated by analysing the response of the mechanistic model to a range of input flow rates (e.g., from the minimum possible to the maximum possible) in combination with a range of input compositions (e.g., from the most rich to the to the most lean) in combination with a range of reactor temperatures (e.g., from the greatest heat input to the lowest heat input). Moreover, uncontrolled variable of the model may be perturbed to provide extra data not achievable from a real chemical plant. The first range may be a range that encompasses extreme values of the control variables and parameter values in a plurality of permutations. The first range may extend beyond the collective ranges of control variables and parameter values defined for the operating modes of the chemical plant 100 which the controller 300 is programmed to implement. In other words, generating first training data using the mechanistic model may comprise generating training data spanning a first range of at least one of the set of parameters using control variables outside the range used in the one or more operating modes of the chemical plant. The step 320B of obtaining second training data involves using a chemical plant (preferably the chemical plant 100 that will use the statistical model) to implement the chemical process. The sensors 120 of the chemical plant 100 are used to sense sets of parameters as the chemical plant is used to carry out a test process. In some cases, it may be appropriate to obtain second training data from a different chemical plant than the chemical plant 100 in which the statistical model will be employed. The test process may be designed to mimic the expected operating range of the control variables within the operating modes of the chemical plant 100 for the chemical process. In other words, obtaining second training data using a chemical plant may comprise obtaining training data spanning a second range of at least one parameter of the set of parameters using control variables within the range used in the one or more operating modes of the chemical plant. In practice, it may be appropriate to sample the data obtained from the chemical plant in order to obtain the second training data. This can be done by using known methods, preferably using QuasiMonte Carlo or Monte Carlo sampling. Accordingly, the first range of the set of parameters is broader than, and encompasses, the second range of the set of parameters. Put another way, for each parameter in the set of parameters, the range of values for that parameter in the first training data is greater than for the second training data. Figure 3B shows a representation in two-dimensional representation of training data for use in the method of Figure 3A. While two dimensions are shown for the sake of a simple depiction, the training data would of course have a much higher number of dimensions. As can be seen in Figure 3B, the first training data spans the state space, while the second training data provides a cluster representing the range of values that would result from an operating mode (in the example shown, only one operating mode was used, whereas multiple clusters may be present for multiple operating modes). The real data, the second training data, can enable the model to be more accurate within the expect ranges for the chemical plant 100, while the data generated from the mechanistic model, the first training data, can enable the model to be more accurate in extrapolating to otherwise unrepresented points in the state space. The first training data therefore represents a broader theoretical span of the state space for the chemical plant 100 and chemical process than the second training data. The second training data is representative of the preferred sub-region of that state space in which the chemical process is expected to produce the best results (e.g., minimising the objective function). Preferably, the number of data points in the first training data is greater than the number of data points in the second training data. The step of training a statistical model of the chemical process via transfer learning using both the first training data and the second training data provides a trained statistical model 340. This step may use any known training methodology for the particular model being used. For example, when a neural network is used as the statistical model, for example, this can simply be trained sequentially using back-propagation on the first training data. The model may then be refined on the second training data. As another example, when a Gaussian process (or multi-objective Gaussian Process) is used as the statistical model, for example, a suitable method fortraining the statistical model using the first and second training data is known in the art as coregionalization (sometimes known as co-kriging). An example of how this may be carried out is described in “Multi-task Gaussian Process Prediction” by Edwin V. Bonilla, Kian Ming A. Chai, and Christopher K. I. Williams. In Advances in Neural Information Processing Systems 20: NIPS'08. As a further available alternative, well known hierarchical Bayesian methods may be employed. Figure 4 shows a flow chart of a method of detecting faulty data 400 for use in the method of Figure 2. The method 400 may be particularly beneficial to remove anomalous data from the sets of the parameter. Anomalous data removal is important for chemical processes that include significant nonlinearities, such as the production of e-methanol. The method of detecting faulty data 400 comprises: projecting the sets of parameters into a lower dimensional space 410; identifying outliers in the lower dimensional space 415; discarding the sets of parameters corresponding to the outliers 420; and then using the remaining sets of parameters in the simulation 425. The step of projecting the sets of parameters into a lower dimensional space 410 provides a set of projected parameters. For example, using training data (for example the first or second training data discussed above, or some combination of both), a dimensionality reduction technique known in mathematics, such as principal component analysis (PCA) or canonical correlation analysis (CCA) may be used to generate a projection function that projects the data into a smaller number of dimensions. PCA, for example, provides an ordered set of orthogonal vectors that sequentially describe the directions of greatest variance in the training data. By removing the lowest order of the vectors, the lower dimension projection provides a representation of the sets of parameter data that best explains the potential variability in the data. CCA provides a similar approach, identifying an ordered set of orthogonal vectors that sequentially describe the directions of greatest correlation among the parameters in the training data. The step of identifying outliers in the lower dimensional space 415 involves projecting the parameters into the lower dimensional space to provide a lower dimensional representation of the set of parameters. The lower dimensional representations are then classified to establish if they are outliers. Classifying a lower dimensional representation of a set of parameters as an outlier, in turn, indicates that the set of parameters from which it was derived is also an outlier. A preferable way to classify the lower dimensional representations as outliers or not, is to use a metric and compare this with a threshold. For example, a metric may be calculated from the lower dimensional representation of the set of parameters, and the sets of parameters discarded as outliers or retained by comparison of the metric with a threshold. Preferred metrics include Hotelling's t-squared statistic or mean squared prediction error. Alternatively, Euclidean distance or Mahalanobis distance may be used. The step of discarding the sets of parameters corresponding to the outliers 420 is thus carried out for each of the sets of parameters based on an analysis of the lower dimensional representation of that set of parameters. Only the sets of parameters not discarded are used in the simulation in step 425. The chemical plant 100 is thus operated in accordance with control variables based on the retained sets of parameters (i.e. the retained sets of parameters after discarding those considered to be outliers) using a simulation of the chemical process to estimate a set of control variables for the chemical plant that optimises an objective function. Put another way, in step 425 (which may represent an implementation of step 240), the chemical process is simulated using only the retained parameters (those that are not outliers) to estimate a set of control variables for the chemical plant that optimises an objective function. Advantageous results have been achieved by carrying out this cleaning of the data in the lower dimensional space. Moreover, while PCA and CCA (along with other such methods) can each provide suitable lower dimensional spaces, the combination of multiple lower dimensional spaces can provide further advantages. The use of a plurality of different lower dimensional spaces can reduce the errors in identifying outliers in the sets of parameters. As shown in Figure 5, the plurality of lower dimensional spaces provide a plurality of projections of the sets of parameters into the respective lower dimensional spaces. The application of metrics set out above can be repeated for each lower dimensional representation. A combined metric can be obtained (for example, a weighted sum of the metric calculated in each lower dimensional space) and compared with a threshold. In an alternative approach, shown in Figure 6, the plurality of lower dimensional spaces provide a plurality of projections of the sets of parameters into the respective lower dimensional spaces. The application of metrics set out above can be repeated for each lower dimensional representation and a respective threshold may be applied to the metrics in each lower dimensional space separately. A system of voting may then be employed across the plurality of lower dimensional spaces, or logic may be employed. For example, the set of parameters may be discarded if it is identified as an outlier in every lower dimensional space (that is, an “AND” test). Figure 5 shows a flow chart of a method of detecting faulty data 500 for use in the method of Figure 2. The method of detecting faulty data 500 comprises the following steps: projecting the sets of parameters into a first lower dimensional space 510A; projecting the sets of parameters into a second lower dimensional space 510B; calculating a first metric in the first lower dimensional space 515A; calculating a second metric in the second lower dimensional space 515B; calculating a combined metric 518; discarding the sets of parameters corresponding to the outliers 520; and then using the remaining sets of parameters in the simulation 525. The combined metric 518 could be a weighted sum of the first and second metrics. The step of discarding the sets of parameters corresponding to the outliers 520 is thus carried out for each of the sets of parameters based on an analysis of the metrics from both lower dimensional representations of that set of parameters. Figure 6 shows a flow chart of a further method of detecting faulty data 600 for use in the method of Figure 2. The method of detecting faulty data 600 comprises the following steps: projecting the sets of parameters into a first lower dimensional space 610A; projecting the sets of parameters into a second lower dimensional space 610B; identifying outliers in the first lower dimensional space 615A; identifying outliers in the second lower dimensional space 615B; discarding the sets of parameters corresponding to the outliers 620; and then using the remaining sets of parameters in the simulation 625. The identification of outliers in the first lower dimensional space in step 615A and the identification of outliers in the second lower dimensional space in step 615B may be carried out as described above for step 420. The step of discarding the sets of parameters corresponding to the outliers 620 may be carried out when the results of steps 615A and 615B both indicate that the respective lower dimensional representation of the set of parameters is an outlier, and thus agree that the corresponding set of parameters is an outlier. The statistical method of step 240 may itself include a step of dimensionality reduction as a preprocessing step to provide model input data in a lower dimensional space. Importantly, it is preferable that the lower dimensional spaces used in steps 410,510A, 510B, 610A, 610B are different from any lower dimensional space forming pre-processing for the statistical model. That is, the dimensional reduction techniques optimal for the methods of detecting faulty data are not necessarily suitable for the statistical model. On the other hand, it is preferable that the second training data used in step 340 fortraining the statistical model is cleaned in step 330 using one of the methods of detecting faulty data 400, 500, 600. Similarly, it is preferable that step 420 of method 400 simulates the chemical process using the model trained in accordance with method 300. The following sets out preferred embodiments in the form of clauses. Clauses: Clause 1. A method of controlling a chemical process, comprising: providing a feedstock to a chemical plant; processing the feedstock in the chemical plant to implement a chemical process to produce a product; repeatedly obtaining a set of parameters of the chemical process using sensors; and operating the chemical plant in accordance with control variables based on the sets of parameters by: discarding outliers from the sets of parameters; and simulating the chemical process using the retained sets of parameters to estimate a set of control variables for the chemical plant that optimises an objective function, wherein the step of discarding outliers from the sets of parameters comprises: projecting the sets of parameters into a first lower dimensional space to provide projected parameters; identifying outliers using the projections of the parameters in the first lower dimensional space; and discarding at least some of the sets of parameters corresponding to the outliers. Clause 2. The method of clause 1, wherein discarding at least some of the sets of parameters corresponding to the outliers comprises: projecting the sets of parameters into a second lower dimensional space to provide projected parameters, wherein the second lower dimensional space is different from the first lower dimensional space; identifying outliers from the projected parameters in the second lower dimensional space; and discarding the sets of parameters corresponding to the outliers in both the first and the second lower dimensional spaces. Clause 3. The method of clause 1 or clause 2, wherein the outliers are identified in the first lower dimensional space by: obtaining training data by sensing a set of parameters of a test process; applying dimensionality reduction to the training data to identify a first lower dimensional space; calculating a metric from the projection of the training data in the first lower dimensional space; deriving a threshold from the projection of the training data in the first lower dimensional space; and applying the threshold to the metric calculated from the projected parameters. Clause 4. The method of clause 2 or clause 3, wherein the metric is one or more of: Hotelling's t-squared statistic; or mean squared prediction error. Clause 5. The method of any preceding clause, further comprising deriving the first lower dimensional space from the sets of parameters using at least one of: principal component analysis; canonical component analysis. Clause 6. The method of any preceding clause, wherein simulating the chemical process using the retained sets of parameters comprises projecting the sets of parameters into a third lower dimensional space to provide input parameters for the simulation, wherein the third lower dimensional space is different from the first lower dimensional space. Clause 7. The method of any preceding clause, wherein simulating the chemical process using the retained sets of parameters comprises using a Gaussian process model. Clause 8. A method of controlling a chemical process, comprising: providing a feedstock to a chemical plant; processing the feedstock in the chemical plant to implement a chemical process to produce a product; repeatedly obtaining a set of parameters of the chemical process using sensors; and operating the chemical plant in accordance with control variables based on the sets of parameters using a simulation of the chemical process to estimate a set of control variables for the chemical plant that optimises an objective function, wherein the simulation is obtained by: providing a mechanistic model of the chemical process; generating first training data using the mechanistic model; obtaining second training data by sensing a set of parameters of a test process implemented by a chemical plant; and training a statistical model of the chemical process using both the first and the second training data to provide a trained statistical model. Clause 9. The method of clause 8, wherein the step of training a statistical model of the chemical process using both the first and the second training data to provide a trained statistical model uses transfer learning. Clause 10. The method of clause 8, wherein the step of training a statistical model of the chemical process using both the first and the second training data to provide a trained statistical model uses co-regionalization. Clause 11. The method of any one of clause 8 to 10, wherein: generating first training data using the mechanistic model comprises generating training data spanning a first range of at least one of the set of parameters; obtaining second training data comprises obtaining training data spanning a second range of the at least one of the set of parameters; and the first range is broader than and encompasses the second range. Clause 12. The method of any one of clause 8 to 11, wherein: generating first training data using the mechanistic model comprises generating training data having a first number of data points; obtaining second training data comprises obtaining training data having a second number of data points; and the first number of data points is greater than the second number of data points. Clause 13. The method of any one of clause 8 to 12, wherein: the chemical plant is configured to operate in one of a plurality of operating modes; each operating mode is defined by a set of control variables being within respective predefine ranges; and generating first training data using the mechanistic model comprises generating training data spanning a first range of at least one of the set of parameters using control variables outside the ranges used in the plurality of operating modes. Clause 14. The method of any one of clause 8 to 13, wherein: obtaining second training data comprises implementing the chemical process in the chemical plant; and sensing a set of parameters of the chemical process using the sensors. Clause 15. The method of any one of clause 8 to 14, wherein: obtaining second training data comprises implementing the chemical process in a further chemical plant; and sensing a set of parameters of the chemical process using a set of sensors of the further chemical plant. Clause 16. The method of any one of clause 8 to 15, wherein operating the chemical plant in accordance with control variables based on the sets of parameters using a simulation of the chemical process comprises discarding outliers from the sets of parameters. Clause 17. A method of controlling a chemical process, wherein operating the chemical plant in accordance with control variables based on the sets of parameters using a simulation of the chemical process to estimate a set of control variables for the chemical plant that optimises an objective function comprises: discarding outliers from the sets of parameters; and simulating the chemical process using the retained sets of parameters to estimate a set of control variables for the chemical plant that optimises an objective function, wherein the step of discarding outliers from the sets of parameters comprises: projecting the sets of parameters into a first lower dimensional space to provide projected parameters; identifying outliers using the projections of the parameters in the first lower dimensional space; and discarding at least some of the sets of parameters corresponding to the outliers. Clause 18. A method of controlling a chemical process, wherein: obtaining second training data by sensing a set of parameters of a test process implemented by a chemical plant comprises discarding outliers from the sets of parameters of the second training data; and the step of discarding outliers from the sets of parameters comprises: projecting the sets of parameters into a first lower dimensional space to provide projected parameters; identifying outliers using the projections of the parameters in the first lower dimensional space; and discarding at least some of the sets of parameters corresponding to the outliers. Clause 19. The method of clause 17 or clause 18, wherein discarding at least some of the sets of parameters corresponding to the outliers comprises: projecting the sets of parameters into a second lower dimensional space to provide projected parameters, wherein the second lower dimensional space is different from the first lower dimensional space; identifying outliers from the projected parameters in the second lower dimensional space; and discarding the sets of parameters corresponding to the outliers in both the first and the second lower dimensional spaces. Clause 20. The method of any one of clauses 17 to 19, wherein the outliers are identified in the first lower dimensional space by: obtaining training data by sensing a set of parameters of a test process; applying dimensionality reduction to the training data to identify a first lower dimensional space; calculating a metric from the projection of the training data in the first lower dimensional space; deriving a threshold from the projection of the training data in the first lower dimensional space; and applying the threshold to the metric calculated from the projected parameters. Clause 21. The method of any one of clauses 18 to 19, wherein the metric is one or more of: Hotelling's t-squared statistic; or mean squared prediction error. Clause 22. The method of any one of clauses 17 to 21, further comprising deriving the first lower dimensional space from the sets of parameters using at least one of: principal component analysis; canonical component analysis. Clause 23. The method of any one of clauses 17 to 22, wherein simulating the chemical process using the retained sets of parameters comprises projecting the sets of parameters into a third lower dimensional space to provide input parameters for the simulation, wherein the third lower dimensional space is different from the first lower dimensional space. Clause 24. The method of any one of clauses 17 to 23, wherein simulating the chemical process using the retained sets of parameters comprises using a Gaussian process model. Clause 25. The method of any one of clause 8 to 24, wherein the statistical model is obtained by Gaussian process regression. Clause 26. The method of any preceding clause, wherein: the chemical plant is configured to operate in one of a plurality of operating modes; and each operating mode is defined by a set of control variables. Clause 27. The method of any preceding clause, wherein the chemical process is powered at least in part by renewable energy. Clause 28. The method of any preceding clause, wherein the product is methanol. Clause 29. The method of any preceding clause, wherein the feedstock is a renewable source or wherein the feedstock is produced using renewable energy. Clause 30. The method of any preceding clause, wherein the chemical plant includes a catalyst. Clause 31. The method of any preceding clause, wherein the feedstock is produced by electrolysis powered by at least one renewable source. 07 05 25

Claims

1. A method of controlling a chemical process, comprising:providing a feedstock to a chemical plant;processing the feedstock in the chemical plant to implement a chemical process to produce a product;repeatedly obtaining a set of parameters of the chemical process using sensors; and operating the chemical plant in accordance with control variables based on the sets of parameters by:discarding outliers from the sets of parameters; andsimulating the chemical process using the retained sets of parameters to estimate a set of control variables for the chemical plant that optimises an objective function, wherein the step of discarding outliers from the sets of parameters comprises:projecting the sets of parameters into a first lower dimensional space to provide projected parameters;identifying outliers using the projections of the parameters in the first lower dimensional space; anddiscarding at least some of the sets of parameters corresponding to the outliers,wherein discarding at least some of the sets of parameters corresponding to the outliers comprises:projecting the sets of parameters into a second lower dimensional space to provide projected parameters, wherein the second lower dimensional space is different from the first lower dimensional space;identifying outliers from the projected parameters in the second lower dimensional space; and discarding the sets of parameters corresponding to the outliers in both the first and the second lower dimensional spaces.

2. The method of claim 1, wherein the outliers are identified in the first lower dimensional space by: obtaining training data by sensing a set of parameters of a test process;applying dimensionality reduction to the training data to identify a first lower dimensional space;calculating a metric from the projection of the training data in the first lower dimensional space;deriving a threshold from the projection of the training data in the first lower dimensional space; andapplying the threshold to the metric calculated from the projected parameters.

3. The method of claim 1 or claim 2, wherein the metric is one or more of: Hotelling's t-squared statistic; or mean squared prediction error.07 05 254. The method of any preceding claim, further comprising deriving the first lower dimensional space from the sets of parameters using at least one of: principal component analysis; canonical component analysis.

5. The method of any preceding claim, wherein simulating the chemical process using the retained sets of parameters comprises projecting the sets of parameters into a third lower dimensional space to provide input parameters for the simulation, wherein the third lower dimensional space is different from the first lower dimensional space.

6. The method of any preceding claim, wherein simulating the chemical process using the retained sets of parameters comprises using a Gaussian process model.

7. The method of any preceding claim, wherein:the chemical plant is configured to operate in one of a plurality of operating modes; and each operating mode is defined by a set of control variables.

8. The method of any preceding claim, wherein the chemical process is powered at least in part by renewable energy.

9. The method of any preceding claim, wherein the product is methanol.

10. The method of any preceding claim, wherein the feedstock is a renewable source or wherein the feedstock is produced using renewable energy.

11. The method of any preceding claim, wherein the chemical plant includes a catalyst.

12. The method of any preceding claim, wherein the feedstock is produced by electrolysis powered by at least one renewable source.

Citation Information

Patent Citations

  • Temporary expanding integrated monitoring network

    US20080103751A1

  • Automated model building and batch model building for a manufacturing process, process monitoring, and fault detection

    US20100057237A1

  • Anomaly detection from aggregate statistics using neural networks

    US20220019863A1

  • Outlier detection and management

    US20230288918A1