Controlling an industrial process

By employing a trained inverse neural network and RL to optimize convergence rates in model-based control systems, the method addresses model inaccuracies and reduces computational costs, improving the efficiency and accuracy of industrial process control.

WO2025242609A1PCT designated stage Publication Date: 2025-11-27ABB (SCHWEIZ) AG
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
PCT/EP2025/063708
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-05-20
Filing Date
2025-05-19
Publication Date
2025-11-27

AI Technical Summary

Technical Problem

Existing model-based control systems, such as Model Predictive Control (MPC), face challenges in accurately representing complex industrial processes due to limited system information and difficulty in determining actuator limitations, leading to inefficient and computationally intensive operations.

Method used

A method utilizing a trained inverse neural network and Reinforcement Learning (RL) to optimize convergence rates of process control signals, considering model mismatches and constraints, reducing computational effort by leveraging low-fidelity models and interactive learning.

Benefits of technology

This approach enhances the efficiency and accuracy of industrial process control by optimizing convergence rates, addressing model inaccuracies and reducing computational costs, making it feasible for various industrial applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure EP2025063708_27112025_PF_FP_ABST
    Figure EP2025063708_27112025_PF_FP_ABST
Patent Text Reader

Abstract

The invention is concerned with a method, process control device, process control system excitation arrangement, control device, process control system, computer program and computer program product for controlling an industrial process. The process control device obtains a prediction of a range of process control signals û(k: k + N) for the industrial process (32), which range of process control signals û(k: k + N) have been predicted using a prediction of a range of process response signals ŷ(k: k + N) in a model of the industrial process (32) and which range of process response signals ŷ(k: k + N) have been predicted based on a range of reference signals r(k: k + N) and applies the predicted process control signals û(k: k + N) in the control of the process (32).
Need to check novelty before this filing date? Find Prior Art

Description

[0001]CONTROLLING AN INDUSTRIAL PROCESS FIELD OF THE INVENTIONThe invention relates to the field of process control systems. The invention moreparticularly concerns a method, process control device, process control system, computerprogram and computer program product for controlling an industrial process.BACKGROUND OF THE INVENTION Model based control is a powerful tool used in process control systems. Model-basedcontrol is for instance described in WO 2007 / 001252.Model-based control - for example model predictive control (MPC) - relies on having amodel that mimics the behaviour of the actual controlled process.Thus, model-based control, such as Model Predictive Control (MPC), is a known tool forthe closed-loop optimal control of industrial processes with constraints and actuatorlimitations. The performance of an MPC algorithm depends on the system modelunderlying the MPC scheme. Designing such an accurate model is challenging because the practically available system information can usually only be used to obtain a simplified, and thus inaccurate, representation of the real complex system. Furthermore, the performance of an MPC algorithm also depends on understanding the actual limitations of actuators, a task that can often be challenging due to the difficulty in precisely determining their values. Similarly, Economic Model Predictive Control ((E)MPC) solutions are standard tools for the closed-loop optimal control of industrial processes with constraints and actuator limitations, and benefit from a rich theory to assess their closed-loop behaviour. However, unfortunately, MPC design process hinges on the quality of the model underlying the control scheme and accurate nonlinear models are often hard to come by. One way to improve on the situation is to adapt the MPC model based on a data-drivenapproach to better align with the actual system. Nonetheless, it is important to note thatsimply fitting the model to the data may not always yield the optimal policy within the (E)MPC scheme, and in some cases, it can even have adverse effects.One practical approach to reducing the model dependency of the (E)MPC solution is touse a (Reinforcement Learning) RL agent to directly tune the MPC scheme incorporatingrate and path constraints to increase the performance on the real system. This way, in fact, MPC will be used as a function approximator in RL to establish a connection between RL and economic MPC. The combination of learning and control techniques has been well studied in literature. However, the solutions derived from those frameworks are all subject to intensive and expensive computation, where economic MPC needs to be called in each episode, meaning exhaustive exploration of these environments would be inefficient and practically unfeasible.Thus, there is a need for an improvement and especially one where the computationaleffort is reduced. SUMMARY OF THE INVENTIONIt is therefore an objective of the invention to improve the control of an industrial processand especially to reduce the computational effort needed in the control.This objective is according to a first aspect achieved through a method for controlling anindustrial process, the method being performed by a process control device and comprising:obtaining a prediction of a range of process control signals û(k: k + N) for the industrialprocess, which range of process control signals û(k: k + N) have been predicted using aprediction of a range of process response signals y^(k: k + N) in a model of the industrialprocess and which range of process response signals y^(k: k + N) have been predictedbased on a range of reference signals r(k: k +N), andapplying one or more of the predicted process control signals û(k: k + N) in the controlof the process,wherein the predictions of the process response signals y^(k: k + N) have been madebased on the reference signals r(k: k +N) and a convergence rate α(k), where theconvergence rate α(k) of the predictions of the process response signals y^(k: k + N) isoptimized, which optimization has been made while considering constraints b of thepredicted process response signals y^(k: k + N), andthe optimization of the convergence rate α(k) has been made based on a considering of a determined first mismatch ϑ^between the model and the actual process together with the constraints b for the predicted process response signals.The objective is according to a second aspect achieved through a process control devicefor controlling an industrial process, the process control device comprising a processoroperative to:obtain a prediction of a range of process control signals û(k: k + N) for the industrialprocess, which range of process control signals û(k: k + N) have been predicted using aprediction of a range of process response signals y^(k: k + N) in a model of the industrialprocess and which range of process response signals y^(k: k + N) have been predictedbased on a range of reference signals r(k: k +N), andapply one or more of the predicted process control signals û(k: k + N) in the control of theprocess,wherein the predictions of the process response signals y^(k: k + N) have been made basedon the reference signals r(k: k +N) and a convergence rate α(k), where the convergence rateα(k) of the predictions of the process response signals y^(k: k + N) is optimized, whichoptimization has been made while considering constraints b of the predicted processresponse signals y^(k: k + N), andthe optimization of the convergence rate α(k) has been made based on a considering of a determined first mismatch ϑ^between the model and the actual process together with the constraints b for the predicted process response signals.The objective is according to a third aspect achieved through a process control systemcomprising the process control device of the second aspect.The objective is according to a fourth aspect achieved through a computer program for controlling an industrial process, the computer program comprising computer program code which when run by a processor causes the processor toobtain a prediction of a range of process control signals û(k: k + N) for the industrialprocess, which range of process control signals û(k: k + N) have been predicted using aprediction of a range of process response signals y^(k: k + N) in a model of the industrialprocess and which range of process response signals y^(k: k + N) have been predictedbased on a range of reference signals r(k: k +N), andapply the predicted process control signals û(k: k + N) in the control of the process,wherein the predictions of the process response signals y^(k: k + N) have been made basedon the reference signals r(k: k +N) and a convergence rate α(k), where the convergence rateα(k) of the predictions of the process response signals y^(k: k + N) is optimized, which optimization has been made while considering constraints b of the predicted processresponse signals y^(k: k + N),, andthe optimization of the convergence rate α(k) has been made based on a considering of a determined first mismatch ϑ^between the model and the actual process together with the constraints b for the predicted process response signals.The objective is according to fifth aspect achieved through a computer program productfor controlling an industrial process, the computer program product comprising a datacarrier with the computer program according to the fourth aspect.The model may comprise fitting parameters used to express the dependency of the processresponse to the process control signal to the actual process. The model may be of a typethat is suitable for MPC where MPC is an acronym for Model Predictive Control.The model may additionally be an MPC model having been trained using a training set ofprocess control signals and process response signals.The obtaining of the range of process control signals û(k: k + N) may comprise:predicting the range of process response signals y^(k: k + N) based on the range ofreference signals r(k: k +N), andpredicting the range of process control signals û(k: k + N) based on the predicted processresponse signals using the model.A predicted process response signal y^(k: k + i) at a time i after a current point in time kmay be determined as the reference signal r(k: k +i) at the time i after the current point intime k plus the convergence rate α(k) times the prediction of the process response signaly^(k: k + i − 1) at a time immediately preceding the time i after the current point in time kplus the reference signal r(k: k +i-1) at the time immediately preceding the time i after the current point in time k. The predicting of the range of process response signals may be performed in a process response predictor and the predicting of a range of process control signals may be performed in an inverse neural network, for instance in a trained inverse neural network.The process response predictor may be implemented by the processor or in the cloud. Alsothe inverse neural network may be implemented by the processor or in the cloud.The processor may thus be operative to implement a process response predictorconfigured to predict the range of process response signals and / or a trained inverse neuralnetwork configured to predict the range of process control signals û(k: k + N).The obtaining of the range of process control signals û(k: k + N) may further compriseoptimising the convergence rate α(k) while considering constraints b of the predicted process response signals. The optimizing may be performed in an optimizer, which may be provided by the processor or in the cloud. The processor may thus be operative to implement an optimizer configured to optimize the convergence rate.The range of reference signals r(k: k +N) may additionally be a range of optimalprocess response signals that the predicted process response signals are to converge against.The obtaining of the range of process control signals û(k: k + N) may further comprise:determining the first mismatch ϑ^ between the model and the actual process, andproviding the first mismatch ϑ^ for being considered together with the constraints forthe predicted process response signals in the optimizing of the convergence rate.The first mismatch ϑ^ may be determined in a Reinforcement Learning agent providedby the processor.The Reinforcement Learning agent may be implemented by the processor or in thecloud. The processor may thus be operative to implement a Reinforcement Learningagent configured to determine the first mismatch ϑ^. The Reinforcement Learningagent may additionally be an on-line Reinforcement Learning agent.Thereby optimization of the convergence rate α(k) may have been made using aReinforcement Learning agent considering the determined first mismatch ϑ^between the model and the actual process together with the constraints b for the predicted process response signals.The optimization of the convergence rate α(k) may also be considering constraints ofthe predicted control signals û(k: k + N).In this case the optimization of the convergence rate α(k) may have been made based on a considering of a determined second mismatch ϑ^between the model and theactual process together with the constraints for the predicted control signals û(k: k +N).The obtaining of the range of process control signals û(k: k + N) may in this casefurther comprise: determining the second mismatch ϑ^between the model and the actual process and supplying the second mismatch ϑ^for being considered together with the constraints for the predicted process control signals in the optimizing of the convergence rate α(k).The second mismatch ϑ^ may also be determined in the Reinforcement Learning agentprovided by the processor.The method may further comprise obtaining a determined mismatch and supplying itto an operator. In this case, the processor may be further operative to obtain adetermined mismatch and supply it to an operator.The predicted process response signals y^(k: k + N) may be process response signalspredicted for a current parametrisation θP(k) of the model.In this case the method may further comprise determining the currentparametrisation θp(k) of the model and supplying it to the inverse neural network forbeing considered in the prediction of a range of process control signals.In this case, the processor may be further operative to determine the currentparametrisation θp(k) of the model and supply it to the inverse neural network forbeing considered in the prediction of a range of process control signals. The determining of the current parametrisation may be made in the ReinforcementLearning agent. The parametrisation is also a capability factor for the inverse neuralnetwork that will limit the capability if the mismatch is high. As has been mentioned above, the model may have been trained using a set of trainingdata comprising a training set of input and output signals in the form of processcontrol signals and process response signals. The training set of input and outputsignals may be a set of historic input and output signals. Alternatively, the set of inputand output signals may be input and output signals obtained via a model of theprocess, such as an MPC model. Furthermore, in the training the set of input andoutput signals may have been fit to the predicted process response signals so that thepredicted process response signals are provided at a distance to corresponding constraints φ.In a variation of the first aspect, the method may further comprise training the modelusing the set of input and output signals. In a corresponding variation of the secondaspect, the processor may be further operative to train the model using the set of inputand output signals. The training may in this case additionally comprise fitting the set of input and outputsignals to the predicted process response signals so that the predicted processresponse signals are distanced from the corresponding constraints φ. This may berealized through the absolute value of the estimated process control signal being higher than or equal to the constraint φ or the difference between the absolute value of the estimated process control signal and the constraint φ being higher than or equal to zero. BRIEF DESCRIPTION OF THE DRAWINGS The subject matter of the invention will now be explained in more detail in the following text with reference being made to preferred exemplary embodiments which are illustrated in the attached drawings, of which:Fig. 1 schematically shows a vessel equipped with a process control device controlling anindustrial process on the vessel,Fig. 2 shows a block schematic of one realization of the process control device,Fig. 3 schematically shows a computer program product comprising computer readablecode which implements a process control function of the process control device,Fig.4 schematically shows the use of experimental data in a model of the industrial process,Fig. 5 schematically shows input signals to and output signals from the model used fortraining, Fig.6 shows a flow chart of a number of method steps in a first embodiment of a method of controlling an industrial process being performed by the process control function,Fig. 7 schematically shows a process response predictor, an inverse neural networkimplementing the process control system model of the process control function togetherwith and an actual process being controlled using a control signal output from the inverseneural network,Fig. 8 schematically shows the use of collected experimental data in the model of theindustrial process, which use is based estimations of the process response being distancedfrom a corresponding constraint,Fig. 9 shows a flow chart of a number of method steps in a second embodiment of amethod of controlling an industrial process being performed by the process control function, andFig. 10 schematically shows the process response predictor, the inverse neural network, anoptimizer and a reinforcement learning agent of the process control function together withthe actual process being controlled using a control signal output from the inverse neural network. DETAILED DESCRIPTION In the following description, for purposes of explanation and not limitation, specific details are set forth such as particular architectures, interfaces, techniques, etc. in order to provide a thorough understanding of the present invention. However, it will be apparent to those skilled in the art that the present invention may be practiced in other embodiments that depart from these specific details. In other instances, detailed descriptions of well-known devices and methods are omitted so as not to obscure the description with unnecessary detail.Fig. 1 schematically shows a vessel V 10 comprising a process control device PCD 12, whichis a device for controlling an industrial process or a part of an industrial process. In thepresent example the controlling is the controlling of one or more processes of the vessel10, such as path planning of the vessel 10 and controlling the vessel to follow the path.Thus, the process control device 12 controls one or more industrial processes of the vessel10.Fig. 2 shows one realization of the process control device PCD 12. The process controldevice 12 may comprise a processor PR 14 and a data storage 16 with computer program instructions 18 that, when executed by the processor 14, implements a process controlfunction PCF 19, perhaps in the form of a vessel control function. There is also a firstcommunication interface, here realized as a radio communication interface RCI 20 and a second communication interface, here realized as a computer communication interface 22. The radio communication interface 20 is also a transmitter that transmits radio signals. Itshould be realized that it is possible that the process control device 10 only comprises theradio communication interface and not the computer communication interface or viceversa. Optionally there is also a display D 24 connected to the processor 14 for displayingdata to an operator of the vessel.In some cases, some of the process control function 19 may also be implemented in thecloud or employ the cloud for processing. Parts of the process control function 19 may thusbe implemented on a computing resource in a data centre, such as on a server blade,where it is additionally possible that this part is being assigned to a computing resource bya function assigning unit. Alternatively, the process control function 19 may connect to thecloud and request processing and receive the result of the processing therefrom.The process control device 12 may thus comprise a processor 14 with associated programmemory 16 including computer program code 18 for implementing the process controlfunction 19.A computer program may also be provided via a computer program product, for instance in the form of a non-transitory computer readable storage medium or data carrier, like a CD ROM or a memory stick, carrying such a computer program with the computerprogram code, which will implement the process control function when being loaded intoa processor. One such computer program product in the form of a CD ROM 26 with theabove-mentioned computer program code 18 is schematically shown in fig.3.As has been described above, the process control device 12 may perform path planning forthe vessel 10.It should here be realized that the path planning function is merely a first example of aprocess control function that the process control device can implement. The processcontrol function 19 can be used for controlling other processes of the vessel 10. In fact, theprocess control function 19 is not limited to being used on the vessel 10. As a secondexample, it can be used in relation to control strategies for chemical reactors. As a thirdexample it can be used for optimizing energy consumption while improving the overall product quality in heat exchanger network (HEN) control. As a fourth example, theprocess control function can be used in control design in robotics.Path planning can in many cases be performed using model-based control, such as modelpredictive control (MPC), for instance in order to consider inevitable modelling errors andenvironmental disturbances such as wind, wave, and current. MPC ensures reliable and accurate trajectory planning, enhancing the vessel’s safety and performance.MPC may in the second example also be used to address the inherent instability of certainchemical processes, such as exothermic reactors, and maintain process variables at desired steady states. MPC can in the third example be used in the challenging task of controlling HENs, which exhibit nonlinear behaviour, complexity, and are prone todisturbances and noise. In the fourth example MPC may play a crucial role in robotics,particularly in scenarios where there are uncertainties in the robot dynamics or when manipulating unknown objects. By accounting for model uncertainties, MPC enables robust and adaptive control strategies for precise and efficient robotics operations.However, model-based control typically requires a large amount of processing. Aspects ofthe present disclosure are directed towards improving on this.To this point, one may investigate the possibility of running fast MPC based on model inversion by deep neural networks (Neural Network Inversion-based Model PredictiveControl) to overcome the expensive computational effort of the traditional algorithms.Neural-network inversion refers to the process of estimating the input sequence given its corresponding outputs. In the context of MPC, it means using neural network to predict the future input to the model for a desired optimal output trajectory. From a control point of view, a direct MPC algorithm tries to find the optimal input sequence to guarantee an optimal output trajectory towards a given reference profile. Alternatively, this problem can be oversimplified if the shape of control input sequence is already assigned. This can be achieved by using the inverse neural network concept.How processing may be done according to a first embodiment will now be described.According to aspects of the present disclosure predictions of process response signals areobtained from a range of given reference signals and these predictions are then used in amodel of an inverse neural network for predicting control signals to be used for controllingan industrial process.Before being put to use, the model may need to be trained.Fig. 4 schematically shows training of a model M 28 of the industrial process using inputand output signals, where the input signals u are control signals used in the control of theindustrial process, i.e. process control signals, and the output signals y are processresponse signals output from the industrial process and fig. 5 schematically showscollected data for use as input signals to and output signals from the model used fortraining and prediction. It is possible that a parameterisation θp of a cost function and constraints of an MPCformulation θC of the process is used, where the model 28 may comprise fittingparameters used to express the dependency of the process response to the process controlsignal to the actual process. The model may be a model that is suitable for MPC and mayadditionally be a low-fidelity model.The parameterized model θp may then be set to generate output data from a sequence ofinputs covering all possible variations of the fitting parameters θP, which output data y(k;k + p) are the process response signals used as input signals to the MPC model and theoutput data are the control signals (u(k; k:p) from the MPC model 28.For instance, if there are k + q historic values of the input signal y and output signal u fordifferent values of the fitting parameters θP, these may be used for training the model,while future signals in a range k – (k + q) may be used for predictions for different valuesof the fitting parameters θP. The range (k + q) – (k + p) may also be a range N in whichpredictions are to be made.Thus, it is possible to retrieve a sequence of controls from a batch of past and futureinputs / outputs for all possible values of the fitting parameters θp, where the past signalsmay be obtained in an interval k – (k + q) and future signals in a range where N = p – q,see fig.5 and 6.After the model 28 has been trained, it is then possible to predict process control signals uto be used in the control. How this may be done according to a first embodiment will nowbe described with reference being made to fig.6 and 7, where fig.6 shows a flow chart of a number of method steps in a first embodiment of a method of controlling an industrialprocess being performed by the process control function and fig. 7 schematically shows theprocess control function PRF 19 controlling an actual process AP 32, which processcontrol function 19 comprises a process response predictor PRP 29, an inverse neuralnetwork INN 30 implementing the model of the industrial process and the actual process32 being controlled using a control signal being output from the inverse neural network30.The operation may be started through the process control function 19 obtaining a range ofreference signals r(k: k N), S100, which reference signals may be or correspond to a rangeof desired process response signals, such as a number of determined positions along apath of the vessel 10.Thereafter, the process control function 19 may predict process response signalsy^(k: k + N) based on the range r(k: k +N) of reference signals, S110, where k is a currentpoint in time and N is a range of time instances following after the current point in time. The predicting may be made by the process response predictor 29 as a number ofequations identifying a range of predictions of the process response signal y^(k) asfunctions of the range of reference signals r(k).As an example, a prediction of a process response signal at a time i after a current point intime k may be formulized as:y^(k + i) = r(k + i) + α(k)(y^(k + i − 1) − r(k + i − 1)), ^ = 1 … Nwhere r(k) represents the reference (planned) profile and ^(^) is the convergence rate attime instant ^, which convergence rate α(k) is a rate with which the process responsesignal converges towards the reference signal.Thereby, a predicted process response signal y^(k + i) at the time i after the current point intime k may be determined as the reference signal r(k +i) at the time i after the currentpoint in time k plus the convergence rate α(k) times the prediction of the process responsesignal y^(k + i − 1) at a time immediately preceding the time i after the current point in time k plus the reference signal r(k +i-1) at the time immediately preceding the time i after the current point in time k. Thus, if the tasks of generating input control sequence is already assigned to the inverse neural network, the optimization problem of the MPC could be reformulated in a way tocalculate an optimal output trajectory (y^(k + i), ^ = 1 … ^) towards the reference profilewith respect to the convergence rate ^(^) for a prediction horizon N looking at a simplefirst order time-series system as indicated above.As an example, it is possible to look at a number of equations associated to the predictionhorizon N to sketch the optimal profile for: ^(^ + ^) = ^(^ + ^) + ^(^(^) − ^(^))^(^ + ^) = ^(^ + ^) + ^(^(^ + ^) − ^(^ + ^)). . . ^(^ + ^) = ^(^ + ^) + ^(^(^ + ^ − ^) − ^(^ + ^ − ^))The range of predictions of the process response signals may then be used by the trainedinverse neural network 30 in the model of the process control system for predicting therange of control signals for the industrial process, S120. The range of predictions of theprocess response signals may more particularly be used in the inverse neural network 30 together with a group of historic process response and control signals for predicting the range of control signals using the parameterised model of the process. If the range isdefined by k + N and the historic interval is defined by k – n, then the range of controlsignals u^(k: k + N)|θ^(k) for current values of the parametrisation θP(k) is predictedthrough the use of the range of predictions of the process response y^(k: k + N)|θ^(k) andthe historic control and process response signals u(k − n: k)|θ^(k) and y(k − n: k)|θ^(k)in the model.The model may here be a parametrized model defining a range of predicted output signalsû(k: k + N) as a function of the predicted input signals y^(k: k + N).Thereafter the range of predicted control signals û(k: k + N) may be used for controllingthe actual process 32. Thus, at least one of the predicted control signals may be applied inthe control of the industrial process 32, S130. The process response y(k) caused by thecontrol as well as the used output signal û(k) are also supplied to the inverse neuralnetwork 30 for being used in future estimations, such as for improving the used fittingparameters.As was mentioned above, one or more of the process response predictor 29 and the inverseneural network 30 may be provided in the cloud. If both are provided in the cloud, it is possible that the range of reference signals is provided to the cloud in a request for processing and the range of predicted process control signals are returned as a response to the request for being used in the control.A second embodiment will now be described with reference being made to fig. 8, 9, 10 and11, where fig.8 shows the training of the model, fig.9 shows a flow chart of a number ofmethod steps in the method of controlling the industrial process being performed by theprocess control function and fig. 10 shows the process control function used to control theactual process, where the process control function comprises the process control predictor, an optimizer, a reinforcement learning agent and the inverse neural network.In this case, the process control function PCF 19 implements an optimizer OP 34 and anonline Reinforcement Learning (RL) agent RLA 36 in addition to the process responsepredictor PRP 29 and inverse neural network INN 30.The optimizer 34 optimizes a control problem that in this case involves the optimizing ofthe convergence factor α. The RL agent 36 in turn determines at least one first mismatchbetween the MPC model and the actual process.The operation may in this case be started through training of the model 28 using atraining set of input and output data, S200, where the training set of input and outputdata may comprise collected historic data for the actual process or data generated throughthe use of an available process model, such as an MPC model. The training of the model 28may in this case also involve predicting the process response signal ^^ using the outputsignal u and the parametrisation θP and applying a constraint φ on this prediction. Thisconstraint φ can impact the performance and effectiveness of the control algorithm. Tothis scope, one can introduce the distance of the prediction ^^ from the constraint φ asanother output data set of the model. Thereby, the training may need to fulfil a criterionthat the absolute value of the prediction ^^ minus the constraint φ should be above zero.Put differently, the absolute value of the prediction should be equal to or higher than the constraint φ.The training may more particularly involve fitting the input-outputs (^(^: ^ + ^), ^(^: ^ +^)) and (^(^: ^ + ^), ^^(^: ^ + ^) − ^ ≥ 0) of the set, where the latter guarantees thenonequality constraints on the output variables (^^).After the training regular operation may be started. Operation may again involve theprocess control function 19 obtaining a range of reference signals r(k: k + N) for aprediction range N, S210, which reference signals may be or correspond to a range ofdesired process response signals that predicted process responses y^(k: k + N) are toconverge against. The reference signals r(k: k + N) are provided to the optimizer 34 as wellas to the process response predictor 29.Furthermore, in operation the RL agent 36 receives observations from the process 32,from which observations it is able to determine a first mismatch ϑ^ between the model 28and the actual process 32, S220, which first mismatch ϑ^may be a mismatch between the MPC formulation of the process and the actual process. The first mismatch may more particularly be a mismatch of a part of the MPC formulation that is related to the processresponse y, The RL agent 36 supplies the first mismatch ϑ^ to the optimizer 34. The RLagent 36 may additionally determine a second mismatch ϑ^, S230, which secondmismatch ϑ^may be a mismatch between the MPC formulation of the process and theactual process. The second mismatch ϑ^ may more particularly be a mismatch of a part ofthe MPC formulation that is related to the process control signal u. The RL agent 36 alsoprovides this second mismatch ϑ^ to the optimizer 34. Thus, the RL agent 36 may be anon-line RL agent 36.As before, the process response signal being predicted by the process response predictor29 is a function of the reference trajectory r and a convergence rate α.The optimizer 34 optimizes the convergence rate α based on a constraint b of the predictedinput signals y^(k), S240. This constraint may additionally be adjusted with the first mismatchϑ^. The first mismatch may for instance be added to the constraint b. Theoptimizer may also optimize the convergence rate α based on a constraint b on thepredicted control signal û(k), which constraint may be adjusted with the second mismatchϑ^. The second mismatch ϑ^ may for instance be added to the constraint.The optimizer 34 aims at finding the optimum convergence rate ^(^) to track thereference trajectory respecting the path constraints on y^. This can be expressed as: where ^^^() represents the trained inverse neural network.Put differently, the optimizing may involve: ^^^ ^ ^ subject to: •y^(k) = r(k) + α(k)(y^(k − 1) − r(k − 1))• |y^(k)| ≤ b + ϑ^• |u^(k)| ≤ b + ϑ^It is possible that additional constraints related to the effort and energy minimization are imposed to deal with non-uniqueness for over-actuated systems.After the convergence rate has been optimized, the process response predictor 29 maypredict a range of process response signals ^^(k: k + N) for the model based on the ranger(k: k + N) of reference signals, S250.The predicting may be made in the form of a number of equations identifying a range ofpredictions of the process response signal as a function of the range of reference signals.Also, here a prediction of a process response signal at a time i after a current point in timek may be formulized as: y^(k + i) = r(k + i) + α(k)(y^(k + i − 1) − r(k + i − 1)), ^ = 1 … NPut differently, a number of equations may be formed as:^(^ + ^) = ^(^ + ^) + ^(^(^) − ^(^))^(^ + ^) = ^(^ + ^) + ^(^(^ + ^) − ^(^ + ^)) . .^(^ + ^) = ^(^ + ^) + ^(^(^ + ^ − ^) − ^(^ + ^ − ^))The RL agent 36 also determines an updated parameter setting θp(k) of the model based on the observations from the actual process 32 and provides to the inverse neural network 30, S260. The range of predictions of the process response signals may then be used by the trained inverse neural network 28 in the model of the process control system for predicting therange of control signals for the industrial process, S270. The range of predictions of theprocess response signals may more particularly be used in the inverse neural network 30 together with a group of historic process response and control signals for predicting the range of control signals using the parameterised model of the process. If the range isdefined by k + N and the historic interval is defined by k – n, then the range of processcontrol signals u^(k: k + N)|θ^(k) for current values of the parametrisation θP(k) ispredicted through the use of the range of predictions of the process responsey^(k: k + N)|θ^(k) and the historic control and process response signals u(k − n: k)|θ^(k)and y(k − n: k)|θ^(k) in the model.Thereafter the predicted range of output signals may be used as control signals forcontrolling the actual process, S280. Thus, one or more of the predicted process controlsignals may be applied in the control of the industrial process.The first and / or the second mismatch may also be provided to the display 24 in order tobe displayed to an operator. Thereby the operator can interpret what the agent is doing.This means that the solution becomes less of a black box. Existence of the mismatchdoesn’t mean that the operator necessarily must do anything with it. However, as it existsas a biproduct, a Graphical User Interface (GUI) could be constructed from which theoperator can see what the agent is up to. Aspects of the present disclosure can also be described in the following way:The inverse neural network is trained on a LoFi model of the process that may becategorized by an uncertain parameter θP, which refers to the uncertainties of the model.The inverse neural network generates N-step ahead control sequences according to aplanned output trajectory and an uncertainty level.The optimizer employs a planning optimization problem to calculate an optimal outputtrajectory towards the reference profile subject to rate and path constraints based on agiven parameters setting. The RL agent is trained to adjust and adapt the MPC and network tuning based on theprocess status for generating a suitable factor for tuning MPC and inverse neural networkaccording to the current process status.Various aspects disclosed herein can be summarized in five steps:1. Develop a fully parameterized MPC scheme: boundaryconditions (model), constraints and actuators limitations based on a low fidelity model of the system.2. Utilize the proposed parameterized model to generate outputdata from a sequence of inputs covering all possible variations of the fitting parameters. (Collecting experimental data in the traditional sense, which means applying a sequence of u(k), to acquire y(k))3. Train an inverse neural network able to retrieve N-step aheadinput prediction û(k: k + N), based on output data (y(k)).4. Introduce a convergence rate towards a planned trajectory andincorporate it with the network by solving the proposed optimization problem.5. Run RL interacting with the fast MPC environment for trainingthe unknown parameters of MPC toward the maximization of the overall performance, rather than replicating the model more precisely.As a summary, there is proposed the use of Artificial intelligence (AI), particularlyReinforcement Learning (RL), for tuning a fast-MPC algorithm for optimal decision making of manipulated variables, which can have a significant impact on the processperformance in terms of disturbance rejection and constraints satisfaction and ultimatelythe profitability of running the process. The path constraints on the target outputs are handled by the optimization problem, while the remaining inequality constraints are all addressed with the network training sessions. The whole algorithm takes advantage of using a low fidelity model to find a global optimal sequence of control inputs based on a predefined (economic) objective function. Some of the advantages of the aspects described herein are:^ The short comings of conventional predictive optimization algorithms arecomplemented, where a high-fidelity model is either not available or too expensive to be redeemed within a moving horizon framework. ^A fast and computationally cheap optimization problem is used, which makes itindustrially feasible for many applications such as marine, pulp and paper, power plants and chemical processes. ^There is a novel constraint handling by incorporating the inequality constraintsinto the training sessions and thereby the computational effort required by theoptimizer is reduced.^ The introduced RL-guided MPC solution provides a dynamic shaping of theoptimization problem according to the actual process variations. ^There is a feasible learning process for the RL agent as it interacts with a fastenvironment thanks to the Neural Network Inversion-based MPC. ^There is a practical and meaningful interpretation of the Agent activities thanks tothe availability of a mismatch from which the operator can see what the agent is upto. While the invention has been described in connection with what is presently considered to be most practical and preferred embodiments, it is to be understood that the invention is not to be limited to the disclosed embodiments, but on the contrary, is intended to covervarious modifications and equivalent arrangements. Therefore, the invention is only to belimited by the following claims.

Claims

1. Claims 1. A method for controlling an industrial process (32), the method being performed bya process control device (12) and comprising the steps of: obtaining a prediction (S120; S260) of a range of process control signals û(k: k + N)for the industrial process (32), which range of process control signals û(k: k + N)have been predicted using a prediction (S110; S250) of a range of process responsesignals y^(k: k + N) in a model (28) of the industrial process (32) and which rangeof process response signals y^(k: k + N) have been predicted based on a range ofreference signals r(k: k +N), andapplying (S130; S270) the predicted process control signals û(k: k + N) in thecontrol of the process (32)wherein the predictions of the process response signals y^(k: k + N) have been madebased on the reference signals r(k: k +N) and a convergence rate α(k), where theconvergence rate α(k) of the predictions of the process response signals y^(k: k + N)is optimized (S240), which optimization has been made while consideringconstraints b of the predicted process response signals y^(k: k + N), andthe optimization of the convergence rate α(k) has been made based on aconsidering of a determined (S220) first mismatch ϑ^between the model (28) and the actual process (32) together with the constraints b for the predicted process response signals.

2. The method according to claim 1, wherein a predicted process responsesignal y^(k: k + i) at a time i after a current point in time k is determined as thereference signal r(k: k +i) at the time i after the current point in time k plus the convergence rate α(k) times the prediction of the process response signal y^(k: k + i − 1) at a time immediately preceding the time i after the current point intime k plus the reference signal r(k: k +i-1) at the time immediately preceding the time i after the current point in time k.

3. The method according to claim 1 or 2, wherein the optimization of the convergencerate α(k) has been made using a Reinforcement Learning agent considering the determined (S220) first mismatch ϑ^between the model (28) and the actual process (32) together with the constraints b for the predicted process response signals.

4. The method according to any previous claim, wherein the optimization (240) of theconvergence rate α(k) has also been made while considering constraints b of thepredicted control signals û(k: k + N) and possibly also together with a consideringof a determined (S230) second mismatch ϑ^ between the model and the actualprocess5. The method according to any previous claim, wherein the range of reference signalsr(k: k +N) is a range of optimal process response signals that the predicted processresponse signals y^(k: k + N) are to converge against.

6. The method according to any previous claim, further comprising obtaining adetermined mismatch and supplying it to an operator.

7. The method according to any previous claim, wherein the predicted processresponse signals y^(k: k + N) are process response signals predicted for a currentparametrisation θP(k) of the model (28).

8. The method according to claim 8, further comprising determining the currentparametrisation θp(k) of the model (28) and supplying the current parametrisation θp(k)forbeing considered in the prediction of a range of process control signals.

9. The method according to any previous claim, wherein the model (28) is an MPC(Model Predictive Control) model having been trained (S200) using a training set of input and output signals in the form of process control signals and process responsesignals.

10. The method according to claim 9, wherein in the training the training set of inputand output signals have been fit to the predicted process response signals in a wayso that the predicted process response signals are provided at a distance to constraints φ on these predicted process response signals.

11. A process control device (12) for controlling an industrial process (32), saidprocess control device comprising a processor (14) operative to:obtain a prediction of a range of process control signals û(k: k + N) for theindustrial process (32), which range of process control signals û(k: k + N) have beenpredicted using a prediction of a range of process response signals y^(k: k + N) in amodel (28) of the industrial process (32) and which range of process responsesignals y^(k: k + N) have been predicted based on a range of reference signals r(k:k+N), and apply the predicted process control signals û(k: k + N) in the control of the process(32) wherein the predictions of the process response signals y^(k: k + N) have been madebased on the reference signals r(k: k +N) and a convergence rate α(k), where the convergence rate α(k) of the predictions of the process response signals y^(k: k + N) isoptimized, which optimization has been made while considering constraints b of the predicted process response signals y^(k: k + N), andthe optimization of the convergence rate α(k) has been made based on a considering of a determined first mismatch ϑ^between the model (28) and the actual process (32) together with the constraints b for the predicted process response signals.

12. The process control device according to claim 11, wherein the processor (14) isoperative to implement a process response predictor (29) configured to predict the range of process response signals y^(k: k + N) and a trained inverse neural network(30) configured to predict the range of process control signals û(k: k + N).

13. A process control system comprising the process control device according to claim11 or 12.

14. A computer program for controlling an industrial process (32), the computerprogram comprising computer program code (18) which when run by aprocessor (14) causes the processor (14) to:obtain a prediction of a range of process control signals û(k: k + N) for theindustrial process (32), which range of process control signals û(k: k + N) have beenpredicted using a prediction of a range of process response signals y^(k: k + N) in amodel (28) of the industrial process and which range of process response signalsy^(k: k + N) have been predicted based on a range of reference signals r(k: k +N),and apply the predicted process control signals û(k: k + N) in the control of the process(32) wherein the predictions of the process response signals y^(k: k + N) have been madebased on the reference signals r(k: k +N) and a convergence rate α(k), where the convergence rate α(k) of the predictions of the process response signals y^(k: k + N) isoptimized, which optimization has been made while considering constraints b of the predicted process response signals y^(k: k + N), andthe optimization of the convergence rate α(k) has been made based on a considering of a determined first mismatch ϑ^between the model (28) and the actual process (32) together with the constraints b for the predicted process response signals.

15. A computer program product for controlling an industrial process (32), thecomputer program product comprising a data carrier (26) with said computer program according to claim 14.

Citation Information

Patent Citations

  • Apparatuses, systems, and methods utilizing adaptive control

    WO2007001252A1