Enhanced model prediction control system, prediction method and training method thereof
By introducing reinforcement learning models into the model prediction control system, online compensation decision deviation is achieved, the problem of deviation between the model and the actual controlled system is solved, prediction accuracy and control stability are improved, and control delay is avoided.
Patent Information
- Application Number
- CN202311689874.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-11-30
- Filing Date
- 2023-12-08
- Publication Date
- 2025-05-30
AI Technical Summary
When the model predictive control system faces environmental changes or uncertainties, the model deviates from the actual controlled system, resulting in incorrect control results, and the prior art requires offline retraining of the model, resulting in control delay.
The reinforced model prediction control system is adopted, combined with the model prediction controller and the reinforced learning model, and the reinforced learning model is used to compensate decision deviations online in real time to improve prediction accuracy without offline retraining of the model.
Improve prediction accuracy and stabilize device control under uncertain conditions, avoiding the occurrence of control delays.
Smart Images

Figure CN120065712A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a control system and its operation method, and more particularly to a reinforced model predictive control system, its prediction method and training method. Background Art
[0002] Industrial equipment can introduce model predictive control technology to control the dynamic change process of a control target under different control inputs, and use an optimization algorithm to find the optimal input that can make the control target reach the expected set value. However, the model used in model predictive control may deviate from the actual controlled system due to environmental changes or the existence of other uncertainties, resulting in incorrect control results.
[0003] To solve this problem, generally, the model must be removed and retrained offline. This method requires a large amount of training time and is prone to control delay. Summary of the Invention
[0004] The present invention relates to a reinforced model predictive control system, its prediction method and training method. When the model predictive controller faces a situation with uncertain problems, it can instantaneously compensate for decision-making deviations online through a reinforcement learning model to improve prediction accuracy and stabilize the control of the equipment, without the need to retrain the model predictive controller offline.
[0005] According to an aspect of the present invention, a reinforced model predictive control system is proposed. The reinforced model predictive control system is used for equipment and includes a model prediction controller (MPC), a reinforcement learning model, and a compensation unit. The model prediction controller is used to output a controller prediction control parameter value. The reinforcement learning model is used to output a learning prediction control parameter value. The compensation unit compensates the controller prediction control parameter value with the learning prediction control parameter value to obtain a compensated prediction control parameter value. The compensated prediction control parameter value is provided to the equipment to output a predicted target parameter value.
[0006] According to another aspect of the present invention, a prediction method for a reinforced model predictive control system is proposed. The prediction method for the reinforced model predictive control system is used for equipment and includes the following steps. Output a controller prediction control parameter value with the model prediction controller. Output a learning prediction control parameter value with the reinforcement learning model. Perform a compensation procedure based on the learning prediction control parameter value and the controller prediction control parameter value to obtain a compensated prediction control parameter value. Provide the compensated prediction control parameter value to the equipment to output a predicted target parameter value.
[0007] According to another aspect of the present invention, a training method for a reinforced model predictive control system is proposed. The reinforced model predictive control system compensates the model predictive controller through a reinforcement learning model. The training method is used for a device and includes the following steps. Train the reinforcement learning model with a first-stage training program. The first-stage training program includes: outputting a plurality of reference predictive control parameter values with a reference controller; providing these reference predictive control parameter values to the device to obtain a plurality of predicted target parameter values; training the reinforcement learning model through these reference predictive control parameter values and these predicted target parameter values. Train the reinforcement learning model with a second-stage training program. The second-stage training program includes: outputting a plurality of controller predictive control parameter values with the model predictive controller; outputting these reference predictive control parameter values with the reference controller; compensating each controller predictive control parameter value with each reference predictive control parameter value to obtain a plurality of corrected predictive control parameter values; providing these corrected predictive control parameter values to the device to obtain these predicted target parameter values; training the reinforcement learning model through these corrected predictive control parameter values and these predicted target parameter values. Train the reinforcement learning model with a third-stage training program. The third-stage training program includes: outputting these controller predictive control parameter values with the model predictive controller; outputting a plurality of learning predictive control parameter values with the reinforcement learning model; performing a compensation program based on these learning predictive control parameter values and these controller predictive control parameter values to obtain a plurality of compensated predictive control parameter values; providing these compensated predictive control parameter values to the device to obtain these predicted target parameter values; training the reinforcement learning model through these compensated predictive control parameter values and these predicted target parameter values.
[0008] For a better understanding of the above and other aspects of the present invention, the following specific embodiments are given and described in detail in conjunction with the accompanying drawings: BRIEF DESCRIPTION OF THE DRAWINGS
[0009] Figure 1 Schematically shows the control actions of a tank device 900 according to an embodiment;
[0010] Figure 2 Schematically shows a block diagram of a reinforced model predictive control system 100 according to an embodiment;
[0011] Figure 3 Schematically shows the prediction method of the reinforced model predictive control system 100;
[0012] Figure 4 Schematically shows the online update process of the reinforcement learning model 140;
[0013] Figure 5A Schematically shows the first-stage training program PS1 of the offline training method of the reinforcement learning model 140;
[0014] Figure 5B Schematically shows the first - stage training program PS1 of the offline training method of the reinforcement learning model 140;
[0015] Figure 6A Schematically shows the second - stage training program PS2 of the offline training method of the reinforcement learning model 140;
[0016] Figure 6B Schematically shows the second - stage training program PS2 of the offline training method of the reinforcement learning model 140;
[0017] Figure 7A Schematically shows the third - stage training program PS3 of the offline training method of the reinforcement learning model 140;
[0018] Figure 7B Schematically shows the third - stage training program PS3 of the offline training method of the reinforcement learning model 140;
[0019] Figure 8 Schematically shows the control results of the bucket - tank liquid - level height y1 of the tank - barrel device 900 by using the reinforcement model - predictive control system 100 and the general model - predictive control system; and
[0020] Figure 9 Schematically shows the control results of the bucket - tank temperature y2 of the tank - barrel device 900 by using the reinforcement model - predictive control system 100 and the general model - predictive control system.
[0021] 100: Reinforcement model - predictive control system;
[0022] 110: Detector;
[0023] 120: Model - predictive controller;
[0024] 121: Prediction model;
[0025] 122: Optimization searcher;
[0026] 130: Compensation judgment unit;
[0027] 140: Reinforcement learning model;
[0028] 150: Compensation unit;
[0029] 160: Range - limiting unit;
[0030] 180: Storage unit;
[0031] 300: Reference controller;
[0032] 800: Device;
[0033] 900: Tank equipment;
[0034] C10, C11, C12, C20, C21, C22: Curves;
[0035] PS1: First-stage training program;
[0036] PS2: Second-stage training program;
[0037] PS3: Third-stage training program;
[0038] S110, S120, S130, S140, S150, S170, S270, S280, S290, S530, S560, S570, S580, S590, S610, S620, S630, S640, S660, S670, S680, S690, S710, S720, S730, S740, S760, S770, S780, S790: Steps;
[0039] u1: Cold water feed flow valve opening;
[0040] u2: Hot water feed flow valve opening;
[0041] u3: Outlet water flow valve opening;
[0042] Ud: Reference predictive control parameter value;
[0043] Ui: Control parameter value;
[0044] Um: Controller predictive control parameter value;
[0045] Umd: Modified predictive control parameter value;
[0046] Umr: Compensated predictive control parameter value;
[0047] Ur: Learning predictive control parameter value;
[0048] Uw: Control parameter value;
[0049] Wc: Cold water;
[0050] Wd: Disturbance warm water;
[0051] Wh: Hot water;
[0052] Wo: Warm water;
[0053] X: Hidden variable value;
[0054] y1: Tank liquid level height;
[0055] y2: Tank temperature;
[0056] Yi, Yw: Target parameter values;
[0057] Yp: Predicted target parameter value. Detailed implementation manners
[0058] The technical terms in this specification refer to the customary terms in this technical field. If this specification explains or defines some terms, the explanations of these terms shall prevail according to the explanations or definitions in this specification. Each embodiment of the present invention has one or more technical features. On the premise of possible implementation, those with ordinary knowledge in this technical field can selectively implement some or all of the technical features in any embodiment, or selectively combine some or all of the technical features in these embodiments.
[0059] The following invention provides different features for implementing some implementation manners or examples of the present invention. The following describes specific examples of components and configurations to simplify some implementation manners of the present invention. Of course, such components and configurations are only examples and are not intended to be restrictive. In addition, some implementation manners of the present invention may repeatedly refer to reference signs and / or letters in various examples. This repetition is for the need of simplicity and clarity, and does not itself indicate the relationship between the various implementation manners and / or configurations discussed.
[0060] Please refer to Figure 1 , which schematically shows the control actions of the tank device 900 according to an embodiment. In the tank device 900, cold water Wc and hot water Wh are injected, and warm water Wo is output. The control parameter value Uw of the tank device 900, for example, includes the opening degree u1 of the cold water feed flow valve, the opening degree u2 of the hot water feed flow valve, and the opening degree u3 of the outlet water flow valve. By controlling the opening degree u1 of the cold water feed flow valve, the opening degree u2 of the hot water feed flow valve, and the opening degree u3 of the outlet water flow valve, the target parameter values Yw can be stabilized, such as the tank level height y1, the tank temperature y2, etc.
[0061] The tank device 900 can introduce the technology of model prediction control (MPC) to control the control parameter value Uw and predict the target parameter value Yw.
[0062] In practical applications, the tank device 900 may be affected by various uncertain interference factors. For example, there is additional interfering warm water Wd injected into the tank device 900, or the ambient temperature around the tank device 900 changes, resulting in inaccurate models when using the technology of model prediction control. In this embodiment, the inaccurate models can be compensated by reinforcement learning technology to improve the prediction accuracy.
[0063] The above Figure 1Taking the tank equipment 900 as an example for illustration, however, in factories such as petrochemical, paper-making, textile, and steel industries that adopt distributed control systems, there are also many devices (such as distillation towers, etc.) that require parameter prediction and monitoring. The technology of the present invention is not limited to the above-mentioned tank equipment 900, and any similar equipment can apply the technology of the present invention.
[0064] Please refer to Figure 2 , which schematically shows a block diagram of a reinforced model predictive control system 100 according to an embodiment. The reinforced model predictive control system 100 includes an observer 110, a model prediction control (MPC) 120, a compensation judgment unit 130, a reinforcement learning model 140, a compensation unit 150, a range limiting unit 160, and a storage unit 180. The observer 110, the model prediction controller 120, the compensation judgment unit 130, the reinforcement learning model 140, and the compensation unit 150 are used to perform various operations, controls, processes, and judgment procedures, such as circuits, wafers, circuit boards, or storage devices storing program codes. Alternatively, in an embodiment, the program codes can be loaded by a processing unit to execute the functions of the observer 110, the model prediction controller 120, the compensation judgment unit 130, the reinforcement learning model 140, and / or the compensation unit 150. The processing unit is, for example, a central processing unit (CPU), or other programmable general-purpose or special-purpose micro control units (MCUs), microprocessors, digital signal processors (DSPs), programmable controllers, application specific integrated circuits (ASICs), graphics processing units (GPUs), image signal processors (ISPs), image processing units (IPUs), arithmetic logic units (ALUs), complex programmable logic devices (CPLDs), field programmable gate arrays (FPGAs), or other similar components or combinations of the above components.
[0065] The storage unit 180 is used to store data, such as any type of fixed or removable random access memory (RAM), read-only memory (ROM), flash memory, hard disk drive (HDD), solid state drive (SSD), or similar components or combinations of the above components, and can be used to store multiple modules or various application programs executed by the processing unit.
[0066] In this embodiment, when the model predictive controller 120 faces a situation with uncertain problems, it can immediately compensate for decision-making biases online through the reinforcement learning model 140 to improve the prediction accuracy and stabilize the control of the device 800. The operations of the above components will be described in detail below with reference to the flowchart.
[0067] Please refer to Figures 2 - 3 , Figure 3 for an example to illustrate the prediction method of the reinforcement model predictive control system 100. In Figure 3 step S110, as Figure 2 shown, the detector 110 outputs the hidden variable value X based on the obtained control parameter values and target parameter values.
[0068] The control parameter values are, for example, the opening degrees u1 of the cold water feed flow valve, u2 of the hot water feed flow valve, and u3 of the outlet water flow valve as described above. The target parameter values are, for example, the bucket tank liquid level height y1 and the bucket tank temperature y2 as described above. During the prediction process of this embodiment, the detector 110 can directly obtain the current compensated predictive control parameter value Umr and the current predictive target parameter value Yp for calculation to obtain the hidden variable value X.
[0069] Next, in Figure 3 step S120, the model predictive controller 120 outputs the controller predictive control parameter value Um. The model predictive controller 120 includes a prediction model 121 and an optimization searcher 122. The prediction model 121 is used to obtain the target parameter value Yi based on various control parameter values Ui and the hidden variable value X. The optimization searcher 122 is used to find the optimal controller predictive control parameter value Um from these control parameter values Ui, so that under the condition of the hidden variable value X, the most compliant predictive target parameter value Yp can be predicted.
[0070] For device 800, the model predictive controller 120 has been offline trained before going online. Generally speaking, after offline training, the controller prediction control parameter value Um deduced by the model predictive controller 120 based on the hidden variable value X can stably make device 800 (such as the aforementioned tank device 900) output a prediction target parameter value Yp that meets the standard. However, after device 800 has been operating for some time, it may be affected by various uncertain interference factors, resulting in the model predictive controller 120 being unable to accurately deduce a suitable controller prediction control parameter value Um any longer.
[0071] To maintain the prediction / control accuracy of the model predictive controller 120, the model predictive controller 120 can be retrained offline, but this will cause the process to stop, which does not meet the operating efficiency. In this embodiment, online compensation can be performed through the following steps without stopping the process, meeting the actual needs of the industry.
[0072] Next, in Figure 3 step S130 of Figure 2 as shown in
[0073] Compensation judgment unit 130 performs a judgment procedure based on the prediction target parameter value Yp to judge whether to enable the reinforcement learning model 140. In one embodiment, the compensation judgment unit 130 judges whether to enable the reinforcement learning model 140 based on, for example, the mean squared error (MSE) between the prediction target parameter value Yp and the actual target parameter value. That is, once it is found that the prediction target parameter value Yp deviates from the standard, the controller prediction control parameter value Um output by the model predictive controller 120 needs to be compensated. Figure 3 Then, in Figure 2 step S140 of
[0074] as shown in Figure 3 the reinforcement learning model 140 outputs a learned prediction control parameter value Ur.
[0075] Next, in Figure 2As shown, after learning the predictive control parameter value Ur and compensating the controller's predictive control parameter value Um through a compensation program, the calculation result is first input to the range limiting unit 160 to limit the output compensated predictive control parameter value Umr within a predetermined range (i.e., between the maximum upper limit value and the minimum lower limit value). For example, the output compensated predictive control parameter value is limited to be between 0 and 1, but this is not limiting.
[0076] In other embodiments, in addition to using summation operations, the compensation unit 150 may also use multiplication operations, integration operations, or matrix operations according to requirements, which are not limited herein.
[0077] The above steps S140 and S150 need to be executed only when the compensation judgment unit 130 determines that the reinforcement learning model 140 needs to be enabled. Additionally, if in Figure 3 step S130, the compensation judgment unit 130 determines that the reinforcement learning model 140 does not need to be enabled, then the above steps S140 and S150 are skipped and directly proceed to step S170.
[0078] In Figure 3 after step S150, it proceeds to step S170. As Figure 2 shown, the range limiting unit 160 outputs the compensated predictive control parameter value Umr to the device 800 to output the predicted target parameter value Yp.
[0079] Then, it returns to step S110 and repeats the above steps to continuously control and predict the parameters of the device 800.
[0080] Through the above predictive method of the reinforcement model predictive control system 100, when the model predictive controller 120 faces a situation with uncertain problems, it can immediately compensate for the decision-making deviation online through the reinforcement learning model 140 to improve the prediction accuracy and stabilize the control of the device 800.
[0081] As Figure 2 shown, both the compensated predictive control parameter value Umr and the predicted target parameter value Yp are transmitted to the storage unit 180 for storage. Once the compensated predictive control parameter value Umr and the predicted target parameter value Yp accumulate to a predetermined quantity, the online update program of the reinforcement learning model 140 can be initiated.
[0082] Please refer to Figure 4 for an example to illustrate the online update program of the reinforcement learning model 140. In Figure 4In step S270, the storage unit 180 continuously stores the compensated predictive control parameter value Umr and the predictive target parameter value Yp. In another embodiment, the compensated predictive control parameter value Umr and the predictive target parameter value Yp may also be stored in a register (not schematically shown) connected to the reinforcement learning model 140.
[0083] Next, in Figure 4 step S280, the storage unit 180 determines whether the compensated predictive control parameter value Umr and the predictive target parameter value Yp have accumulated a predetermined number. If the compensated predictive control parameter value Umr and the predictive target parameter value Yp have not accumulated to the predetermined number, the process returns to step S270. Conversely, if the compensated predictive control parameter value Umr and the predictive target parameter value Yp have accumulated the predetermined number, the process proceeds to step S290.
[0084] In step S290, an online update procedure is performed on the reinforcement learning model 140 based on these compensated predictive control parameter values Umr and these predictive target parameter values Yp. In one embodiment, the reinforcement learning model 140 may adopt the Deep Deterministic Policy Gradient Algorithm (DDPG) or the Deep Q Network Algorithm (DQN) according to the continuity of the control actions of the device 800. For example, if the control actions of the device 800 are continuous, the Deep Deterministic Policy Gradient Algorithm may be adopted; if the control actions of the device 800 are discrete, the Deep Q Network Algorithm may be adopted.
[0085] Step S290 is an online update procedure. During the operation of step S290, the reinforcement model predictive control system 100 does not need to stop operating and there will be no problem of control delay.
[0086] After the reinforcement learning model 140 completes the online update procedure, the temporarily stored compensated predictive control parameter value Umr and the predictive target parameter value Yp can be cleared. After the newly added compensated predictive control parameter value Umr and the predictive target parameter value Yp accumulate to the predetermined number again, the online update procedure of step S290 will be started again.
[0087] Through the online update procedure of the above reinforcement learning model 140, the reinforcement learning model 140 can be adaptively updated to maintain the stability of online operation.
[0088] In addition, before the reinforcement learning model 140 goes online, it must undergo offline training to correctly compensate the model predictive controller 120. However, directly performing offline training on the reinforcement learning model 140 is a rather difficult task. The present invention proposes a three-stage training procedure in the following description to gradually obtain the parameters of the reinforcement learning model 140, thereby significantly shortening the training time.
[0089] Please refer to Figures 5A - 5B , Figure 5A which schematically shows the first-stage training procedure PS1 of the offline training method for the reinforcement learning model 140. Figure 5B The following is an example to illustrate the first-stage training procedure PS1 of the offline training method for the reinforcement learning model 140. The first-stage training procedure PS1 includes steps S530, S560, S570, S580, and S590. In Figure 5A step S530, as Figure 5B shown, a reference controller 300 (for example, a Proportional–Integral–Derivative (PID) controller) outputs a plurality of reference predictive control parameter values Ud.
[0090] Next, in Figure 5A step S560, as Figure 5B shown, the reference controller 300 provides these reference predictive control parameter values Ud to the device 800 to obtain a plurality of predicted target parameter values Yp.
[0091] Then, in Figure 5A step S570, as Figure 5B shown, the storage unit 180 temporarily stores these reference predictive control parameter values Ud and the corresponding predicted target parameter values Yp.
[0092] Next, in Figure 5A step S580, as Figure 5B shown, the storage unit 180 determines whether the temporarily stored reference predictive control parameter values Ud and the corresponding predicted target parameter values Yp have accumulated to the expected quantity. If the reference predictive control parameter values Ud and the corresponding predicted target parameter values Yp have not accumulated to the expected quantity, return to step S530. Conversely, if the temporarily stored reference predictive control parameter values Ud and the corresponding predicted target parameter values Yp have accumulated to the expected quantity, enter step S590.
[0093] Then, in Figure 5A step S590, as Figure 5BAs shown, the reinforcement learning model 140 is trained with these reference prediction control parameter values Ud and these predicted target parameter values Yp. In the first-stage training program PS1, in step S590, the reinforcement learning model 140 is trained by supervised learning.
[0094] In the first-stage training program PS1, the reference controller 300 adopted is a single-loop controller without a feedback mechanism, which can only achieve basic parameter setting and control. However, these reference prediction control parameter values Ud and these predicted target parameter values Yp can enable the reinforcement learning model 140 to quickly converge to a prototype.
[0095] Please refer to Figures 6A - 6B , Figure 6A which schematically shows the second-stage training program PS2 of the offline training method of the reinforcement learning model 140. Figure 6B An example illustrates the second-stage training program PS2 of the offline training method of the reinforcement learning model 140. The second-stage training program PS2 includes steps S610, S620, S630, S640, S660, S670, S680, S690. In Figure 6A step S610 of Figure 6B as shown, the detector 110 outputs the hidden variable value X.
[0096] Next, in Figure 6A step S620 of Figure 6B as shown, the model predictive controller 120 outputs the controller predictive control parameter value Um.
[0097] In Figure 6A step S630 of
[0098] the reference controller 300 (such as a PID controller) outputs a plurality of reference prediction control parameter values Ud. The above steps S620 and S630 can be executed synchronously. Figure 6A Then, in Figure 6B step S640 of
[0099] as shown, the compensation unit 150 compensates each controller predictive control parameter value Um with each reference prediction control parameter value Ud to obtain a plurality of corrected prediction control parameter values Umd. Figure 6A Next, in Figure 6B step S660 of
[0100] as shown, the compensation unit 150 provides these corrected prediction control parameter values Umd to the device 800 to obtain a plurality of predicted target parameter values Yp. Figure 6A Then, in Figure 6BAs shown, the storage unit 180 temporarily stores these corrected predictive control parameter values Umd and the corresponding predictive target parameter values Yp.
[0101] Next, in Figure 6A step S680, as Figure 6B shown, the storage unit 180 determines whether the temporarily stored corrected predictive control parameter values Umd and the corresponding predictive target parameter values Yp have accumulated to the predicted quantity. If the corrected predictive control parameter values Umd and the corresponding predictive target parameter values Yp have not accumulated to the predicted quantity, it returns to steps S610 and S630 again. Conversely, if the temporarily stored corrected predictive control parameter values Umd and the corresponding predictive target parameter values Yp have accumulated to the predicted quantity, it proceeds to step S690.
[0102] Then, in Figure 6A step S690, as Figure 6B shown, the reinforcement learning model 140 is trained using these corrected predictive control parameter values Umd and these predictive target parameter values Yp.
[0103] In the second-stage training program PS2, the compensation for the model predictive controller 120 is considered, enabling the reinforcement learning model 140 to converge towards the compensation mechanism.
[0104] Please refer to Figures 7A - 7B , Figure 7A which schematically shows the third-stage training program PS3 of the offline training method for the reinforcement learning model 140, Figure 7B illustrating the third-stage training program PS3 of the offline training method for the reinforcement learning model 140 by way of example. The third-stage training program PS3 includes steps S710, S720, S730, S740, S760, S770, S780, S790. In Figure 7A step S710, as Figure 7B shown, the detector 110 outputs the hidden variable value X.
[0105] Next, in Figure 7A step S720, as Figure 7B shown, the model predictive controller 120 outputs the controller predictive control parameter value Um.
[0106] In Figure 7A step S730, the reinforcement learning model 140 outputs multiple learning predictive control parameter values Ur. The above steps S720 and S730 can be executed synchronously.
[0107] Then, in Figure 7A step S740, as Figure 7BAs shown, the compensation unit 150 compensates each controller predictive control parameter value Um with respective learning predictive control parameter values Ur to obtain a plurality of compensated predictive control parameter values Umr.
[0108] Next, in Figure 7A step S760, as Figure 7B shown, the compensation unit 150 provides these compensated predictive control parameter values Umr to the device 800 to obtain a plurality of predictive target parameter values Yp.
[0109] Then, in Figure 7A step S770, as Figure 7B shown, the storage unit 180 temporarily stores these compensated predictive control parameter values Umr and the corresponding predictive target parameter values Yp.
[0110] Next, in Figure 7A step S780, as Figure 7B shown, the storage unit 180 determines whether the temporarily stored compensated predictive control parameter values Umr and the corresponding predictive target parameter values Yp have accumulated to a predicted quantity. If the compensated predictive control parameter values Umr and the corresponding predictive target parameter values Yp have not accumulated to the predicted quantity, it returns to steps S710 and S730. Conversely, if the temporarily stored compensated predictive control parameter values Umr and the corresponding predictive target parameter values Yp have accumulated to the predicted quantity, it proceeds to step S790.
[0111] Then, in Figure 7A step S790, as Figure 7B shown, the reinforcement learning model 140 is trained with these compensated predictive control parameter values Umr and these predictive target parameter values Yp.
[0112] In the third-stage training program PS3, the compensation of the reinforcement learning model 140 for the model predictive controller 120 is directly considered, enabling the reinforcement learning model 140 to be trained.
[0113] According to the above first-stage training program PS1, second-stage training program PS2, and third-stage training program PS3, the offline training of the reinforcement learning model 140 can gradually converge, and its convergence speed can be effectively improved.
[0114] Please refer to Figure 8 , which comparatively illustrates the control results of the reinforced model predictive control system 100 and the general model predictive control system for the tank level height y1 of the tank device 900. Curve C10 is the ideal curve of the tank level height, curve C11 is the tank level height curve of the general model predictive control system, and curve C12 is the tank level height curve of the reinforced model predictive control system 100. FromFigure 8 It can be seen that curve C12 is closer to curve C10. That is to say, the enhanced model predictive control system 100 can control the liquid level height y1 of the tank more accurately.
[0115] Please refer to Figure 9 , which compares and explains the control results of the enhanced model predictive control system 100 and the general model predictive control system for the tank temperature y2 of the tank equipment 900. Curve C20 is the ideal curve of the tank temperature, curve C21 is the tank temperature curve using the general model predictive control system, and curve C22 is the tank temperature curve using the enhanced model predictive control system 100. From Figure 9 It can be seen that curve C12 is as close to curve C10 as curve C11. That is to say, the enhanced model predictive control system 100 can also control the tank temperature y2 more accurately.
[0116] According to the above embodiments, when the model predictive controller 120 faces a situation with uncertain problems, it can immediately compensate for the decision-making deviation online through the reinforcement learning model 140 to improve the prediction accuracy and stabilize the control of the device 800, without the need to retrain the model predictive controller 120 offline, thus avoiding the occurrence of control delay.
[0117] In summary, although the present invention has been described above with embodiments, it is not intended to limit the present invention. Those with ordinary knowledge in the technical field to which the present invention pertains can make various changes and modifications without departing from the spirit and scope of the present invention. Therefore, the protection scope of the present invention shall be subject to the protection scope defined in the scope of this patent application.
Claims
1. A reinforced model predictive control system for a device, and comprises: a model predictive controller for outputting a controller predictive control parameter value; a reinforcement learning model for outputting a learning predictive control parameter value; and a compensation unit that compensates the controller predictive control parameter value with the learning predictive control parameter value to obtain a compensated predictive control parameter value, and the compensated predictive control parameter value is provided to the device to output a predictive target parameter value.
2. The reinforced model predictive control system according to claim 1, wherein the compensation unit performs an addition operation on the learning predictive control parameter value and the controller predictive control parameter value respectively at a first ratio and a second ratio to obtain the compensated predictive control parameter value.
3. The reinforced model predictive control system according to claim 2, wherein the sum of the first ratio and the second ratio is 1, and the first ratio is greater than 0.
2.
4. The reinforced model predictive control system according to claim 1, further comprises: a storage unit for storing the compensated predictive control parameter value and the predictive target parameter value to provide the reinforcement learning model to execute an immediate update program.
5. The reinforced model predictive control system according to claim 1, further comprises: a compensation judgment unit for performing a judgment program based on the predictive target parameter value to judge whether to enable the reinforcement learning model.
6. The reinforced model predictive control system according to claim 5, wherein the compensation judgment unit judges whether to enable the reinforcement learning model based on the mean square error between the predictive target parameter value and the actual target parameter value.
7. The reinforced model predictive control system according to claim 1, further comprises: a range limiting unit for limiting the compensated predictive control parameter value within a predetermined range.
8. A prediction method for a reinforced model predictive control system for a device, and comprises: outputting a controller predictive control parameter value with a model predictive controller; outputting a learning predictive control parameter value with a reinforcement learning model; performing a compensation program based on the learning predictive control parameter value and the controller predictive control parameter value to obtain a compensated predictive control parameter value; and providing the compensated predictive control parameter value to the device to output a predictive target parameter value.
9. The prediction method for the reinforced model predictive control system according to claim 8, wherein the compensation program performs an addition operation on the learning predictive control parameter value and the controller predictive control parameter value respectively at a first ratio and a second ratio to obtain the compensated predictive control parameter value.
10. The prediction method for the reinforced model predictive control system according to claim 9, wherein the sum of the first ratio and the second ratio is 1, and the first ratio is greater than 0.
2.
11. The prediction method for the reinforced model predictive control system according to claim 9, wherein the sum of the first ratio and the second ratio is 1, the first ratio is 0.5, and the second ratio is 0.
5.
12. The prediction method for the reinforced model predictive control system according to claim 9, further comprises: performing a judgment program based on the predictive target parameter value to judge whether to enable the reinforcement learning model.
13. The prediction method of the enhanced model predictive control system according to claim 12, wherein the determination program determines whether to enable the reinforcement learning model based on the mean square error between the predicted target parameter value and the actual target parameter value.
14. The prediction method of the enhanced model predictive control system according to claim 8, wherein the compensated predicted control parameter value is limited within a predetermined range.
15. The prediction method of the enhanced model predictive control system according to claim 8, further comprising: Performing an immediate update program on the reinforcement learning model.
16. The prediction method of the enhanced model predictive control system according to claim 8, wherein the reinforcement learning model adopts a deep deterministic policy gradient algorithm or a deep Q-network algorithm.
17. A training method of an enhanced model predictive control system, the enhanced model predictive control system compensates a model predictive controller through a reinforcement learning model and is used for a device, the training method comprising: Training the reinforcement learning model with a first-stage training program, the first-stage training program comprising: Outputting a plurality of reference predicted control parameter values with a reference controller; Providing these reference predicted control parameter values to the device to obtain a plurality of predicted target parameter values; and Training the reinforcement learning model through these reference predicted control parameter values and these predicted target parameter values; Training the reinforcement learning model with a second-stage training program, the second-stage training program comprising: Outputting a plurality of controller predicted control parameter values with the model predictive controller; Outputting these reference predicted control parameter values with the reference controller; Compensating each of the controller predicted control parameter values with each of the reference predicted control parameter values to obtain a plurality of corrected predicted control parameter values; Providing these corrected predicted control parameter values to the device to obtain these predicted target parameter values; Training the reinforcement learning model through these corrected predicted control parameter values and these predicted target parameter values; and Training the reinforcement learning model with a third-stage training program, the third-stage training program comprising: Outputting these controller predicted control parameter values with the model predictive controller; Outputting a plurality of learned predicted control parameter values with the reinforcement learning model; Performing a compensation program based on these learned predicted control parameter values and these controller predicted control parameter values to obtain a plurality of compensated predicted control parameter values; Providing these compensated predicted control parameter values to the device to obtain these predicted target parameter values; and Training the reinforcement learning model through these compensated predicted control parameter values and these predicted target parameter values.
18. The training method of the enhanced model predictive control system according to claim 17, wherein in the first-stage training program, the reinforcement learning model is trained by supervised learning.