Efficient MPC decision enhancement method based on teaching learning
By decomposing the optimal sub-problem, designing Lagrangian functions and constructing inverse KKT conditions, the problems of insufficient decision constraint identification and low learning efficiency of MPC when learning decision behavior are solved, and the efficient decision control performance improvement and infinite time domain performance compensation of MPC are achieved.
Patent Information
- Application Number
- CN202510622637.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-15
- Publication Date
- 2025-06-13
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
When existing MPC controllers learn the decision-making behavior of excellent operation experts, there are problems such as insufficient identification of decision constraints, low learning efficiency, and difficulty in compensating performance in infinite time domains.
By collecting dynamic data during the teaching process of human experts, using the Belmann optimality principle to decompose the optimal sub-problem, design the Lagrangian function, obtain the first-order necessity conditions, establish candidate decision constraints, and construct inverse KKT conditions and nonlinear planning propositions to realize the efficient decision constraints of MPC and the learning of cost function weight parameters.
It improves the learning efficiency and decision-making control performance of MPC, and can compensate for infinite time domain performance in the finite time domain, so that MPC has the ability to learn from manual teaching and continuously improve decision-making performance.
Smart Images

Figure CN120146226A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of improving the performance of MPC controllers in human-machine interaction, and in particular, to an efficient MPC decision enhancement method based on teaching learning. Background Art
[0002] MPC (Model Predictive Control) is based on an internal process prediction model and can obtain an efficient control action by solving a non-linear optimization problem with decision constraints and a quadratic cost function. The design of a suitable control law (cost function and decision constraints) has a crucial impact on the closed-loop control performance of MPC, which involves the design of multiple weight parameters in the cost function, multiple decision variables, and their increment constraints, and requires combining specific decision task scenarios and rich decision-making experience. In typical fields such as continuous process industries and discrete manufacturing industries, excellent operation experts usually have rich mechanism knowledge and practical experience, and their decision-making behaviors are rational, far-sighted, and clear in goal. In some task scenarios, the decisions of these operation experts can even be considered optimal or relatively optimal. However, the experiential knowledge, decision-making intentions, and decision-making patterns possessed by excellent operation experts are all implicit and cannot be directly used in the design process of MPC machine algorithms. Therefore, it is necessary to solve the problems of how to extract appropriate control law parameters from the explicit decision data of excellent operators and how to enable MPC to have the ability to learn from excellent human teaching data and for the machine to have the ability to learn from people.
[0003] Teaching learning, also known as expert teaching, inverse optimal control, imitation learning, target learning, etc., its main purpose is to enable a machine (machine algorithm) to learn the teaching decision-making behaviors of excellent experts, so that the machine has the ability to implement optimal or approximately optimal control strategies possessed by the experts. However, the methods in existing research and practical applications still have three aspects of problems: (1) First, most teaching learning methods only consider scenarios without decision constraints or with known constraints. Decision constraints, as key indicators reflecting operation behaviors, effectively identifying constraints will help to more comprehensively characterize the decision-making behaviors of excellent experts into machine algorithms; (2) Second, currently most methods require teaching data of the complete decision-making process. For task scenarios with long decision-making cycles and high decision-making space dimensions, these methods are less efficient in the process of learning control laws; (3) Third, most teaching learning methods are carried out within the framework of optimal control. For the rolling optimization strategy of MPC, it is unrealistic to solve an infinite-time open-loop optimal control problem at each step, and only the corresponding finite-time problem can be considered, which inevitably leads to a loss of control performance and loses the guarantee of algorithm stability.
[0004] Therefore, how to study fast, efficient, and accurate control law parameter and decision constraint learning methods under the framework of MPC by combining human teaching data, while compensating for the infinite-horizon performance of MPC, is an important technical problem that needs to be overcome for autonomously enhancing the decision-making ability of MPC and forming a human-machine interaction system where machines learn from humans. Summary of the Invention
[0005] The object of the present invention is to provide an efficient MPC decision enhancement method based on teaching learning for the deficiencies of the existing technology. This method can not only more efficiently learn the effective decision constraints and cost function weight parameters of MPC simultaneously, but also compensate for the infinite-horizon performance of MPC, improve the overall control effect of MPC, and enable MPC to have the ability to learn from human teaching and continuously improve its decision-making performance.
[0006] The object of the present invention is achieved through the following technical solutions: providing an efficient MPC decision enhancement method based on teaching learning, including the following steps: (1) Collect multiple sets of dynamic data during the teaching process of human experts, and form an observation data set with completely aligned time steps after data preprocessing; (2) Use the Bellman optimality principle to decompose the observation data set to obtain the optimal sub-problems of the optimal control problem during the teaching process of human experts; (3) According to the optimal sub-problems, design multiple sets of Lagrange multipliers to obtain the corresponding Lagrangian functions; (4) Obtain the first-order necessary conditions of the Lagrangian function; (5) Establish candidate decision constraints and revise the original complementary slackness conditions; (6) Construct equality and inequality equations to form the inverse KKT conditions for solving the cost function parameters and effective decision constraints; (7) Relax the inverse KKT conditions and normalize the weight parameters to obtain a non-linear constraint optimization proposition that can simultaneously estimate the weight parameters and decision constraints; (8) Construct a teaching error state set based on the observed expert teaching data, that is, the terminal constraint set; (9) Construct the value function at time, and use the error state to construct its terminal cost function within the terminal constraint set; (10) Based on the constructed terminal constraint set and terminal cost function, construct a non-linear programming proposition containing 0-1 integers to compensate for the infinite-horizon performance of the finite-horizon MPC.
[0007] Specifically, the multiple sets of dynamic data in the human expert teaching process in step (1) include data of operation variables, controlled variables, operation variable increments, and state variables. The data preprocessing is to completely align the time steps of all variable data of operation variables, controlled variables, operation variable increments, and state variables, so as to obtain the preprocessed observation data set.
[0008] Specifically, in step (3), according to the optimal sub-problem, multiple sets of Lagrange multipliers are designed to obtain the corresponding Lagrangian function , where the Lagrange multipliers corresponding to the inequality constraints are greater than or equal to 0, and the upper and lower limits of the operation variables and operation variable increments in the Lagrangian function are parameters to be estimated; the expression is as follows: ; where is the starting time corresponding to the optimal sub-problem, represents the time index, is the length of the teaching data corresponding to the optimal sub-problem, , , and are the state variable, controlled variable, operation variable, and operation increment variable of the process respectively, and and , and , and , and are the lower and upper limits of the state variable, controlled variable, operation variable, and operation increment variable respectively; , and represent the dimensions of the state variable, controlled variable, and operation variable respectively, represents the set of real numbers; represents the Lagrange multiplier vector of the state model equation, and represent the Lagrange multiplier vectors of the initial state and the termination state respectively, represents the Lagrange multiplier vector of the output model equation, represents the Lagrange multiplier vector of the control increment equation, and represent the Lagrange multiplier vectors of the lower limit constraint and the upper limit constraint of the state variable respectively, and represent the Lagrange multiplier vectors of the lower limit constraint and the upper limit constraint of the output variable respectively, and represent the Lagrange multiplier vectors of the lower limit constraint and the upper limit constraint of the control variable respectively, and represent the Lagrange multiplier vectors for the lower and upper bound constraints of the control variable increments, respectively; represents the predicted value of the controlled variable at time ; and represent the predicted values of the manipulated variable and its increment variable at time ; represents the predicted value of the manipulated variable at time ; represents the predicted value of the state variable at time ; represents the predicted value of the state variable at time ; represents the value of the state variable at time ; represents the initial value of the state variable at time ; represents the output model expression of the process; represents the non - linear performance cost function, expressed as follows: ; where and are the weight parameters in the cost function, represents the tracking target; represents the state - space model expression of the process.
[0009] Furthermore, step (4) is specifically: performing a partial derivative operation on the Lagrangian function to obtain the first - order necessary conditions of the Lagrangian function, namely, the stationarity condition, the feasibility condition of the optimal sub - problem, the feasibility condition of the dual problem, the complementary slackness condition, and the primal complementary slackness condition.
[0010] Specifically, the candidate decision constraints constructed in step (5) include: the maximum value of the observed manipulated variable, the minimum value of the manipulated variable, the maximum value of the manipulated variable increment, and the minimum value of the manipulated variable increment.
[0011] Specifically, the inverse KKT conditions in step (6) are composed of a system of equations and inequalities, and each weight parameter of the model predictive control (MPC) is greater than or equal to 0.
[0012] Specifically, in step (7), the weight parameters are normalized such that the sum of all weight parameters is 1 to obtain a unique solution, and the constructed non-linear constrained optimization problem is a quadratic programming problem.
[0013] Specifically, the set of teaching error states in step (8) is calculated from the difference between the controlled variable and its target, and the set of teaching error states is a control invariant set.
[0014] Further, the terminal cost function in step (9) is: when there is an error state belonging to the terminal constraint set, the value function of the teaching error state corresponding to the error state in the terminal constraint set must be the minimum value function in the terminal constraint set.
[0015] Specifically, the sum of all integer variables in the non-linear programming problem containing 0-1 integers in step (10) is 1, and only the first beat control action of the solved control action will be executed. Then, this problem will be re-solved according to the latest feedback data in each control cycle.
[0016] The beneficial effects of the present invention are as follows: The efficient MPC decision enhancement method based on teaching learning of the present invention enables the MPC machine algorithm to have the ability to learn from excellent human teaching, efficiently improving the decision control performance. It can not only learn the optimal cost function weight parameters and effective decision constraints by only using partial teaching data, greatly improving the learning efficiency of MPC and reducing the computational burden. Further, the method proposed by the present invention can also compensate for the infinite-time domain performance of finite-time domain MPC according to the observed excellent human teaching data. This method is of great significance for the design of MPC algorithms and the improvement of MPC decision performance. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Figure 1 is the logical framework diagram of this method. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0018] The present invention will be described in detail below with reference to the drawings.
[0019] The present invention discloses an efficient MPC decision enhancement method based on teaching learning, as Figure 1 shown, including the following steps: Step 1: Construct a data set; Collect multiple groups of dynamic data during the teaching process of human experts, and through data preprocessing, align the time steps of all states, operations, operation increments, and controlled variables to form a complete set of expert teaching data : ; where respectively represent the teaching data corresponding to the operation increment, operation, state, and controlled variable, Indicates the total amount of data collected.
[0020] Step 2: Construct the optimal sub-problem ; Decompose the observed data set according to the form of the original optimal control problem, using the Bellman optimality principle to obtain the optimal sub-problem of the optimal control problem during the human expert teaching process (and ensure that the observed data set corresponding to the solution of the optimal sub-problem is a subset of the observed data of the solution of the original problem ): ; ; where is the starting time corresponding to the optimal sub-problem, represents the time index, is the length of the teaching data corresponding to the optimal sub-problem, is the solution of the operation increment and operation variable corresponding to the optimal sub-problem, , , and are the state variable, controlled variable, operation variable, and operation increment variable of the process respectively, and , and , and , and are the lower and upper limits of the corresponding variables; , and represent the dimensions of the state variable, controlled variable, and operation variable respectively, represents the set of real numbers. represents the predicted value of the controlled variable at time for time and represent the predicted values of the operation variable and its increment variable at time for time represents the predicted value of the operation variable at time for time represents the predicted value of the state variable at time for time represents the predicted value of the state variable at time for time represents The value of the state variable at a moment denotes the initial value of the state variable at a moment denotes the value of the state variable at a moment; the constraints in the last two lines of the above formula indicate that the starting and ending state of this sub - problem are both states in the teaching . Denotes the output model expression of the process; is the operation increment variable, is the tracking error, is the tracking target of the controlled variable, is the non - linear performance cost function, in the following quadratic form: ; where and are the weight parameters in the cost function, is the transpose symbol.
[0021] Step 3: Construct the Lagrangian function corresponding to the optimal sub - problem ; According to the optimal sub - problem , design multiple groups of Lagrange multipliers and construct the Lagrangian function : ; where, denotes the Lagrange multiplier vector of the state model equation, and respectively denote the Lagrange multiplier vectors of the initial state and the termination state, denotes the Lagrange multiplier vector of the output model equation, denotes the Lagrange multiplier vector of the control increment equation, and respectively denote the Lagrange multiplier vectors of the lower bound constraint and the upper bound constraint of the state variable, and respectively denote the Lagrange multiplier vectors of the lower bound constraint and the upper bound constraint of the output variable, and respectively denote the Lagrange multiplier vectors of the lower bound constraint and the upper bound constraint of the control variable, and respectively denote the Lagrange multiplier vectors of the lower bound constraint and the upper bound constraint of the control variable increment. Denotes the non - linear performance cost function, expressed as follows: ; where and are the weight parameters in the cost function, Denote the tracking target; Denote the state - space model expression of the process object.
[0022] Step 4: Obtain the first - order necessary condition of the Lagrangian function ; Perform partial - derivative operations on the Lagrangian function to obtain the Lagrangian stationarity condition, the optimal sub - problem feasibility condition, the dual - problem feasibility condition, the complementary slackness condition, and the primal complementary slackness condition. Among them, the Lagrangian stationarity condition is: ; where Denote the control - variable increment matrix, Denote the control - variable matrix, Denote the state - variable matrix, Denote the output - variable matrix; is the partial - derivative operation symbol.
[0023] Step 5: Construct candidate decision constraints and revise the original complementary slackness condition.
[0024] Utilize the maximum and minimum values of all observed manipulated variables and manipulated - variable increments to construct candidate decision constraints , , and , and revise the original complementary slackness condition: ; where and respectively denote the sets of control variables and enhanced control variables, and respectively denote the maximum and minimum values of the observed data of the corresponding variables; Denote taking the maximum value in the set, Denote taking the minimum value in the set, Denote taking the maximum value in the set, Denote taking the minimum value in the set. Further, obtain the revised complementary slackness condition: ; Step 6: Construct equality and inequality equations to form the inverse KKT conditions for solving the cost - function parameters and effective decision constraints.
[0025] Substitute the observed solution set corresponding to the optimal sub - problem into the following inverse KKT conditions: ; where and are the dimensions of the controlled and manipulated variables respectively, and are the weight parameters of the error state variable and the control increment variable in the cost function respectively; and are the values of the manipulated variable and its increment in the observation solution set, and are the values of the state variable and the controlled variable in the observation solution set.
[0026] Relax the inverse KKT conditions in step six and add the equality constraint that the sum of all weight parameters in the cost function is 1 to obtain the nonlinear constrained optimization proposition : ; where represents the square of the norm, represent the estimated values of the weight parameters of the error state variable and the control increment variable in the cost function respectively, represents the Lagrange multiplier corresponding to the inequality equation.
[0027] Step eight: Construct the terminal constraint set of the teaching error state; Use the expert teaching data set to calculate the difference between the controlled variable and its target to obtain the teaching error state set : ; where is based on the difference between the controlled variable and its target and satisfies .
[0028] Construct the value function at time , and use the error state to construct its terminal cost function within the terminal constraint set : ; where, ; where is an integer variable that can only take 0 or 1, is the teaching error state and
[0029]
[0029] Step ten: Construct the MPC decision enhancement optimization proposition; Use the terminal constraint set and the terminal cost function constructed in steps eight and nine to construct a nonlinear programming proposition containing 0-1 integers and add an equality constraint that the sum of all integer variables is 1: ; where is the length of the prediction horizon, is the estimated performance cost function, is an integer variable that can only take 0 or 1, is the teaching error state; is the operation increment corresponding value function; is the optimized operation variable and its increment sequence, and only the first beat is implemented. In the next control cycle, the above optimization proposition will be solved again with the latest feedback information.
[0030] The principle of this method is simple, the steps are clear, it has strong flexibility, is easy to be programmed and implemented on a computer, can efficiently learn the effective decision constraints and cost function weight parameters of MPC at the same time, can also compensate for the infinite-horizon performance of MPC, improve the overall control effect of MPC, and enable MPC to have the ability to learn from manual teaching and continuously improve decision-making performance.
[0031] The above-described implementation steps are only a preferred solution of the present invention, but it is not intended to limit the present invention. Those of ordinary skill in the relevant technical field can also make various changes and modifications without departing from the spirit and scope of the present invention. Therefore, all technical solutions obtained by adopting the means of equivalent replacement or equivalent transformation fall within the protection scope of the present invention.
Claims
1. An efficient MPC decision enhancement method based on teaching learning, characterized in that: The following steps are involved: (1) Collect multiple sets of dynamic data from the teaching process of human experts, and after data preprocessing, form a set of observation data with completely aligned time steps; (2) Using the Bellman optimality principle to decompose the observed data set, we can obtain the optimal sub-problem of the optimal control problem during the human expert teaching process; (3) According to the optimal subproblem, design multiple sets of Lagrangian multipliers and obtain the corresponding Lagrangian function; (4) Obtaining the first-order necessary conditions for the Lagrangian function; (5) Establish constraints on the operational variables and their incremental candidate decisions and revise the original complementary relaxation conditions; (6) Construct equality and inequality equations to form inverse KKT conditions for solving cost function parameters and effective decision constraints; (7) Relax the inverse KKT condition and normalize the weight parameters to obtain a nonlinear constrained optimization problem that can simultaneously estimate the weight parameters and decision constraints; (8) Constructing a teaching error state set, i.e., a terminal constraint set, based on the observed expert teaching data; (9) Construction The value function at the moment, and use the error state to construct its terminal cost function within the terminal constraint set; (10) Based on the constructed terminal constraint set and terminal cost function, a nonlinear programming proposition containing 0-1 integers is constructed to compensate the infinite-horizon performance of the finite-horizon MPC.
2. The efficient MPC decision enhancement method based on teaching learning according to claim 1 is characterized in that: The multiple sets of dynamic data in the human expert teaching process in step (1) include data of operating variables, controlled variables, operating variable increments and state variables. The data preprocessing is to completely align the time steps of all variable data of operating variables, controlled variables, operating variable increments and state variables, thereby obtaining a preprocessed observation data set.
3. The efficient MPC decision enhancement method based on teaching learning according to claim 1 is characterized in that: In step (3), multiple groups of Lagrangian multipliers are designed according to the optimal subproblem to obtain the corresponding Lagrangian function , the Lagrange multiplier corresponding to the inequality constraint is greater than or equal to 0, and the upper and lower limits of the manipulated variables and manipulated variable increments in the Lagrangian function are the parameters to be estimated; the expression is as follows: ; in, is the starting time corresponding to the optimal subproblem, Represents the time index, is the length of the teaching data corresponding to the optimal subproblem, , , and are the state variables, controlled variables, manipulated variables, and manipulated increment variables of the process, and and , and , and , and They are the lower and upper limits of the state variable, controlled variable, manipulated variable, and manipulated increment variable respectively; , and Represent the dimensions of state variables, controlled variables, and manipulated variables, respectively. represents the set of real numbers; represents the Lagrange multiplier vector of the state model equation, and denote the Lagrange multiplier vectors of the initial and final states, respectively, represents the Lagrange multiplier vector of the output model equation, represents the Lagrange multiplier vector governing the incremental equations, and are the Lagrange multiplier vectors representing the lower and upper bound constraints on the state variables, respectively. and are the Lagrange multiplier vectors representing the lower and upper bound constraints on the output variables, respectively. and are the Lagrange multiplier vectors representing the lower and upper constraints on the control variables, respectively. and The Lagrange multiplier vectors representing the lower and upper constraints of the control variable increments; express Always The predicted value of the controlled variable at time , and express Always The predicted values of the moment-by-moment operational variables and their incremental variables, express Always The predicted value of the operational variable at time express Always The predicted value of the state variable at time express Always The predicted value of the state variable at time express The value of the state variable at time express The initial value of the state variable at time , express The value of the state variable at the moment; An output model expression representing the process; Represents a nonlinear performance cost function, expressed as follows: ; in and is the weight parameter in the cost function, Indicates the tracking target; Represent the state-space model expression of the process.
4. The efficient MPC decision enhancement method based on teaching learning according to claim 1 is characterized in that: The step (4) is specifically as follows: performing partial derivative operation on the Lagrangian function to obtain the first-order necessary conditions of the Lagrangian function, namely, the stationarity condition, the optimal subproblem feasibility condition, the dual problem feasibility condition, the complementary relaxation condition and the original complementary relaxation condition.
5. The efficient MPC decision enhancement method based on teaching learning according to claim 1 is characterized in that: The candidate decision constraints constructed in step (5) include: the maximum value of the observed operating variable, the minimum value of the operating variable, the maximum value of the operating variable increment, and the minimum value of the operating variable increment.
6. The efficient MPC decision enhancement method based on teaching learning according to claim 1 is characterized in that: The inverse KKT condition in step (6) is composed of a set of equations of equality and inequality, and each weight parameter of the model predictive control MPC is greater than or equal to 0.
7. The efficient MPC decision enhancement method based on teaching learning according to claim 1 is characterized in that: In the step (7), the weight parameters are normalized, which means that the sum of all weight parameters is 1, so as to obtain a unique solution, and the constructed nonlinear constrained optimization problem is a quadratic programming problem.
8. The efficient MPC decision enhancement method based on teaching learning according to claim 1 is characterized in that: The teaching error state set in step (8) is obtained by calculating the difference between the controlled variable and its target, and the teaching error state set is a control invariant set.
9. The efficient MPC decision enhancement method based on teaching learning according to claim 1 is characterized in that: When there is an error state that belongs to the terminal constraint set, the value function of the teaching error state corresponding to the error state in the terminal constraint set must be the minimum value function in the terminal constraint set.
10. The efficient MPC decision enhancement method based on teaching learning according to claim 1 is characterized in that: The sum of all integer variables in the nonlinear programming proposition containing 0-1 integers in step (10) is 1, and only the first-beat control action of the solved control action will be executed, and then the proposition will be re-solved according to the latest feedback data in each control cycle.