Strategy gradient embedded enhancement model predictive control method
By constructing a terminal cost function and a policy gradient feedback path, and combining the policy gradient and regularization term to optimize the objective function, the problem of unifying real-time performance and long-term performance of multi-degree-of-freedom robot control systems in complex environments is solved, thereby improving adaptability and robustness.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- BEIJING RES INST OF PRECISE MECHATRONICS CONTROLS
- Filing Date
- 2025-12-19
- Publication Date
- 2026-04-21
AI Technical Summary
In existing technologies, multi-degree-of-freedom robot control systems struggle to balance real-time performance, stability, and control accuracy in complex disturbance environments. Traditional MPC is limited by a finite prediction time domain and lacks long-term performance guidance. Furthermore, the combination of reinforcement learning and MPC fails to achieve deep synergy, resulting in a lack of global foresight in control strategies for complex tasks.
A terminal cost function is constructed, a terminal value function is obtained through parameterization, a terminal gradient feedback path is established, and a composite optimization objective function containing policy gradient and regularization term is constructed. Policy parameters are updated in real time to achieve dynamic optimization, forming a policy gradient embedded augmented model predictive control method.
It achieves a unified framework for short-term control and long-term performance, improves the system's adaptability and robustness in complex environments, ensures the real-time performance and long-term optimization consistency of the control strategy, and is suitable for real-time operation on embedded platforms.
Smart Images

Figure CN121900249A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to a policy gradient embedded enhancement model predictive control method, belonging to the field of robot motion control. Background Technology
[0002] Multi-degree-of-freedom (DOF) robots, due to their high degrees of freedom, strong state coupling, and complex constraints, have always been a key research focus and challenge in the field of robot control, especially floating base systems such as bipedal, quadrupedal robots, and attached robotic arms. Model Predictive Control (MPC) exhibits unique advantages in the motion control of such systems due to its ability to explicitly handle system constraints. However, traditional MPC is limited by a finite prediction time domain, making it difficult to effectively guide long-term performance, and the control framework cannot achieve policy self-learning through interactive experience. These problems make it difficult to balance real-time performance, stability, and control accuracy in systems operating under complex disturbances.
[0003] To enhance the capabilities of MPC in policy adaptation and long-term performance guidance, existing research has attempted to combine reinforcement learning (RL) mechanisms with MPC. The patent "A Trajectory Tracking Method for a Variable Load Mobile Robotic Arm Based on Reinforcement Learning and Model Predictive Control" (CN119748459B) (publication date: 2025.04.04) achieves weighted fusion of control inputs from MPC and RL through dynamic adjustment coefficients, improving system adaptability and robustness to some extent. However, its parallel architecture prevents deep coordination between MPC optimization and RL policy updates, making it difficult for the control law to simultaneously achieve both short-term accuracy and policy adaptability in complex tasks. The patent "Inverse Reinforcement Learning Utilizing Model Predictive Control" (CN112906882A) (publication date: 2021.06.04) further embeds reinforcement learning deeply into the MPC cost function, determining the MPC cost function through neural network feature recognition capabilities, thus enhancing task adaptability and policy generalization capabilities. However, this method still mainly relies on rolling optimization in the finite time domain and lacks a long-term performance guidance mechanism, which makes the control strategy lack global foresight in complex tasks and makes it difficult to achieve the unity of short-term control optimization and long-term decision optimization.
[0004] Therefore, there is an urgent need for a unified framework that can integrate short-term prediction optimization with long-term performance guidance. Summary of the Invention
[0005] The technical problem solved by this invention is: in order to address the existing problems of difficulty in unifying short-term optimization and long-term performance goals, as well as the insufficient adaptive capability of control strategies, a policy gradient embedded enhancement model predictive control method is proposed.
[0006] The present invention solves the above-mentioned technical problem through the following technical solution: A policy gradient embedded augmentation model predictive control method includes: Construct the terminal cost function, and use the terminal cost function to construct the terminal gradient feedback path; Based on the terminal cost function and the terminal gradient feedback path, a composite optimization objective function containing policy gradient and regularization term is constructed. A policy gradient embedded enhancement model is constructed for the structural transformation of the composite optimization objective function. The policy gradient embedded enhancement model is solved according to the modeling requirements to obtain the optimal control input sequence. Construct the gradient expression of the composite optimization objective function with respect to the policy parameters, and achieve dynamic optimization by updating the policy parameters in real time.
[0007] The method for constructing the terminal cost function is as follows: The control strategy generated through online optimization using model predictive control Construct the terminal cost function Used to describe the predicted terminal state The expected cumulative return in the future; Using parameterized form for the terminal cost function Perform modeling and obtain parameters The terminal value function of the control linear function approximation model
[0008] Control strategies are generated online through model predictive control optimization. That is, solving the finite-time domain Optimal control sequence within :
[0009]
[0010] In the formula, For stage cost function; For the terminal cost function; For the system dynamics model; This represents the current system state. From terminal status With control strategy The terminal cost function constituted for:
[0011] Using parameterized form for the terminal cost function To perform modeling, Terminal value function
[0012]
[0013] In the formula, Indicates by parameters The control terminal performance evaluation function; For a learnable weight parameter vector, The feature mapping function for the terminal state. It is preferable to use a polynomial expansion form for construction.
[0014] Using terminal value functions The method for constructing the terminal gradient feedback path is as follows: Calculate terminal value function Regarding terminal status The gradient; Construct the current performance gradient and apply it to the current control input using the chain rule. The transmission path.
[0015] The terminal value function Regarding terminal status The gradient calculation method is as follows: ; Using the chain rule, construct the performance gradient towards the current control input. transmission path ; The calculation method is as follows:
[0016]
[0017] In the formula, Let be the Jacobian matrix of the state-feature mapping; This represents the partial derivative of the current control input with respect to the predicted terminal state.
[0018] The method for constructing a composite optimization objective function is as follows: Construct the basic cost term, the policy gradient guiding term, and the policy gradient regularization term. Based on the basic cost term, the policy gradient guiding term, and the policy gradient regularization term, construct a composite objective function. ,in: Basic cost function for:
[0019] In the formula, For the first Step state error; This is the weighting matrix for the state error; To control the weighting matrix of the input, To predict the step size; Policy gradient guiding term for:
[0020] In the formula, Indicates the increment of the current control input; Policy gradient regularization term for:
[0021] Based on the basic cost item Policy gradient guidance term and gradient regularization term Constructing a composite objective function The method is as follows:
[0022]
[0023] In the formula, As a dynamic weighting factor, The regularization coefficient is . Basic proportional coefficient; is the numerical stability constant.
[0024] The method for constructing a policy gradient embedded augmentation model is as follows: Establish the system prediction equation represented by the discrete state-space model, express it as a function of the current state and the control input sequence according to the modeling requirements, construct the state recursion relationship, and determine the system prediction state sequence and the control input sequence to be optimized; Construct a standard quadratic form of the composite optimization objective function with respect to the control variables; Construct control constraints and state constraints; Based on the control constraints, state constraints, and the rewritten form of the composite optimization objective function, a constrained policy gradient embedded enhancement model is constructed according to the modeling requirements, and the optimal control input sequence is obtained.
[0025] The system prediction equation represented by the discrete state-space model is:
[0026] In the formula, , These are the discretized state transfer matrix and control input matrix, respectively; The state recursion, expressed as a function of the current state and the control input sequence, is as follows:
[0027] In the formula, For the present The system state at any given moment; Predict the state sequence for the system; The input sequence for the control to be optimized; , These are the expansion matrices of the system state and control input, respectively.
[0028] The composite optimization objective function Rewritten as about control variables The standard quadratic form is:
[0029] In the formula, This is the coefficient matrix of the quadratic terms; A vector of first-order terms; The control constraints are: For each control input Apply upper and lower limits to reflect the actuator's capability boundaries;
[0030] The state constraints are: Predicting the state of the system Limited to a safe and feasible area;
[0031] Rewrite the control constraints and state constraints in matrix inequality form:
[0032] In the formula, the matrix with vector Constructed based on the constraints.
[0033] The constrained policy gradient embedded enhancement model is a standard quadratic programming model with constraints, as follows:
[0034] The optimal control input sequence is obtained by solving the policy gradient embedded enhancement model. The optimal control input sequence is:
[0035] Among them, the first control input of the sequence is selected. After the actual control action is applied at the current moment, it is updated on a rolling basis at each sampling moment to form a closed-loop predictive control.
[0036] In the standard quadratic form of the composite optimization objective function, the coefficient matrix of the quadratic terms... Linear term vector Its composition includes basic cost items and policy gradient guidance term ,in:
[0037] ;
[0038]
[0039]
[0040]
[0041] .
[0042] The gradient expression of the composite optimization objective function with respect to the policy parameters is as follows:
[0043] in, .
[0044] Let the objective function be about The gradient is zero, thus obtaining the optimal analytical solution for the policy parameters. Specifically: when:
[0045] The optimal analytical solution for the policy parameters is obtained. for:
[0046] when When irreversibility is not possible, the Moore-Penrose generalized inverse is used as a substitute, and the gradient expression is as follows: ; By updating the optimal analytical solution in real time Enables dynamic optimization and updating of strategy parameters.
[0047] The advantages of this invention compared to the prior art are: (1) The present invention provides a strategy gradient embedded enhancement model predictive control method, which establishes the gradient relationship between control input and terminal performance, constructs gradient guidance terms, and explicitly embeds the optimization objective function to ensure that short-term control always moves toward long-term optimum, and realizes the collaborative optimization of process and terminal performance within a unified framework. It has the advantages of unified objectives, gradient guidance, and collaborative enhancement. (2) This invention designs an adaptive fusion mechanism for the objective function. By introducing a dynamic weighting factor based on real-time state feedback, the weight of the terminal gradient guidance term is autonomously adjusted to achieve an autonomous trade-off between short-term response and long-term guidance during the rolling optimization process. This mechanism has the advantages of strong environmental adaptability, excellent global-local balance, and high numerical robustness in the optimization process. (3) This invention embeds long-term performance goals into short-term rolling optimization through explicit feedback of terminal performance gradient, which overcomes the "short-sighted" problem of traditional MPC, so that the system can take into account long-term performance in each local optimization. At the same time, based on the dynamic weighting factor of state feedback, it can autonomously adjust the long and short time domain optimization weights, improve the adaptability and robustness of complex scenarios, and have the ability to adapt to dynamic environments. (4) This invention transforms the composite optimization objective function into a standard quadratic programming problem, taking into account both constraint handling and efficient solution, ensuring real-time performance and numerical stability, and ensuring efficient and stable optimization process. At the same time, it uses analytical methods to update policy parameters, avoiding iterative convergence problems, ensuring that the control optimization direction is consistent with the long-term performance goal, and achieving improved policy learning efficiency and consistency. (5) The present invention updates the terminal cost function online, and the system can improve the strategy autonomously based on real-time interaction, reduce the dependence on offline training and accurate models, realize online self-learning function, and based on an efficient numerical optimization framework, it does not require complex neural network calculations, making it more suitable for embedded platforms with limited computing resources to run in real time, and has the advantages of embedded deployment. Attached Figure Description
[0048] Figure 1 A flowchart of the control method provided by the present invention; Figure 2 The overall architecture diagram of the control system provided by this invention; Figure 3 The strategy gradient embedded enhancement MPC control block diagram provided by this invention; Figure 4 A comparison chart of the center-of-gravity displacement tracking performance of the bipedal robot provided by this invention; Figure 5 A comparison chart of the attitude angle tracking performance of the bipedal robot base provided by this invention. Detailed Implementation
[0049] A policy gradient embedded enhancement model predictive control method is proposed. It utilizes the gradient guidance of terminal performance to integrate long-term performance goals into short-term optimization decisions. Through iterative updates of control law and policy parameters, it simultaneously optimizes instantaneous control accuracy and long-term policy performance. Furthermore, based on real-time optimization using dynamic weighting factors, it enables the control law to self-adjust according to environmental changes. Thus, while ensuring the real-time performance of the system, it comprehensively improves its dynamic adaptive capability and global performance.
[0050] The policy gradient embedded augmentation model predictive control method includes the following steps: Construct the terminal cost function, and use the terminal cost function to construct the terminal gradient feedback path; Based on the terminal cost function and the terminal gradient feedback path, a composite optimization objective function containing policy gradient and regularization term is constructed. A policy gradient embedded enhancement model is constructed for the structural transformation of the composite optimization objective function. The policy gradient embedded enhancement model is solved according to the modeling requirements to obtain the optimal control input sequence. Construct the gradient expression of the composite optimization objective function with respect to the policy parameters, and achieve dynamic optimization by updating the policy parameters in real time.
[0051] The method for constructing terminal value functions is as follows: The control strategy generated through online optimization using model predictive control Construct the terminal cost function Used to describe the predicted terminal state The expected cumulative return in the future; Using parameterized form for the terminal cost function Perform modeling and obtain parameters The terminal value function of the control linear function approximation model
[0052] Control strategies are generated online through model predictive control optimization. That is, solving the finite-time domain Optimal control sequence within :
[0053]
[0054] In the formula, For stage cost function; For the terminal cost function; For the system dynamics model; This represents the current system state. From terminal status With control strategy The terminal cost function constituted for:
[0055] Using parameterized form for the terminal cost function To perform modeling, Terminal value function
[0056]
[0057] In the formula, Indicates by parameters The control terminal performance evaluation function; For a learnable weight parameter vector, the feature mapping function The feature mapping function for the terminal state. It is preferable to use a polynomial expansion form for construction.
[0058] Using terminal value functions The method for constructing the terminal gradient feedback path is as follows: Calculate terminal value function Regarding terminal status The gradient; Construct the current performance gradient and apply it to the current control input using the chain rule. The transmission path.
[0059] Terminal value function Regarding terminal status The gradient calculation method is as follows: ; Using the chain rule, construct the performance gradient towards the current control input. transmission path ; The calculation method is as follows:
[0060]
[0061] In the formula, Let be the Jacobian matrix of the state-feature mapping; This represents the partial derivative of the current control input with respect to the predicted terminal state.
[0062] The method for constructing a composite optimization objective function is as follows: Construct the basic cost term, the policy gradient guiding term, and the policy gradient regularization term. Based on the basic cost term, the policy gradient guiding term, and the policy gradient regularization term, construct a composite objective function. ,in: Basic cost function for:
[0063] In the formula, For the first Step state error; This is the weighting matrix for the state error; To control the weighting matrix of the input, To predict the step size; Policy gradient guiding term for:
[0064] In the formula, Indicates the increment of the current control input; Policy gradient regularization term for:
[0065] Based on the basic cost item Policy gradient guidance term and gradient regularization term Constructing a composite objective function The method is as follows:
[0066]
[0067] In the formula, As a dynamic weighting factor, The regularization coefficient is . Basic proportional coefficient; is the numerical stability constant.
[0068] The method for constructing a policy gradient embedded augmentation model is as follows: Establish the system prediction equation represented by the discrete state-space model, express it as a function of the current state and the control input sequence according to the modeling requirements, construct the state recursion relationship, and determine the system prediction state sequence and the control input sequence to be optimized; Construct a standard quadratic form of the composite optimization objective function with respect to the control variables; Construct control constraints and state constraints; Based on the control constraints, state constraints, and the rewritten form of the composite optimization objective function, a constrained policy gradient embedded enhancement model is constructed according to the modeling requirements, and the optimal control input sequence is obtained.
[0069] The system prediction equation represented by the discrete state-space model is:
[0070] In the formula, , These are the discretized state transfer matrix and control input matrix, respectively; The state recursion, expressed as a function of the current state and the control input sequence, is as follows:
[0071] In the formula, For the present The system state at any given moment; Predict the state sequence for the system; The input sequence for the control to be optimized; , These are the expansion matrices of the system state and control input, respectively.
[0072] Composite optimization objective function Rewritten as about control variables The standard quadratic form is:
[0073] In the formula, This is the coefficient matrix of the quadratic terms; A vector of first-order terms; The control constraints are: For each control input Apply upper and lower limits to reflect the actuator's capability boundaries;
[0074] The state constraints are: Predicting the state of the system Limited to a safe and feasible area;
[0075] Rewrite the control constraints and state constraints in matrix inequality form:
[0076] In the formula, the matrix with vector Constructed based on the constraints.
[0077] The constrained policy gradient embedded augmentation model is a standard quadratic programming model with constraints, as follows:
[0078] The optimal control input sequence is obtained by solving the policy gradient embedded enhancement model. The optimal control input sequence is:
[0079] Among them, the first control input of the sequence is selected. After the actual control action is applied at the current moment, it is updated on a rolling basis at each sampling moment to form a closed-loop predictive control.
[0080] In the standard quadratic form of the composite optimization objective function, the coefficient matrix of the quadratic terms... Linear term vector Its composition includes basic cost items and policy gradient guidance term ,in:
[0081] ;
[0082]
[0083]
[0084]
[0085] .
[0086] The gradient expression of the policy objective function with respect to the policy parameters is:
[0087] in, .
[0088] Let the objective function be about The gradient is zero, thus obtaining the optimal analytical solution for the policy parameters. Specifically: when:
[0089] The optimal analytical solution for the policy parameters is obtained. for:
[0090] when When irreversibility is not possible, the Moore-Penrose generalized inverse is used as a substitute, and the gradient expression is as follows: ; By updating the optimal analytical solution in real time It enables dynamic optimization and updating of strategy parameters.
[0091] The following description, in conjunction with the accompanying drawings and preferred embodiments, provides further details: In the current embodiment, the following is proposed: Figure 1 The control flow shown specifically includes: Step S1: Construct the terminal value function and its gradient feedback structure.
[0092] This step aims to construct a parameterized and continuously differentiable terminal value function to replace the pre-defined fixed terminal cost function in traditional model predictive control, thereby more accurately representing the system's long-term performance objective. Based on this value function, a gradient propagation path from the state to the control input is constructed, realizing an explicit feedback mechanism between the policy and the control input, supporting subsequent policy guidance and parameter update processes.
[0093] Step S11: Construct the terminal cost function based on the MPC implicit policy Solving the finite-time domain problem using quadratic programming. The optimal control sequence within the range is then used to construct a continuously differentiable terminal cost function. Used to describe the system's ability to predict terminal states. The expected cumulative return in the future. Specifically, the optimal control sequence:
[0094] Satisfy system dynamics constraints,
[0095] in, For stage cost function; For the terminal cost function; For the system dynamics model; This represents the current system state.
[0096] Terminal cost function Explicitly defined in the predicted terminal state However, its value depends on the optimal control sequence obtained by solving the model predictive control. This control sequence defines the control strategy at the current moment. Specifically, the terminal cost function can be viewed as a composite function of the terminal state and the control strategy, i.e.:
[0097] To further explain, although the strategy From complete control sequence The terminal performance is defined and modeled, but in actual control execution, only the first control input of the sequence is taken. As the control action at the current moment, that is:
[0098] The control strategy is solved again in the next time step through rolling time-domain optimization, so as to achieve real-time updates and closed-loop control.
[0099] It should be noted that the terminal status It is obtained by progressively propagating the control input sequence derived from the strategy in the prediction model, satisfying the following dynamic mapping relationship:
[0100] in, ,and This represents the state sequence obtained through iterative analysis of the prediction model. This represents the mapping that propagates iteratively to the terminal moment within the rolling prediction window based on the system dynamics model.
[0101] To ensure that the objective function retains good continuity and differentiability after embedding the terminal cost function, and to reduce the computational complexity of modeling, a parameterized form is adopted for the terminal cost function. Modeling:
[0102] in, Indicates by parameters The control terminal performance evaluation function; The weight parameter vector is a learnable vector that can be updated in real time based on the rolling optimization structure during task execution. To ensure the function has analytical differentiability, a linear function approximation model is preferred, with the following form:
[0103] in, This is a feature mapping function for the terminal state, used to improve the model's ability to fit nonlinear performance metrics. The feature mapping function... The preferred approach is to construct the feature using a polynomial extension. Taking a quadratic extension feature as an example, let:
[0104] in, The number of state variables. The extended feature is constructed as follows:
[0105] The linear, squared, and pairwise interaction terms of the state are combined to form a high-dimensional feature vector, which is then coupled with learnable weight coefficients. Linear combinations are used to approximate terminal performance metrics.
[0106] Step S12: Construct the terminal gradient feedback path. To guide the optimization of the current control input based on the terminal performance index, a terminal cost function is used. Parametric modeling results Construct a differentiable gradient propagation path from it to the control input.
[0107] First, calculate the terminal value function. Regarding terminal status The gradient is:
[0108] This gradient reflects the sensitivity of terminal state perturbations to future performance gains.
[0109] Secondly, using the chain rule, this performance gradient is constructed to the current control input. The propagation path:
[0110] in, Let be the Jacobian matrix of the state-feature mapping; The partial derivative of the current control input with respect to the predicted terminal state can be obtained recursively using the chain rule based on the system dynamics model:
[0111] The derivative can be solved using automatic differentiation or analytical methods.
[0112] Compared to the traditional problem of fixed or difficult-to-optimize terminal costs, this method constructs a gradient feedback path from control input to terminal performance, thereby achieving explicit differentiable modeling of the terminal cost function and effective response of control variables, improving the directionality of policy optimization and the ability to regulate long-term performance.
[0113] Step S2: Construct an optimization objective function containing policy gradients and regularization terms.
[0114] This step aims to construct a composite optimization objective function that integrates trajectory tracking performance, control cost, policy gradient guidance term, and gradient regularization term.
[0115] Step S21: Construct the basic cost term Based on the deviation between the system's predicted state and the reference trajectory, as well as the energy consumption and amplitude limitations of the control input, a basic cost function is constructed. :
[0116] in, For the first Step state error; This is the weighting matrix for the state error; To control the input weighting matrix, The prediction step size is used. This cost term ensures that the system can effectively track the reference trajectory within a finite prediction time domain and constrains the stationarity of the control input and energy consumption.
[0117] Step S22: Construct the policy gradient guiding term To improve the control strategy's responsiveness to long-term performance, a policy gradient guiding term is introduced into the optimization objective function. This approach explicitly embeds the gradient information of the terminal cost function with respect to the predicted terminal state, and guides the current control input to be adjusted in a direction that improves long-term performance based on the influence path of the control input on the terminal state.
[0118] A linear approximation of the current control input using the predicted terminal state is adopted:
[0119] in, This represents the increment of the current control input. Based on this, the policy gradient guidance term is defined as:
[0120] Step S23: Construct the policy gradient regularization term To enhance the numerical stability and convergence of the optimization process, the gradient of the predicted terminal state is... The norm is used to constrain the gradient, thereby suppressing the adverse effects of excessively large gradient magnitudes during policy updates. This is different from directly controlling the input gradient using regularization. Using the state gradient regularization term reduces computational complexity and can be expressed as:
[0121] Step S24: Construct a composite objective function In steps S21 to S23, the basic cost terms were constructed respectively. Policy gradient guidance term and gradient regularization term .
[0122] To reasonably adjust the weights of the policy gradient guiding term and the gradient regularization term in the overall optimization objective, dynamic weighting factors are introduced respectively. and regularization coefficient Given the basic cost terms at different times and policy gradient guidance term Differences in magnitude, dynamic weighting factor Optimize the ratio structure and combine it with the prediction step size. Normalization is performed to ensure numerical stability across different time domains. Dynamic weighting factor. The specific calculation formula is as follows:
[0123] in, Basic proportional coefficient; It is a numerical stability constant used to prevent numerical instability caused by an excessively small denominator.
[0124] This design makes It can automatically adjust according to the system state. When in a highly disturbed environment or during a fast-moving task, the system state changes significantly, and the basic cost items... Increase, dynamic weighting factor Automatically increasing policy gradient-guided terms have a more significant impact on the objective function; conversely, in steady-state tracking tasks, where state changes are smaller, the basic cost term... Smaller, dynamic weighting factor Automatic reduction and a greater reliance on short-term rolling optimization in the control law enable adaptive adjustment of policy gradient information under different task and environmental conditions.
[0125] Gradient regularization term The regularization coefficient is used to suppress numerical instability caused by excessively large gradient terms, rather than directly affecting control performance. Using a fixed value can avoid interfering with the main objectives of short-term control and long-term performance while ensuring optimization stability, and at the same time simplify the objective function structure. This fixed value is usually determined through numerical simulation or empirical parameter tuning, so that the regularization term has the least impact on the main control objective while maintaining optimization stability.
[0126] Finally, a composite optimization objective function is formed. :
[0127] The objective function has a continuously differentiable structure, making it suitable for automatic differentiation and online numerical solutions. It takes into account the system's tracking performance, long-term performance guidance, and optimization stability, effectively improving the overall performance and robustness of the control strategy.
[0128] Step S3: Construct a policy gradient embedded augmentation model for predictive control optimization problem.
[0129] This step aims to transform the composite optimization objective function of the integrated policy gradient guidance mechanism into a standard quadratic programming (QP) structure, thereby improving the online solution efficiency and the numerical stability of the control policy.
[0130] Step S31: Establish the system prediction equations. Based on the discrete state-space model:
[0131] in, , These are the discretized state transfer matrix and control input matrix, respectively. To adapt to the requirements of QP modeling, the system's predicted state sequence is represented as a function of the current state and the control input sequence, and the state recursion relationship is constructed as follows:
[0132] in, For the present The system state at any given moment; Predict the state sequence for the system; The input sequence for the control to be optimized; , These are the expansion matrices of the system state and control input, respectively.
[0133] Step S32: Construct the quadratic form of the objective function. To improve the efficiency of numerical optimization, the composite objective function will be used. Rewritten as about control variables Standard quadratic form:
[0134] in: This is the coefficient matrix of the quadratic terms; This is a linear term vector, specifically structured as follows: (1) Basic cost item
[0135] Based on state recursion structure and reference trajectory stacking vector Expanding the basic cost term as follows:
[0136] in, ; constant term This can be ignored during the optimization process.
[0137] (2) Policy gradient guiding term
[0138] This term is linearly dependent on the change in the current control input, and is expressed as:
[0139] (3) Gradient regularization term
[0140] This depends only on the strategy parameters. , and control variables Irrelevant, therefore for optimization variables Since the gradient is always zero, the objective function is not included in this QP modeling step. The direct impact.
[0141] Finally, the QP coefficient matrix and linear terms are obtained as follows:
[0142]
[0143] Step S33: Construct state and control constraints. To ensure that the control inputs and state variables satisfy the actual execution and safety boundaries, construct the following linear inequality constraints: Control constraints: For each control input Upper and lower limits are imposed to reflect the actuator's capability boundaries, in the form of:
[0144] State constraints: Predicting the state of the system Restricted to a safe and feasible area to avoid outbound or unsafe operations, in the form of:
[0145] The above constraints are uniformly rewritten in compact matrix inequality form:
[0146] Among them, matrix with vector It is constructed based on the specific system model and constraints.
[0147] Step S34: Solve the quadratic programming problem. The final constrained standard quadratic programming problem is as follows:
[0148] Solving the above QP problem yields the optimal control input sequence:
[0149] Select the first control input in the sequence The actual control action is applied to the system at the current moment, and then updated on a rolling basis at each sampling moment to form a closed-loop predictive control.
[0150] Step S4: Update online strategy parameters.
[0151] This step aims to improve the real-time responsiveness of the policy guidance mechanism to long-term performance. This is achieved by constructing a composite optimization objective function with respect to the policy parameters. The gradient expression is derived and combined with the analytical update method to achieve adaptive and efficient adjustment of policy parameters, thereby enhancing the dynamic adaptability of the policy guidance term to the optimization direction.
[0152] Step S41: Construct the objective function with respect to the policy parameters The gradient. Due to the fundamental cost term. Do not explicitly depend on policy parameters Therefore, it can be ignored in gradient calculation; only the policy gradient guiding term is considered. and gradient regularization term For parameters The impact of this on constructing the overall objective function. about The gradient expression is as follows:
[0153] make:
[0154] The gradient expression can then be further simplified to:
[0155] It should be noted that the dynamic weighting factor Although its construction indirectly depends on However, its value has been determined in the current optimization round by predicting the trajectory and control input, so it can be regarded as a constant in gradient calculation and there is no need to backpropagate it.
[0156] Step S42: Calculate the analytical update expression for the policy parameters. To obtain the analytical form of the policy parameters, let the objective function be related to... The gradient is zero, that is:
[0157] The optimal analytical solution for the policy parameters is obtained. for:
[0158] when When irreversibility is not possible, the Moore-Penrose generalized inverse can be used as an alternative:
[0159] This update method has a closed-form analytical solution, high computational efficiency, and low implementation difficulty. It is beneficial to improve the stability and timeliness of policy parameter updates, and is suitable for rapid deployment of online control tasks. At the same time, it avoids the problems of slow convergence speed, oscillation instability and hyperparameter sensitivity commonly found in traditional gradient iterative updates.
[0160] Example 1: Using a bipedal robot with 30 degrees of freedom as the controlled object, the enhanced model predictive control method of this invention is applied. Its control system architecture is as follows: Figure 2 As shown, it mainly includes: gait planner, enhanced MPC controller, whole-body controller (WBC) and joint controller.
[0161] The enhanced MPC controller employs a linearized centroid dynamics model, where the system state variables are defined as follows: Including the center of mass angle ,Location angular velocity and linear velocity The control variable is defined as follows: Includes the contact force between the left and right soles and the ground. and torque .
[0162] Control Flow Description: Within each control cycle, the following steps are performed: (1) State acquisition and trajectory prediction: Estimating the current state of the robot's center of mass. Predicting the future based on dynamic models Step state sequence ; (2) Reference trajectory matching: Obtain the centroid reference trajectory at the current moment from the gait planner. And extend to the entire prediction time domain; (3) Terminal gradient feedback construction: based on predicted terminal state Construct parameterized terminal value functions And calculate its equivalent control input. The gradient feedback structure; (corresponding to step S1) (4) Construction of composite optimization objective function: Constructing a function that includes the basic cost term Policy gradient guidance term and gradient regularization term Optimization objective function (Corresponding to step S2) (5) Optimization solution: The objective function is transformed into a quadratic programming (QP) problem. Under the constraints of dynamics and friction cone, the optimal foot force sequence is solved. (Corresponding to step S3) (6) Strategy parameter update: Update terminal value function parameters online using analytical solution. (Corresponding to step S4) (7) Control execution: Input the first control term of the optimal sequence. The data is sent to the Whole Body Motion Controller (WBC), which breaks it down into the desired torques for each joint. ,Location and speed The instructions are then sent to the joint controller, which drives the actuator to complete the action.
[0163] (8) Rolling optimization: Enter the next control cycle and repeat the above process to achieve continuous closed-loop control.
[0164] Comparison of implementation results: like Figure 4 , Figure 5 As shown, through Figure 4 and Figure 5The simulation results were compared and analyzed. Under the same uneven terrain, the bipedal robot using the method of this invention achieved smooth and continuous walking, and the fluctuation amplitudes of its base roll angle, pitch angle, and yaw angle were stabilized at [values to be filled in]. , and Within this range, it exhibits good tracking stability and control robustness. In contrast, robots using traditional MPC control methods, lacking long-term performance guidance and online policy adaptation capabilities, experience continuous attitude instability during walking. Their roll and pitch angles diverge continuously after disturbances, and their center of gravity... The displacement error in the direction continued to increase to 0.5m, eventually causing the robot to fall and be unable to complete a full gait cycle.
[0165] Although the present invention has been disclosed above with reference to preferred embodiments, it is not intended to limit the present invention. Any person skilled in the art can make possible changes and modifications to the technical solutions of the present invention by utilizing the methods and techniques disclosed above without departing from the spirit and scope of the present invention. Therefore, any simple modifications, equivalent changes and alterations made to the above embodiments based on the technical essence of the present invention without departing from the content of the technical solutions of the present invention shall fall within the protection scope of the technical solutions of the present invention.
[0166] The contents not described in detail in this specification are common knowledge to those skilled in the art.
Claims
1. A policy gradient embedded augmentation model predictive control method, characterized in that... include: Construct the terminal cost function, and use the terminal cost function to construct the terminal gradient feedback path; Based on the terminal cost function and the terminal gradient feedback path, a composite optimization objective function containing policy gradient and regularization term is constructed. A policy gradient embedded enhancement model is constructed for the structural transformation of the composite optimization objective function. The policy gradient embedded enhancement model is solved according to the modeling requirements to obtain the optimal control input sequence. Construct the gradient expression of the composite optimization objective function with respect to the policy parameters, and achieve dynamic optimization by updating the policy parameters in real time.
2. The policy gradient embedded augmentation model predictive control method according to claim 1, characterized in that: The method for constructing the terminal cost function is as follows: The control strategy generated through online optimization using model predictive control Construct the terminal cost function Used to describe the predicted terminal state The expected cumulative return in the future; Using parameterized form for the terminal cost function Perform modeling and obtain parameters The terminal value function of the control linear function approximation model .
3. The policy gradient embedded augmentation model predictive control method according to claim 2, characterized in that: Control strategies are generated online through model predictive control optimization. That is, solving the finite-time domain Optimal control sequence within : In the formula, For stage cost function; For the terminal cost function; For the system dynamics model; This represents the current system state. From terminal status With control strategy The terminal cost function constituted for: Using parameterized form for the terminal cost function To perform modeling, Terminal value function In the formula, Indicates by parameters The control terminal performance evaluation function; For a learnable weight parameter vector, The feature mapping function for the terminal state. It is preferable to use a polynomial expansion form for construction.
4. The policy gradient embedded augmentation model predictive control method according to claim 3, characterized in that: Using terminal value functions The method for constructing the terminal gradient feedback path is as follows: Calculate terminal value function Regarding terminal status The gradient; Construct the current performance gradient and apply it to the current control input using the chain rule. The transmission path.
5. The policy gradient embedded augmentation model predictive control method according to claim 4, characterized in that: The terminal value function Regarding terminal status The gradient calculation method is as follows: ; Using the chain rule, construct the performance gradient towards the current control input. transmission path ; The calculation method is as follows: In the formula, Let be the Jacobian matrix of the state-feature mapping; This represents the partial derivative of the current control input with respect to the predicted terminal state.
6. The policy gradient embedded augmentation model predictive control method according to claim 4, characterized in that: The method for constructing a composite optimization objective function is as follows: Construct the basic cost term, the policy gradient guiding term, and the policy gradient regularization term. Based on the basic cost term, the policy gradient guiding term, and the policy gradient regularization term, construct a composite objective function. ,in: Basic cost function for: In the formula, For the first Step state error; This is the weighting matrix for the state error; To control the weighting matrix of the input, To predict the step size; Policy gradient guiding term for: In the formula, Indicates the increment of the current control input; Policy gradient regularization term for: 。 7. The policy gradient embedded augmentation model predictive control method according to claim 6, characterized in that: Based on the basic cost item Policy gradient guiding term and gradient regularization term Constructing a composite objective function The method is as follows: In the formula, As a dynamic weighting factor, The regularization coefficient is . Basic proportional coefficient; is the numerical stability constant.
8. The policy gradient embedded augmentation model predictive control method according to claim 7, characterized in that: The method for constructing a policy gradient embedded augmentation model is as follows: Establish the system prediction equation represented by the discrete state-space model, express it as a function of the current state and the control input sequence according to the modeling requirements, construct the state recursion relationship, and determine the system prediction state sequence and the control input sequence to be optimized; Construct a standard quadratic form of the composite optimization objective function with respect to the control variables; Construct control constraints and state constraints; Based on the control constraints, state constraints, and the rewritten form of the composite optimization objective function, a constrained policy gradient embedded enhancement model is constructed according to the modeling requirements, and the optimal control input sequence is obtained.
9. The policy gradient embedded augmentation model predictive control method according to claim 8, characterized in that: The system prediction equation represented by the discrete state-space model is: In the formula, , These are the discretized state transfer matrix and control input matrix, respectively; The state recursion, expressed as a function of the current state and the control input sequence, is as follows: In the formula, For the present The system state at any given moment; Predict the state sequence for the system; The input sequence for the control to be optimized; , These are the expansion matrices of the system state and control input, respectively.
10. The policy gradient embedded augmentation model predictive control method according to claim 8, characterized in that: The composite optimization objective function Rewritten as about control variables The standard quadratic form is: In the formula, This is the coefficient matrix of the quadratic terms; A vector of first-order terms; The control constraints are: For each control input Apply upper and lower limits to reflect the actuator's capability boundaries; The state constraints are: Predicting the state of the system Limited to a safe and feasible area; Rewrite the control constraints and state constraints in matrix inequality form: In the formula, the matrix with vector Constructed based on the constraints.
11. The policy gradient embedded augmentation model predictive control method according to claim 8, characterized in that: The constrained policy gradient embedded enhancement model is a standard quadratic programming model with constraints, as follows: The optimal control input sequence is obtained by solving the policy gradient embedded enhancement model. The optimal control input sequence is: Among them, the first control input of the sequence is selected. After the actual control action is applied at the current moment, it is updated on a rolling basis at each sampling moment to form a closed-loop predictive control.
12. The policy gradient embedded augmentation model predictive control method according to claim 9, characterized in that: In the standard quadratic form of the composite optimization objective function, the coefficient matrix of the quadratic terms... Linear term vector Its composition includes basic cost items and policy gradient guidance term ,in: ; 。 13. The policy gradient embedded augmentation model predictive control method according to claim 9, characterized in that: The gradient expression of the composite optimization objective function with respect to the policy parameters is as follows: in, .
14. Let the objective function be about The gradient is zero, thus obtaining the optimal analytical solution for the policy parameters. Specifically: when: The optimal analytical solution for the policy parameters is obtained. for: when When irreversibility is not possible, the Moore-Penrose generalized inverse is used as a substitute, and the gradient expression is as follows: ; By updating the optimal analytical solution in real time Enables dynamic optimization and updating of strategy parameters.
Citation Information
Patent Citations
Inverse reinforcement learning with model predictive control
CN112906882A
Trajectory Tracking Method for Variable-Load Mobile Manipulator Based on Reinforcement Learning and Model Predictive Control
CN119748459B