Interference-considered terminal strategy optimization heuristic MPC control method and interference-considered terminal strategy optimization heuristic MPC control system
By constructing a nonlinear dynamic model and optimizing the terminal cost function using a disturbance predictor, the problems of estimation error accumulation and poor adaptability of MPC control under time-varying disturbances are solved, thereby improving the stability and real-time performance of the unmanned system.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- BEIHANG UNIV
- Filing Date
- 2026-02-26
- Publication Date
- 2026-05-12
AI Technical Summary
Existing MPC control technology is difficult to effectively adapt to time-varying interference in complex time-varying interference scenarios, resulting in the accumulation of terminal cost estimation errors, poor online adaptability, high algorithm conservatism, and difficulty in real-time deployment.
A discrete-time nonlinear dynamic model with time-varying disturbances is constructed. A disturbance predictor is introduced to fit the model parameters through historical disturbance data. The optimal parameterized terminal cost function is obtained by supervised learning and embedded into the MPC framework. The terminal MPC optimization problem is solved by combining the discrete-time nonlinear dynamic model and the disturbance prediction results.
It achieves precise control under time-varying disturbance scenarios, reduces control errors caused by disturbances and uncertainties, improves the online control stability and reliability of unmanned systems, reduces solution time fluctuations, and adapts to rapid autonomous control tasks in complex environments.
Smart Images

Figure CN122018319A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of control, in particular to an MPC control method and system considering terminal strategy optimization heuristics of interference. BACKGROUND
[0002] As a kind of efficient nonlinear system control method, model predictive control (MPC) is widely used in unmanned system, spacecraft autonomous detection, on-orbit service autonomous control and other complex control scenarios due to its advantages in handling multiple constraints and multiple variables. Terminal cost / terminal value, as a core component of the MPC framework, directly determines the solution efficiency, control performance and system closed-loop stability of the MPC optimization problem. Its design rationality and adaptability are the key to improving the robustness and online adaptability of unmanned systems in complex environments. However, there are structured and time-varying disturbances in complex control scenarios, and the system is easily affected by white noise and modeling errors. The traditional fixed design of terminal cost has been difficult to meet the control requirements of high precision and high real-time. Therefore, how to construct a terminal cost that can adapt to time-varying disturbances, reduce estimation errors, and balance real-time and robustness has become a technical problem to be solved in the field of MPC control technology.
[0003] To solve the above problems of terminal cost design, in recent years, there have been many improvement attempts in related fields. On the one hand, some studies use data-driven or reinforcement learning methods to construct and approximate the terminal cost / terminal value of MPC, learning the terminal cost as a function that can shorten the prediction domain and improve local decision quality, trying to optimize the adaptability of the terminal cost through the adaptive characteristics of data-driven; on the other hand, traditional robust MPC and adaptive MPC algorithms rely on mature uncertainty processing theory, through the introduction of robust constraints, adaptive adjustment mechanism and other ways, to cope with the influence of environmental uncertainty and external interference on the control performance of MPC, in order to improve the anti-interference ability of the system.
[0004] Despite the improvements made by existing technologies, significant shortcomings remain in complex time-varying disturbance scenarios and real-time, resource-constrained deployment scenarios, making it difficult to meet practical application needs. First, research employing data-driven or reinforcement learning methods typically designs terminal costs based on static or stationary assumptions about environmental disturbances, failing to fully consider the dynamic characteristics of time-varying disturbances in complex environments. This leads to the accumulation of terminal cost estimation errors over time, insufficient online adaptability, and an inability to effectively adapt to system control requirements under time-varying disturbances. Second, improved algorithms such as traditional robust MPC and adaptive MPC suffer from two core practical defects when dealing with structured, time-recurring disturbances: one, to ensure system robustness, they often introduce significant conservative designs, resulting in reduced online solution efficiency for MPC optimization problems and a significant decrease in system control performance; two, the computational and implementation costs of these algorithms are high, often requiring online or offline computation of time-varying invariant sets or frequent re-estimation of model parameters and cost parameters, making them difficult to deploy in space environments with high real-time requirements and limited resources, such as autonomous spacecraft exploration and on-orbit servicing autonomous control. Summary of the Invention
[0005] To address the shortcomings of existing technologies, this application provides an MPC control method and system inspired by terminal strategy optimization that takes into account interference. This solves the technical problems of existing technologies, such as the inability to effectively address the accumulation of MPC terminal cost estimation errors, poor online adaptability, high algorithm conservatism, and difficulty in real-time deployment under complex time-varying interference.
[0006] To achieve the above objectives, this application provides the following technical solution: In a first aspect, embodiments of this application provide an MPC control method inspired by terminal strategy optimization considering interference. This MPC control method includes: constructing a discrete-time nonlinear dynamic model containing time-varying disturbances for the control process of a nonlinear unmanned system, to characterize the evolution of the system state vector with control input, time-varying disturbances, system white noise, and modeling error; and introducing a disturbance predictor. g The parameters of the disturbance model are obtained by fitting historical disturbance data. ; in the k The step involves estimating the current disturbance through sensing or identification. And using a perturbation predictor g Predicting future steps i The disturbance was analyzed, and the disturbance prediction results were obtained. The parameterized terminal cost function is obtained through supervised learning. Optimal parameters The optimal terminal cost function is obtained. Construct a terminal model predictive control (MPC) framework, formulate the corresponding terminal model predictive control (MPC) optimization problem, and determine the optimal terminal cost function. Embedded terminal MPC framework, replacing traditional terminal costs; combining discrete-time nonlinear dynamics model and disturbance prediction results. At any moment k Solve the constrained terminal MPC optimization problem, output a control law to output control input to the unmanned system based on the control law, and complete online control.
[0007] Secondly, embodiments of this application provide an MPC control system inspired by terminal strategy optimization considering interference. The MPC control system inspired by terminal strategy optimization considering interference includes: a model building module, a prediction module, a parameter learning module, a framework building module, and a solution control module.
[0008] Specifically, the model building module is used to construct a discrete-time nonlinear dynamic model with time-varying disturbances for the control process of nonlinear unmanned systems, so as to characterize the evolution of the system state vector with control input, time-varying disturbances, system white noise, and modeling error; the prediction module is used to introduce a disturbance predictor. g The parameters of the disturbance model are obtained by fitting historical disturbance data. ; in the k The step involves estimating the current disturbance through sensing or identification. And using a perturbation predictor g Predicting future steps i The disturbance was analyzed, and the disturbance prediction results were obtained. The parameter learning module is used to obtain the parameterized terminal cost function through supervised learning. Optimal parameters The optimal terminal cost function is obtained. The framework building module is used to construct the terminal model predictive control (MPC) framework, forming the corresponding terminal model predictive control (MPC) optimization problem, and determining the optimal terminal cost function. The embedded terminal MPC framework replaces the traditional terminal cost; the solution control module combines discrete-time nonlinear dynamics models and disturbance prediction results. At any moment k Solve the constrained terminal MPC optimization problem, output a control law to output control input to the unmanned system based on the control law, and complete online control.
[0009] Thirdly, embodiments of this application provide an electronic device, which includes: a processor, a memory, and a program stored in the memory and executable on the processor. When the program is executed by the processor, it implements the MPC control method inspired by terminal strategy optimization considering interference as described in the first aspect.
[0010] Fourthly, embodiments of this application provide a computer-readable storage medium storing a program or instructions that, when executed by a processor, implement the MPC control method inspired by terminal strategy optimization considering interference as described in the first aspect.
[0011] This application provides an MPC control method and system inspired by terminal strategy optimization considering interference. Compared with the prior art, it has the following advantages: This application focuses on nonlinear unmanned systems. By constructing a discrete-time nonlinear dynamic model with time-varying disturbances, it can accurately characterize the evolution of the system state vector with control input, time-varying disturbances, system white noise, and modeling errors, providing accurate model support for the formulation of subsequent control strategies. Furthermore, it introduces a disturbance predictor. g By fitting parameters using historical disturbance data and predicting future disturbances, disturbance information can be obtained in advance, effectively mitigating the adverse effects of time-varying disturbances on system control performance. This application obtains the optimal terminal cost function through supervised learning and embeds it into the MPC framework to replace the traditional terminal cost, optimizing the design rationality and adaptability of the terminal cost. Combining the discrete-time nonlinear dynamics model and disturbance prediction results, the constrained terminal MPC optimization problem is solved, and the control law is output, realizing online control of the unmanned system. This application can adapt to time-varying disturbance scenarios, reduce control errors caused by disturbances and various uncertainties, and improve the stability, accuracy, and reliability of online control of nonlinear unmanned systems. Attached Figure Description
[0012] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0013] Figure 1 This is a flowchart illustrating an MPC control method inspired by terminal strategy optimization considering interference, provided in an embodiment of this application. Figure 2 This is a schematic diagram of the structure of an MPC control system inspired by terminal strategy optimization considering interference, provided in an embodiment of this application. Figure 3 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation
[0014] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention are described clearly and completely. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0015] It should be noted that, in this document, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes the element.
[0016] This application provides an MPC control method and system inspired by terminal strategy optimization that takes into account interference. This solves the technical problems of existing technologies, such as the inability to effectively address the accumulation of terminal cost estimation errors, poor online adaptability, high algorithm conservatism, and difficulty in real-time deployment under complex time-varying interference.
[0017] To better understand the above technical solutions, the following will provide a detailed explanation of the technical solutions in conjunction with the accompanying drawings and specific implementation methods.
[0018] The following section first introduces an MPC control method inspired by terminal strategy optimization that takes into account interference, provided in the embodiments of this application.
[0019] This application provides a flowchart illustrating an MPC control method inspired by terminal strategy optimization considering interference, as shown in the embodiments below. Figure 1 As shown, the MPC control method inspired by terminal strategy optimization considering interference may include the following steps S110-S150.
[0020] S110. For the control process of nonlinear unmanned systems, a discrete-time nonlinear dynamic model with time-varying disturbances is constructed to characterize the evolution of the system state vector with control input, time-varying disturbances, system white noise and modeling error.
[0021] Understandably, this application addresses the control process of nonlinear unmanned systems by constructing a discrete-time nonlinear dynamic model with time-varying disturbances. This model can accurately characterize the evolution of the system state vector with control input, time-varying disturbances, system white noise, and modeling errors, and depict the changing patterns of the system state. This provides an accurate and reliable model foundation for subsequent disturbance prediction, terminal cost function optimization, and MPC optimization problem solving, avoiding control errors caused by incomplete model representation and ensuring the rationality of subsequent control steps.
[0022] S120, Introducing a disturbance predictor g The parameters of the disturbance model are obtained by fitting historical disturbance data. ; in the k The step involves estimating the current disturbance through sensing or identification. And using a perturbation predictor g Predicting future steps i The disturbance was analyzed, and the disturbance prediction results were obtained. .
[0023] Understandably, this application introduces a disturbance predictor. g By fitting historical disturbance data to obtain disturbance model parameters, the accuracy of disturbance prediction can be optimized by fully utilizing historical disturbance information. This application is in its first... k The step involves estimating the current disturbance through sensing or identification, and then utilizing this disturbance predictor. g Predicting future steps i By identifying disturbances and obtaining disturbance prediction results, we can obtain future disturbance information in advance, realize the prediction of time-varying disturbances, provide disturbance reference for solving subsequent MPC optimization problems, effectively weaken the adverse effects of time-varying disturbances on system control performance, and improve the system's adaptability to time-varying disturbances.
[0024] S130. Obtain the parameterized terminal cost function through supervised learning. Optimal parameters The optimal terminal cost function is obtained. .
[0025] Understandably, this application obtains the optimal parameters of the parameterized terminal cost function through supervised learning, thereby determining the optimal terminal cost function. This makes the terminal cost function more adaptable to the control requirements of nonlinear unmanned systems and time-varying disturbance scenarios. Compared with the traditional fixed terminal cost, it has stronger adaptability and provides terminal cost support for subsequent embedding of the terminal model predictive control (MPC) framework and optimization of MPC control performance.
[0026] It should be noted that this application constructs a parameterized terminal cost function through supervised learning. Compared with traditional MPC methods that rely on linearized models, invariant set calculations, or analytical terminal Lyapunov functions, this significantly reduces the manual workload and model dependence in terminal cost design. The terminal cost function construction time of this application is lower than that of traditional analytical construction methods, and it eliminates the need to solve for the complex stability sets of high-dimensional systems, thus achieving rapid generation of the terminal cost function, a wider state coverage, and higher adaptability. The error of the terminal cost function approximation decreases exponentially with the number of training steps, reaching an accuracy level suitable for MPC within a finite number of iterations, significantly outperforming traditional static terminal terms based on linear approximation.
[0027] S140. Construct a terminal model predictive control (MPC) framework, forming a corresponding terminal model predictive control (MPC) optimization problem, and then optimize the terminal cost function. Embedded terminal MPC framework, replacing the cost of traditional terminals.
[0028] Understandably, this application constructs a terminal model predictive control (MPC) framework to form a corresponding terminal MPC optimization problem, clarifying the core framework and optimization objective of MPC control. Then, this application embeds the optimal terminal cost function into the constructed MPC framework, replacing the traditional terminal cost. This optimizes the solution direction and quality of the MPC optimization problem, overcomes the poor adaptability of traditional terminal costs, and improves the accuracy of solving the MPC optimization problem.
[0029] S150, combining discrete-time nonlinear dynamics model and disturbance prediction results At any moment k Solve the constrained terminal MPC optimization problem, output a control law to output control input to the unmanned system based on the control law, and complete online control.
[0030] Understandably, this application combines discrete-time nonlinear dynamics models and perturbation prediction results at time [time value missing]. k Solving the constrained terminal MPC optimization problem fully utilizes the model's accurate representation capabilities and disturbance prediction information, resulting in a control law that better reflects the actual operating state of the system and is more capable of handling time-varying disturbances. This application uses this control law to output control inputs to an unmanned system and complete online control, achieving real-time and precise control of the nonlinear unmanned system. This effectively reduces control deviations caused by time-varying disturbances, system white noise, and modeling errors, improving the stability and reliability of the unmanned system's online control. Furthermore, this application relies on explicit solution and execution logic, ensuring the feasibility of online control.
[0031] Furthermore, this application achieves fast convergence performance of MPC under time-varying disturbance conditions. Unlike traditional robust MPC methods that require increased conservative margins, leading to a decrease in solution speed, this invention significantly reduces the sensitivity of the solution process to disturbance changes by learning the variation law of disturbances. Under typical high-dynamic state change test conditions, the MPC solution time fluctuation amplitude of this application is reduced, the average number of convergence iterations is reduced, and it can still maintain normal convergence even under strong disturbance injection (the maximum disturbance amplitude is 1.5 times the rated disturbance of the system), while traditional methods show non-convergence under the same scenario. Therefore, this application achieves significantly better solution stability and disturbance resistance performance than the background technology in time-varying disturbance environments.
[0032] It should be noted that the MPC control method inspired by terminal strategy optimization considering interference provided in this application is applied to rapid autonomous control tasks in complex and extreme environments considering external interference. Its application scenarios include, but are not limited to, high-precision path tracking of unmanned vehicles, motion operations of intelligent manufacturing robotic arms, flight control of UAV swarms, on-orbit servicing operations of space robots, and complex unmanned systems for flight and landing trajectory control of deep space probes.
[0033] In some embodiments, the discrete-time nonlinear dynamics model satisfies the expression: ; In the formula, k For discrete time steps, and They represent the first k Discrete time step and the first The system state vector at discrete time steps. To control the input, For modelable time-varying recurring disturbances, It represents at least one of the system white noise and modeling error; the system state vector includes at least one of relative position, velocity, attitude angle, and angular velocity; the control input includes at least one of thruster pulse, torque, and force distribution. It is a nonlinear state transition function that characterizes the evolution of the system state; It is continuously differentiable on its domain and satisfies .
[0034] In the embodiments of this application, it is understood that the discrete-time nonlinear dynamic model accurately characterizes the first... through explicit expressions and parameter definitions. k Discrete time step and the first The evolution relationship of the discrete-time step system state vector; considering modelable time-varying recurring disturbances and uncertainties such as system white noise / modeling errors, so that the influence of disturbances and various errors can be accurately identified.
[0035] Furthermore, the nonlinear state transition function Continuously differentiable and satisfying f The property of (0,0)=0 ensures the smoothness of the model and the rationality of the zero initial state, providing a foundation for solving the model, subsequent disturbance prediction, and solving the MPC optimization problem. Meanwhile, this application clarifies that the system state vector includes key parameters such as relative position and velocity, and the control input includes core control quantities such as thruster pulses and torque. This enables the model to accurately match the actual operating characteristics of the nonlinear unmanned system, closely aligning with real-world control scenarios. This improves the accuracy and practicality of the model representation, providing precise and realistic model support for the smooth implementation of subsequent control steps and the improvement of control performance. It effectively avoids control deviations caused by fuzzy model parameters or representations that do not accurately reflect reality.
[0036] In some embodiments, the terminal model predictive control (MPC) framework satisfies the expression: ; In the formula, Indicates the current moment k For the future The predicted value made by the system state vector of the step; Indicates the current moment k For the future k+N The predicted value made by the system state vector of the step; N It is a positive integer; Indicates the current moment k For the future The predicted value made by the control input of the step; The parameter to be estimated; This is the stage cost function, used to characterize the time domain of prediction. Track the degree of deviation from the desired state and control input. The control resources consumed; The terminal cost function is parameterized to quantify the overall cost of the system's future operation in the predicted end-time state. The optimal cost function is used to characterize the cost at the current discrete time step. k When the system is in state At that time, the optimal cumulative cost that can be achieved by solving the terminal model predictive control MPC optimization problem is determined.
[0037] In the embodiments of this application, it is understood that the terminal model predictive control (MPC) framework defines the relationship between the predicted state, predictive control input, and cost function at each future step within the prediction time domain, providing a basis and guidance for solving the MPC optimization problem. This application explicitly distinguishes the current time step. k Asynchronous with the future ( step, The predicted state and predicted control input (step) can accurately characterize the evolution trajectory of the system state and control input in the predicted time domain, which facilitates subsequent targeted optimization of control strategies.
[0038] Furthermore, the stage cost function can accurately characterize the state tracking deviation and resource consumption of control inputs within the prediction time domain, while the parameterized terminal cost function can quantify the comprehensive cost of the system's future operation in the predicted end-time state. The combination of these two functions makes the MPC optimization objective more comprehensive and aligned with actual control requirements. The optimal cost function can directly provide the optimal judgment basis for the control law output, effectively improving the accuracy and relevance of solving the terminal model predictive control MPC optimization problem, and further ensuring the stability and control performance of the online control of the nonlinear unmanned system.
[0039] In some embodiments, the aforementioned terminal model predictive control (MPC) framework includes a state set. and input constraint set And introduce optimal parameters As a terminal item.
[0040] The constraints of the terminal model predictive control (MPC) framework satisfy the following expression: ; In the formula, ; For a set of states, For the input constraint set; For the disturbance prediction results, Indicates the current moment k For the future The predicted value made by the system state vector of the step; Represents the discrete time step at the current moment. k When making predictions about the future, the prediction state of the 0th prediction step corresponds to the starting point of the prediction time domain.
[0041] In this embodiment, it is understood that the terminal model predictive control (MPC) framework, by introducing a state set and an input constraint set, and using the optimal parameters as terminal terms, while setting explicit constraint expressions, delineates reasonable boundaries for solving the terminal model predictive control (MPC) optimization problem, ensuring that the solution process is orderly and controllable. Specifically, the state set and input constraint set can effectively constrain the system state and control input, preventing the system state from exceeding the safe range and the control input from exceeding hardware performance limitations, thus ensuring the safety and compliance of unmanned system control. Based on the disturbance prediction results and the future... The introduction of the step prediction state and the prediction time domain starting point (the prediction state of the 0th prediction step) enables the constraints of the terminal model predictive control (MPC) framework to fully combine disturbance prediction information and prediction state evolution law, adapt to time-varying disturbance scenarios, and improve the rationality and adaptability of the constraints.
[0042] In some embodiments, this application will include a disturbance predictor. g By incorporating the perturbation features into MPC, a predictive perturbation model is obtained, which satisfies the expression: In the formula, The parameters for the perturbation model can be fitted using historical perturbation data, such as recurrent neural networks, AR models, or low-dimensional periodic decomposition. In actual online operation, at the [missing information]... k Step estimate current disturbance From sensing or identification, and using a disturbance predictor g Predicting the future Therefore, the expression for the terminal model predictive control MPC framework (MPC-θ) contains... and use in the optimization objective If a hard guarantee of closed-loop robustness is required, soft constraints / constraint tightening can be introduced, or adjustments can be made during optimization. To make the uncertain set robust, for example, by adding a conservative margin to the state in the constraints.
[0043] In some embodiments, the parameterized terminal cost function obtained through supervised learning is described above. Optimal parameters The optimal terminal cost function is obtained. Specifically, the aforementioned S130 may include the following steps: S210. Obtain an expert dataset covering interference samples from different times and phases. In the formula, For training state samples, The baseline control input obtained for long-domain MPC. Indicates baseline controller, for The corresponding long domain cost, This represents the number of samples.
[0044] Understandably, this application obtains an expert dataset covering interference samples of different times and phases. This expert dataset includes training state samples, baseline control inputs, long-domain costs, and the number of samples, which can fully cover various time-varying interference scenarios. It provides comprehensive and representative training data support for supervised learning, ensuring that the terminal cost function obtained from training can be adapted to different interference conditions and avoiding the problem of insufficient adaptability of the terminal cost function due to single samples and incomplete coverage.
[0045] S220, Parameterize the terminal cost function Let it be in the form of a linear parameterized basis function: In the formula, For the parameter to be estimated, Indicates transpose. For predefined basis function vectors.
[0046] In one example, the predefined basis function vector is at least one of radial basis function vector and orthogonal polynomial basis function vector.
[0047] Understandably, this application sets the parameterized terminal cost function in the form of a linear parameterized basis function, clearly defining the relationship between the parameters to be estimated and the predefined basis function vectors. This simplifies the structure of the parameterized terminal cost function, reduces the difficulty and computational complexity of solving the parameters to be estimated, and avoids the problem of difficulty in solving parameters and difficulty in practical application caused by the complexity of the function structure.
[0048] S230, in the parameter to be estimated θ Solve the SMP convex optimization problem with descent constraints to obtain the optimal parameters. In order to determine the optimal terminal cost function The SMP convex optimization problem with descent constraints is used to optimize the parameterized terminal cost function during the offline supervised learning phase. The parameters to be estimated .
[0049] In one example, with the stage cost being quadratic, the SMP convex optimization problem with descent constraints is a convex quadratic constraint problem with a quadratic objective and linear inequality constraints, which is solved offline using a convex optimization solver.
[0050] Understandably, this application involves parameters to be estimated. Solve the SMP convex optimization problem with descent constraints to obtain the optimal parameters. And determine the optimal terminal cost function. Furthermore, this SMP convex optimization problem with descent constraints is specifically designed for optimizing the parameters to be estimated during the offline supervised learning phase. It can accurately optimize the parameters to be estimated in the parameterized terminal cost function, ensuring that the obtained optimal terminal cost function has both good fitting accuracy and reasonableness. Simultaneously, relying on the characteristics of SMP convex optimization, this application can guarantee the efficiency and global optimality of parameter solving, avoiding getting trapped in local optima, further improving the performance of the terminal cost function, and providing support for subsequently embedding the optimal terminal cost function into the MPC framework and improving MPC control performance.
[0051] Based on this, this application employs a Value Function Learning (VFL) method to optimize and approximate the terminal value function (terminal cost function) in the MPC framework through supervised learning. Based on an expert dataset, this application solves a convex optimization problem with descent constraints in SMP to optimize the parameters of the parameterized terminal cost function. Optimization is performed to obtain the optimal terminal value function that accurately approximates the real future cumulative cost. In other words, this application uses a value function learning method to optimize the parameterized terminal cost function based on an expert dataset to obtain the optimal terminal value function. This learning process solves a SMP convex optimization problem with descent constraints, which improves the accuracy of the value function approximation while ensuring the closed-loop stability of the system. This enables short-domain MPC to achieve control performance comparable to long-domain MPC, effectively solving the technical problems of insufficient control performance and difficult online deployment of traditional MPC in time-varying disturbance scenarios.
[0052] In some embodiments, the SMP convex optimization problem with descent constraints satisfies the expression: ; ; in, This is a regularization coefficient used to prevent overfitting of the parameterized terminal cost function. For regularization terms; The total number of training samples, These are training state samples; For the first Control input for each sample; For the first Perturbation estimation for each sample; for The corresponding long-domain cost; Indicates the parameter to be estimated belong d 3D real vector space; Represents the real number field.
[0053] Let be the stage cost function, and characterize the . The instantaneous operating cost under each sample is used to quantify state tracking deviation and control energy consumption; Indicates the first Based on training state samples, for each training sample... Control input and disturbance estimation The calculated terminal cost function value corresponding to the predicted state at the next time step.
[0054] In the embodiments of this application, it is understood that the SMP convex optimization problem with descent constraints, through explicit expressions and parameter definitions, provides a precise and feasible path for optimizing the parameters to be estimated in the offline supervised learning stage, ensuring that the parameter optimization process is orderly, controllable, and conforms to actual training needs. Specifically, this is achieved by setting regularization coefficients. and the corresponding regularization term This can prevent overfitting of the parameterized terminal cost function during training, improve the generalization ability of the parameterized terminal cost function under unknown interference conditions, and avoid the decrease in the adaptability of the parameterized terminal cost function due to overfitting of training samples.
[0055] Furthermore, this application clarifies the total number of training samples, training state samples, and the number of training samples. The parameters, such as the control input, perturbation estimation, and long-domain cost for each sample, enable the optimization problem to fully utilize various information from the expert dataset, aligning with the training requirements of supervised learning and ensuring the fitting accuracy of parameter optimization. This application determines the parameters to be estimated. belong d The definition of the 3D real vector space clarifies the parameters to be estimated. The data type and dimensions provide clear boundaries for parameter solving, adapting to the solution characteristics of SMP convex optimization and ensuring the smoothness of the solution process; in addition, the stage cost function It can quantify state tracking deviation and control energy consumption, and combine the parameterized terminal cost function value corresponding to the predicted state at the next moment. This allows the objective of the optimization problem to take into account both instantaneous operating performance and the comprehensive cost of subsequent states, ensuring that the parameters to be estimated obtained by optimization can make the terminal cost function both accurate and reasonable.
[0056] Based on this, the SMP convex optimization problem with descent constraints can accurately and efficiently optimize the parameters to be estimated, providing support for determining the optimal terminal cost function, ensuring the control performance of the subsequent terminal model predictive control (MPC) framework, and adapting to the time-varying disturbance control requirements of nonlinear unmanned systems.
[0057] In some embodiments, this application compares and verifies the method provided in this application with traditional reinforcement learning (RL) methods through simulation, as follows: (1) Baseline method Expert long-horizon MPC (Baseline): Prediction length T (e.g., T=50) generates training data and a performance baseline as an "expert".
[0058] Traditional reinforcement learning (RL) (based on Actor-Critic or DDPG / TD3) directly learns the policy: it is trained in the same environment, and the learned policy is deployed online as a comparison object.
[0059] Traditional short-domain MPC (no learned terminal cost): MPC that uses a prediction length N, such as N=1 or N=5, but without the cost of learning terminal.
[0060] (2) Simulation model, disturbance and parameters State dimension (n=6), relative position ,speed Input (m=3) triaxial thrust. Sampling period Ts=0.5s. The discretized system dynamics model is nonlinear, including a contact force coupling model, or a high-fidelity physical model is used.
[0061] Interference model (recurrent).
[0062] The number of training samples M is selected according to the scenario method, and the required samples are calculated based on the parameter dimension. This can be set. As the experimental group.
[0063] (3) Evaluation indicators 1. Online solution time (average / maximum / variance): Measure the average solution time of the MPC solver at each step.
[0064] 2. Convergence speed: The number of time steps or iterations required to reach a steady state or achieve the contact force target.
[0065] 3. Control performance: tracking error.
[0066] 4. Value approximation accuracy is used to characterize the accuracy with which the parameterized terminal cost function approximates the baseline cost obtained from long-domain MPC.
[0067] 5. Descent constraint empirical violation fraction and scenario-based theoretical value.
[0068] 6. Robustness (disturbance resistance): Evaluate the attenuation / fluctuation of the above indicators under multiple disturbance frequency / amplitude scenarios, and compare the performance degradation of different methods when the disturbance changes.
[0069] (4) Comparative experimental design Experimental Group 1 (This Invention): Obtained through offline scene synthesis Use short domain For example: N=1 or N=5.
[0070] Experimental group 2 (no terminal): Short domain MPC was used as a direct comparison.
[0071] Experimental Group 3 (Traditional RL Strategy): The policy / value was directly learned using RL and deployed as a control policy (or used as an initial value for MPC).
[0072] Experimental Group 4 (Experts): Long-range MPC(T) is used as the upper limit performance baseline.
[0073] Each group runs under the same batch of random initial conditions (covering the state space boundary) and the same perturbation samples. , For the number of trials, such as 50-200, record the evaluation indicators mentioned above, and calculate the mean and confidence interval.
[0074] (5) Expected / Reference Benchmark The terminal cost of learning proposed in this invention can reduce the average solution time of short domain (N=1) by about 20% compared with expert MPC, and approach the convergence of expert MPC (empirical violations decrease with the number of training samples).
[0075] Under time-varying recurrence perturbation conditions, by introducing a perturbation predictor and covering the perturbation phase in the training data, this method is expected to significantly outperform direct short-domain MPC in terms of solving time stability (variance) and closed-loop performance degradation, and outperform direct RL (the latter may degrade in performance or require more online retraining when the perturbation distribution changes).
[0076] Statistical tests: Compare the average solution time, RMSE, and convergence time of each group, and perform t-tests / Wilcoxon tests to prove significance.
[0077] In some embodiments, this application provides an MPC control system 300 inspired by terminal strategy optimization considering interference, such as... Figure 2 As shown, the MPC control system 300, inspired by the terminal strategy optimization considering interference, may include the following modules: The model building module 310 is used to construct a discrete-time nonlinear dynamic model with time-varying disturbances for the control process of nonlinear unmanned systems, so as to characterize the evolution of the system state vector with control input, time-varying disturbances, system white noise and modeling error. Prediction module 320 is used to introduce the disturbance predictor. g The parameters of the disturbance model are obtained by fitting historical disturbance data. ; in the k The step involves estimating the current disturbance through sensing or identification. And using a perturbation predictor g Predicting future steps i The disturbance was analyzed, and the disturbance prediction results were obtained. ; Parameter learning module 330 is used to obtain the parameterized terminal cost function through supervised learning. Optimal parameters The optimal terminal cost function is obtained. ; Framework building module 340 is used to build the terminal model predictive control (MPC) framework, form the corresponding terminal model predictive control (MPC) optimization problem, and determine the optimal terminal cost function. Embedded terminal MPC framework, replacing the cost of traditional terminals; Solver control module 350 is used to combine discrete-time nonlinear dynamics model and disturbance prediction results. At any moment k Solve the constrained terminal MPC optimization problem, output a control law to output control input to the unmanned system based on the control law, and complete online control.
[0078] According to embodiments of this application, any and multiple modules among the model building module 310, prediction module 320, parameter learning module 330, framework building module 340, and solution control module 350 can be merged into one module, or any one of these modules can be split into multiple modules. Alternatively, at least some of the functions of one or more of these modules can be combined with at least some of the functions of other modules and implemented in one module.
[0079] Figure 2 Each module in the system shown has the function of implementing each step in the MPC control method inspired by the terminal strategy optimization considering interference, and can achieve its corresponding technical effect. For the sake of brevity, it will not be elaborated here.
[0080] In some embodiments, this application provides an electronic device, the structural schematic of which is shown below. Figure 3 As shown.
[0081] The electronic device may include a processor 410 and a memory 420 storing computer program instructions.
[0082] Specifically, the processor 410 may include a central processing unit (CPU), an application-specific integrated circuit (ASIC), or one or more integrated circuits that can be configured to implement the embodiments of this application.
[0083] Memory 420 may include mass storage for data or instructions. For example, and not limitingly, memory 420 may include a hard disk drive (HDD), floppy disk drive, flash memory, optical disk, magneto-optical disk, magnetic tape, or Universal Serial Bus (USB) drive, or a combination of two or more of these. Where appropriate, memory 420 may include removable or non-removable (or fixed) media. Where appropriate, memory 420 may be internal or external to the integrated gateway disaster recovery device. In a particular embodiment, memory 420 is non-volatile solid-state memory.
[0084] Memory 420 may include read-only memory (ROM), random access memory (RAM), disk storage media devices, optical storage media devices, flash memory devices, electrical, optical, or other physical / tangible memory storage devices. Therefore, typically, memory 420 includes one or more tangible (non-transitory) computer-readable storage media (e.g., memory devices) encoded with software including computer-executable instructions, and when the software is executed (e.g., by one or more processors), it can perform the operations described in any of the interference-considered terminal strategy optimization-inspired MPC control methods in the above embodiments.
[0085] The processor 410 reads and executes computer program instructions stored in the memory 420 to implement any of the interference-considering terminal strategy optimization-inspired MPC control methods in the above embodiments.
[0086] In one example, the electronic device may also include a communication interface 430 and a bus 400. For example, Figure 3 As shown, the processor 410, memory 420, and communication interface 430 are connected via bus 400 and communicate with each other.
[0087] The communication interface 430 is mainly used to realize communication between various modules, devices, units and / or equipment in the embodiments of this application.
[0088] Bus 400 includes hardware, software, or both, that couples components of an online data traffic metering device together. For example, and not limitingly, the bus may include an Accelerated Graphics Port (AGP) or other graphics bus, an Enhanced Industry Standard Architecture (EISA) bus, a Front Side Bus (FSB), HyperTransport (HT) interconnect, an Industry Standard Architecture (ISA) bus, an Infinite Bandwidth Interconnect, a Low Pin Count (LPC) bus, a memory bus, a Microchannel Architecture (MCA) bus, a Peripheral Component Interconnect (PCI) bus, a PCI-Express (PCI-X) bus, a Serial Advanced Technology Attachment (SATA) bus, a Video Electronics Standards Association Local (VLB) bus, or other suitable buses, or combinations of two or more of these. Where appropriate, bus 400 may include one or more buses. Although specific buses are described and illustrated in embodiments of this application, any suitable bus or interconnect is contemplated herein.
[0089] Furthermore, in conjunction with the interference-considered terminal strategy optimization-inspired MPC control method in the above embodiments, this application embodiment can provide a computer storage medium for implementation. The computer storage medium stores computer program instructions; when these computer program instructions are executed by a processor, they implement any of the interference-considered terminal strategy optimization-inspired MPC control methods in the above embodiments.
[0090] It should be clarified that this application is not limited to the specific configurations and processes described above and shown in the figures. For the sake of brevity, detailed descriptions of known methods are omitted here. In the above embodiments, several specific steps are described and shown as examples. However, the method process of this application is not limited to the specific steps described and shown. Those skilled in the art can make various changes, modifications, and additions, or change the order of steps, after understanding the spirit of this application.
[0091] The functional blocks shown in the above block diagram can be implemented as hardware, software, firmware, or a combination thereof. When implemented in hardware, they can be, for example, electronic circuits, application-specific integrated circuits (ASICs), appropriate firmware, plug-ins, function cards, etc. When implemented in software, the elements of this application are programs or code segments used to perform the required tasks. Programs or code segments can be stored on a machine-readable medium or transmitted over a transmission medium or communication link via data signals carried on a carrier wave. "Machine-readable medium" can include any medium capable of storing or transmitting information. Examples of machine-readable media include electronic circuits, semiconductor memory devices, ROM, flash memory, erasable ROM (EROM), floppy disks, CD-ROMs, optical disks, hard disks, fiber optic media, radio frequency (RF) links, etc. Code segments can be downloaded via computer networks such as the Internet, intranets, etc.
[0092] It should also be noted that the exemplary embodiments mentioned in this application describe methods or systems based on a series of steps or apparatus. However, this application is not limited to the order of the above steps; that is, the steps can be performed in the order mentioned in the embodiments, or in a different order, or several steps can be performed simultaneously.
[0093] The aspects of this disclosure have been described above with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this disclosure. It should be understood that each block in the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that these instructions, executable via the processor of the computer or other programmable data processing apparatus, enable the implementation of the functions / actions specified in one or more blocks of the flowchart illustrations and / or block diagrams. Such a processor can be, but is not limited to, a general-purpose processor, a special-purpose processor, a special application processor, or a field-programmable logic circuit. It is also understood that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can also be implemented by special-purpose hardware performing the specified functions or actions, or can be implemented by a combination of special-purpose hardware and computer instructions.
[0094] The above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. An MPC control method inspired by terminal strategy optimization considering interference, characterized in that, include: To address the control process of nonlinear unmanned systems, a discrete-time nonlinear dynamic model with time-varying disturbances is constructed to characterize the evolution of the system state vector with control input, time-varying disturbances, system white noise, and modeling error. Introducing a disturbance predictor g The parameters of the disturbance model are obtained by fitting historical disturbance data. ; in the k The step involves estimating the current disturbance through sensing or identification. And using a perturbation predictor g Predicting future steps i The disturbance was analyzed, and the disturbance prediction results were obtained. ; The parameterized terminal cost function is obtained through supervised learning. Optimal parameters The optimal terminal cost function is obtained. ; Construct a terminal model predictive control (MPC) framework, form a corresponding terminal model predictive control (MPC) optimization problem, and apply the optimal terminal cost function. Embedded in the terminal MPC framework, replacing the traditional terminal cost; Combining discrete-time nonlinear dynamics model and disturbance prediction results At any moment k Solve the constrained terminal MPC optimization problem, output a control law, and output control input to the unmanned system based on the control law to complete online control.
2. The MPC control method inspired by terminal strategy optimization considering interference as described in claim 1, characterized in that, The discrete-time nonlinear dynamic model satisfies the expression: ; In the formula, k For discrete time steps, and They represent the first k Discrete time step and the first The system state vector at discrete time steps. To control the input, For modelable time-varying recurring disturbances, This indicates at least one of the system's white noise and modeling error; It is a nonlinear state transition function that characterizes the evolution of the system state; It is continuously differentiable on its domain and satisfies ; The system state vector includes at least one of relative position, velocity, attitude angle, and angular velocity; the control input includes at least one of thruster pulse, torque, and force distribution.
3. The MPC control method inspired by terminal strategy optimization considering interference as described in claim 2, characterized in that, The terminal model predictive control (MPC) framework satisfies the expression: ; In the formula, Indicates the current moment k For the future The predicted value made by the system state vector of the step; Indicates the current moment k For the future k+N The predicted value made by the system state vector of the step; N It is a positive integer; Indicates the current moment k For the future The predicted value made by the control input of the step; The parameter to be estimated; This is the stage cost function, used to characterize the time domain of prediction. Track the degree of deviation from the desired state and control input. The control resources consumed; The terminal cost function is parameterized and used to quantify the comprehensive cost of the system's future operation in the predicted end-time state. The optimal cost function is used to characterize the cost at the current discrete time step. k When the system is in state At that time, the optimal cumulative cost that can be achieved by solving the terminal model predictive control MPC optimization problem is determined.
4. The MPC control method inspired by terminal strategy optimization considering interference as described in claim 3, characterized in that, The terminal model predictive control (MPC) framework includes a state set. and input constraint set And introduce optimal parameters As a terminal item; The constraints of the terminal model predictive control (MPC) framework satisfy the expression: ; In the formula, ; For a set of states, For the input constraint set; For the disturbance prediction results, Indicates the current moment k For the future The predicted value made by the system state vector of the step; Represents the discrete time step at the current moment. k When making predictions about the future, the prediction state of the 0th prediction step corresponds to the starting point of the prediction time domain.
5. The MPC control method inspired by terminal strategy optimization considering interference as described in claim 4, characterized in that, The parameterized terminal cost function is obtained through supervised learning. Optimal parameters The optimal terminal cost function is obtained. ,include: Obtain an expert dataset covering interference samples from different times and phases. In the formula, For training state samples, The baseline control input obtained for long-domain MPC. Indicates baseline controller, for The corresponding long domain cost, The number of samples; Parameterized terminal cost function Let it be in the form of a linear parameterized basis function: In the formula, For the parameter to be estimated, Indicates transpose. For predefined basis function vectors; In the parameters to be estimated θ Solve the SMP convex optimization problem with descent constraints to obtain the optimal parameters. In order to determine the optimal terminal cost function The SMP convex optimization problem with descent constraints is used to optimize the parameterized terminal cost function during the offline supervised learning phase. The parameters to be estimated .
6. The MPC control method inspired by terminal strategy optimization considering interference as described in claim 5, characterized in that, The SMP convex optimization problem with descent constraints satisfies the expression: ; ; in, This is a regularization coefficient used to prevent overfitting of the parameterized terminal cost function. For regularization terms; The total number of training samples, These are training state samples; For the first Control input for each sample; For the first Perturbation estimation for each sample; for The corresponding long-domain cost; Indicates the parameter to be estimated belong d 3D real vector space; Represents the real number field; Let be the stage cost function, and characterize the . The instantaneous operating cost under each sample is used to quantify state tracking deviation and control energy consumption; Indicates the first Based on training state samples, for each training sample... Control input and disturbance estimation The calculated terminal cost function value corresponding to the predicted state at the next time step.
7. The MPC control method inspired by terminal strategy optimization considering interference as described in claim 5, characterized in that, The predefined basis function vector is at least one of radial basis function vector and orthogonal polynomial basis function vector; When the stage cost is quadratic, the SMP convex optimization problem with descent constraints is a convex quadratic constraint problem with a quadratic objective and linear inequality constraints, and is solved offline using a convex optimization solver.
8. An MPC control system inspired by terminal strategy optimization considering interference, characterized in that, include: The model building module is used to construct a discrete-time nonlinear dynamic model with time-varying disturbances for the control process of nonlinear unmanned systems, so as to characterize the evolution of the system state vector with control input, time-varying disturbances, system white noise and modeling error. Prediction module, used to introduce disturbance predictor g The parameters of the disturbance model are obtained by fitting historical disturbance data. ; in the k The step involves estimating the current disturbance through sensing or identification. And using a perturbation predictor g Predicting future steps i The disturbance was analyzed, and the disturbance prediction results were obtained. ; The parameter learning module is used to obtain the parameterized terminal cost function through supervised learning. Optimal parameters The optimal terminal cost function is obtained. ; The framework construction module is used to build the terminal model predictive control (MPC) framework, form the corresponding terminal model predictive control (MPC) optimization problem, and define the optimal terminal cost function. Embedded in the terminal MPC framework, replacing the traditional terminal cost; The solution control module is used to combine the discrete-time nonlinear dynamics model and the disturbance prediction results. At any moment k Solve the constrained terminal MPC optimization problem, output a control law, and output control input to the unmanned system based on the control law to complete online control.
9. An electronic device, characterized in that, include: A processor, a memory, and a program stored in the memory and executable on the processor, wherein the program, when executed by the processor, implements an MPC control method inspired by interference-considering terminal strategy optimization as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a program or instructions that, when executed by a processor, implement the MPC control method inspired by interference-considering terminal strategy optimization as described in any one of claims 1 to 7.