Optimal control method and device for nonlinear systems with unknown models

By building a data-driven model in a nonlinear system with unknown models, using higher-order ordinary differential equations and iterative algorithms, the problems of computational complexity and model dependence are solved, and efficient optimal control is achieved.

CN116382093BActive Publication Date: 2025-08-19BEIJING UNIV OF POSTS & TELECOMM
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202310559968.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-05-18
Publication Date
2025-08-19
Estimated Expiration
2043-05-18

AI Technical Summary

Technical Problem

The prior art has computational complexity problems (dimensional disasters) and dependence on precise system models when dealing with nonlinear systems with unknown models, making it difficult to design efficient optimal controllers.

Method used

Establish the optimal cost function of a nonlinear system, expand it into higher-order ordinary differential equations through partial differential equations, build a data-driven model, and determine the optimal control through iterative processing, and optimize it using differential dynamic programming algorithm.

Benefits of technology

It effectively solves the problem of dimensional disaster with large computations, improves the efficiency of nonlinear system control and algorithm convergence speed, and overcomes the time-varying behavior of the HJB equation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116382093B_ABST
    Figure CN116382093B_ABST
Patent Text Reader

Abstract

This paper provides an optimal control method and device for nonlinear systems with unknown models. For such a nonlinear system, an optimal cost function is established for the system, and a partial differential equation is determined to solve the optimal cost function. Based on an empirical dataset for the nonlinear system, the partial differential equation is expanded to obtain a higher-order ordinary differential equation. This higher-order ordinary differential equation is then applied to function approximation to obtain a data-driven model for the nonlinear system. Based on pre-defined constraints, the data-driven model is iteratively processed to determine the optimal control for the nonlinear system. This method effectively addresses the "curse of dimensionality" problem caused by large computational complexity, and the algorithm converges quickly, improving the efficiency of nonlinear system control.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This paper belongs to the field of control planning, specifically to optimal control methods and devices for nonlinear systems with unknown models. Background Art

[0002] Control theory in control systems engineering is a subfield of mathematics that deals with the control of continuously operating dynamic systems in engineering processes and machines. The goal is to develop control strategies that optimally apply control actions to control such systems without delay or overshoot while ensuring control stability.

[0003] For example, optimization-based control and estimation techniques such as model predictive control (MPC) allow for a model-based design framework in which system dynamics and constraints can be directly considered. MPC is used in many applications to control dynamic systems of various complexities. Examples of such systems include production lines, automotive engines, robotics, CNC machining, electric motors, satellites, and generators. However, in many cases, the model of the controlled system is nonlinear and may be difficult to design, use in real time, or may be inaccurate. Examples of such situations are common in robotics, building controls (HVAC), smart grids, factory automation, transportation, self-regulating machines, and traffic networks. In addition, even if a nonlinear model is fully available, designing an optimal controller is an inherently challenging task because of the need to solve partial differential equations known as the Hamilton-Jacobi-Bellman (HJB) equations.

[0004] Finding the optimal control law for general nonlinear systems requires solving the Hamilton-Jacobi-Bellman (HJB) partial differential equation, hereafter referred to as the HJB equation. While various traditional approaches exist for optimal control of dynamic systems with performance metrics or so-called cost functions, these approaches suffer from two drawbacks. On the one hand, solving the HJB equation is inherently computationally complex, increasing exponentially with the number of state dimensions—a phenomenon known as the "curse of dimensionality." On the other hand, traditional approaches rely on precise system models and are therefore inapplicable to difficult-to-model systems. Consequently, optimal control problems that are independent of mathematical models remain a hot topic of research. Summary of the Invention

[0005] In view of the above problems in the prior art, the purpose of this paper is to provide a method and device for optimal control of nonlinear systems with unknown models, which can improve the computational efficiency of optimal control of nonlinear systems.

[0006] In order to solve the above technical problems, the specific technical solutions of this article are as follows:

[0007] In one aspect, this paper provides an optimal control method for a nonlinear system with an unknown model, the method comprising:

[0008] For a nonlinear system with an unknown model, an optimal cost function of the system is established, and a partial differential equation for solving the optimal cost function is determined;

[0009] Expanding the partial differential equation according to the empirical data set of the nonlinear system to obtain a high-order ordinary differential equation;

[0010] Introducing the high-order ordinary differential equation into function approximation to obtain a data-driven model of the nonlinear system;

[0011] The data driven model is iteratively processed according to set constraints to determine the optimal control of the nonlinear system.

[0012] Furthermore, for a nonlinear system with an unknown model, an optimal cost function of the system is established, and a partial differential equation for solving the optimal cost function is determined, including:

[0013] Establish the state equation of a nonlinear system with unknown model: x(t0)=x0, where is the state variable of the system, is the control input of the system, is the system dynamic equation, is the system input state equation;

[0014] Determine an optimal cost function of the nonlinear system based on the dynamic constraints of the nonlinear system and the state equation of the nonlinear system: Among them, t∈[t0,t f ] and u[t,t f ] indicates that the control input u is limited to the time interval [t,t f ]Inside;

[0015] Determine the partial differential equation that solves the optimal cost function:

[0016] Furthermore, the partial differential equation is expanded based on the empirical data set of the nonlinear system to obtain a high-order ordinary differential equation, including:

[0017] Establishing the empirical data set based on historical input and output data of the nonlinear system;

[0018] According to the empirical data set and the partial differential equation, a saturated state equation and a saturated partial differential equation based on the state trajectory and the control input are determined as follows:

[0019] in is the initial value The state trajectory of It is defined as unknown optimal control;

[0020] According to the saturated state equation and the saturated partial differential equation, as well as the state equation and the cost function, a high-order ordinary differential equation of the optimal cost function is obtained through a differential dynamic programming algorithm, as shown below: in, S 12 =V xxg , S 22 =W uu , The boundary conditions are

[0021] Furthermore, the high-order ordinary differential equation is introduced into function approximation to obtain a data-driven model of the nonlinear system, including:

[0022] Determine the basis function that approximates the optimal cost function as follows:

[0023]

[0024] Determining an estimation function for optimal control based on the basis function;

[0025] The estimation function is introduced into the high-order ordinary differential equation to obtain a high-order differential dynamic approximation.

[0026] Furthermore, the estimation function is introduced into the high-order ordinary differential equation to obtain a high-order differential dynamic approximation, and then further includes:

[0027] Determining an algebraic matrix equation satisfied by the weight approximation based on the high-order differential dynamic approximation;

[0028] When the continuous excitation condition is met, the algebraic matrix equation is optimized to obtain a target algebraic matrix equation for calculating the weight approximation.

[0029] Furthermore, the data driven model is iteratively processed according to set constraints to determine the optimal control of the nonlinear system, including:

[0030] Define the approximate Hamiltonian and control input;

[0031] Based on the definition of the approximately estimated Hamiltonian and the control input, a differential dynamics iterative process is performed on the high-order differential dynamics approximation to determine the optimal control of the nonlinear system.

[0032] Furthermore, according to the definition of the approximate estimated Hamiltonian and the control input, a differential dynamics iterative process is performed on the high-order differential dynamics approximation to determine the optimal control of the nonlinear system, including:

[0033] Step 1: Set initial parameter values and calculate the initial cost function value, wherein the initial parameter values include at least the initial control input and the initial state variable;

[0034] Step 2: Calculate the iterative control input of the nonlinear system according to the initial parameter value;

[0035] Step 3: determining the iterative state variables of the nonlinear system according to the iterative control input and the state equation;

[0036] Step 4: Calculate the iterative weight approximation based on the target algebraic matrix equation and the number of iterations, and determine whether the first convergence condition is met. If not, return to step 3. If so, proceed to step 5.

[0037] Step 5: Based on the iterative state variable and the optimal cost function, calculate the iterative cost function value and determine whether the iterative cost function value meets the second convergence condition. If not, bring the iterative control input into step 2 to perform control input iteration. If not, determine the iterative control input as the target control input.

[0038] On the other hand, this article also provides an optimal control device for a nonlinear system with an unknown model, the device comprising:

[0039] A partial differential equation determination module is used to establish an optimal cost function for a nonlinear system with an unknown model and determine a partial differential equation for solving the optimal cost function;

[0040] a high-order ordinary differential equation determination module, configured to expand the partial differential equation according to the empirical data set of the nonlinear system to obtain a high-order ordinary differential equation;

[0041] a data-driven model determination module, configured to introduce the high-order ordinary differential equation into function approximation to obtain a data-driven model of the nonlinear system;

[0042] The optimal control module is used to iteratively process the data-driven model according to set constraints to determine the optimal control of the nonlinear system.

[0043] In another aspect, an optimal control device for a nonlinear system with an unknown model is provided, the device comprising:

[0044] an input interface configured to receive a state trajectory of the nonlinear system;

[0045] Memory;

[0046] A processor configured to execute the method described above and generate control instructions;

[0047] An output interface is configured to send the control command to an actuator of the nonlinear system to control the operation of the system.

[0048] Finally, this document provides a computer-readable storage medium storing a computer program. When the computer program is executed by a processor, the method described above is implemented.

[0049] Using the above technical solution, this paper describes a method and device for optimal control of nonlinear systems with unknown models. For a nonlinear system with an unknown model, an optimal cost function is established for the system and a partial differential equation is determined to solve the optimal cost function. Based on an empirical data set for the nonlinear system, the partial differential equation is expanded to obtain a higher-order ordinary differential equation. This higher-order ordinary differential equation is then introduced into function approximation to obtain a data-driven model of the nonlinear system. Based on pre-defined constraints, the data-driven model is iteratively processed to determine the optimal control for the nonlinear system. This paper effectively solves the "curse of dimensionality" problem caused by large computational complexity, and the algorithm converges quickly, improving the efficiency of nonlinear system control.

[0050] In order to make the above and other purposes, features and advantages of this article more obvious and easy to understand, the following specifically cites preferred embodiments and provides detailed descriptions in conjunction with the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0051] In order to more clearly illustrate the technical solutions in the embodiments of this article or the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of this article. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0052] Figure 1 A schematic overview of the principles used by some embodiments for controlling the operation of the system is shown.

[0053] Figure 2 A schematic diagram of the steps of the optimal control method for a nonlinear system with an unknown model provided in an embodiment of this invention is shown;

[0054] Figure 3 A flow chart of an optimal control method for a nonlinear system with unknown model provided in an embodiment of the present invention is shown;

[0055] Figure 4 The following diagrams show the state trajectory comparison of the system under initial control and optimal control inputs in one embodiment of this article;

[0056] Figure 5 A comparison diagram of the state trajectories of another embodiment of the present invention under initial control and optimal control inputs is shown;

[0057] Figure 6 The initial cost function V0 and the optimal cost function V of the system in this embodiment are shown. 17 ;

[0058] Figure 7 A schematic structural diagram of an optimal control device for a nonlinear system with an unknown model provided in an embodiment of this invention is shown;

[0059] Figure 8 A schematic structural diagram of the control device provided in the embodiments of this article is shown.

[0060] Description of the accompanying symbols:

[0061] 100, control device; 102, system; 104, model; 106, control instruction;

[0062] 701. Partial differential equation determination module; 702. High-order ordinary differential equation determination module; 703. Data-driven model determination module; 704. Optimal control module. DETAILED DESCRIPTION

[0063] The following will be combined with the accompanying drawings to clearly and completely describe the technical solutions in the embodiments of this document. Obviously, the embodiments described are only part of the embodiments of this document, not all of the embodiments. Based on the embodiments of this document, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this document.

[0064] It should be noted that the terms "first," "second," and the like in the specification and claims herein and in the accompanying drawings are used to distinguish similar objects and are not necessarily used to describe a particular order or precedence. It should be understood that the terms used in this manner are interchangeable where appropriate so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having," as well as any variations thereof, are intended to cover non-exclusive inclusions. For example, a process, method, apparatus, product, or device comprising a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units that are not explicitly listed or that are inherent to these processes, methods, products, or devices.

[0065] Figure 1 A schematic overview of the principles used by some embodiments for controlling the operation of a control system is shown. Some embodiments provide a control device 100 configured to control a system 102. For example, the device 100 can be configured to control a continuously operating dynamic system 102 in an engineering process or machine. Hereinafter, "control device" and "device" may be used interchangeably and have the same meaning. Hereinafter, "continuously operating dynamic system" and "system" may be used interchangeably and have the same meaning. Examples of systems 102 are HVAC systems, LIDAR systems, condensing units, production lines, self-tuning machines, smart grids, automotive engines, robots, CNC machining, electric motors, satellites, generators, transportation networks, and the like. Some embodiments are based on the recognition that the device 100 is developed to optimally use control actions to control control instructions 106 of the system 102 without delay or overshoot and ensuring control stability.

[0066] In some embodiments, the device 100 uses model-based and / or optimization-based control and estimation techniques, such as model predictive control (MPC), to develop control commands 106 for the system 102. Model-based techniques can be advantageous for controlling dynamic systems. For example, MPC allows for a model-based design framework in which the dynamics and constraints of the system 102 can be directly considered. MPC develops the control commands 106 based on a model 104 of the system. The model 104 of the system refers to a description of the dynamics of the system 102 using differential equations. In some embodiments, the model 104 is nonlinear and can be difficult to design and / or difficult to use in real time. For example, even if a nonlinear model is fully available, estimating the optimal control commands 106 is an inherently challenging task because it requires solving the partial differential equations (PDEs) describing the dynamics of the system 102, known as the Hamilton-Jacobi-Bellman (HJB) equations, which are computationally challenging.

[0067] Some embodiments use data-driven control techniques to design model 104. Data-driven techniques utilize operational data generated by system 102 in order to construct a feedback control strategy that stabilizes system 102.

[0068] Furthermore, the embodiments of this article provide an optimal control method for nonlinear systems with unknown models, which can effectively solve the "dimensionality curse" problem caused by large computational complexity, and the algorithm converges quickly. Figure 2 This is a schematic diagram of the steps of an optimal control method for a nonlinear system with an unknown model provided in the embodiment of this article. This specification provides the method operation steps described in the embodiment or flowchart, but may include more or fewer operation steps based on conventional or non-creative work. The order of steps listed in the embodiment is only one way of executing the order of many steps and does not represent the only execution order. When the actual system or device product is executed, it can be executed in the order or in parallel according to the method shown in the embodiment or the accompanying drawings. Specifically, Figure 2 As shown, the method may include:

[0069] S201: For a nonlinear system with an unknown model, establishing an optimal cost function of the system, and determining a partial differential equation for solving the optimal cost function;

[0070] S202: Expanding the partial differential equation according to the empirical data set of the nonlinear system to obtain a high-order ordinary differential equation;

[0071] S203: Introducing the high-order ordinary differential equation into function approximation to obtain a data-driven model of the nonlinear system;

[0072] S204: performing iterative processing on the data-driven model according to set constraints to determine optimal control of the nonlinear system.

[0073] It can be understood that for nonlinear systems, the partial differential equation (Hamilton-Jacobi-Bellman, HJB) is expanded into a higher-order ordinary differential equation, namely the (Differential Dynamic Programming, DDP) expansion, combined with the differential dynamic programming (DDP) technology. Then, function approximation is introduced into the DDP expansion to form an actor-critic structure and construct a data-driven model. Based on the data-driven model, a DDP iterative algorithm with strict convergence proof is developed. The new algorithm proposed in this patent overcomes technical barriers and solves the time-varying behavior of the HJB partial differential equation under the condition of a finite time domain cost function.

[0074] In the embodiment of this specification, for a nonlinear system with an unknown model, an optimal cost function of the system is established, and a partial differential equation for solving the optimal cost function is determined, including:

[0075] Establish the state equation of a nonlinear system with unknown model: x(t0)=x0, where is the state variable of the system, is the control input of the system, is the system dynamic equation, is the system input state equation;

[0076] Determine an optimal cost function of the nonlinear system based on the dynamic constraints of the nonlinear system and the state equation of the nonlinear system: Among them, t∈[t0,t f ] and u[t,t f ] indicates that the control input u is limited to the time interval [t,t f ]Inside;

[0077] Determine the partial differential equation that solves the optimal cost function: Illustratively, the nonlinear system can be as follows:

[0078]

[0079] in, is the state variable of the system, is the control input of the system, is the system dynamic equation, is the system input state equation. Assuming f(x)+g(x)u satisfies the Lipschitz continuity condition, is a closed bounded set of all saturated inputs, where γ is a constraint. For a fixed time interval T = [t0,t f ], we define the cost function associated with the system shown in formula (1) as:

[0080]

[0081] Where Q(x) is a positive definite function, W(u) is a nonnegative integrand, and τ is the independent variable of integration.

[0082] The goal of the optimal control problem is to design a constrained optimal control Make the cost function (2) satisfy:

[0083] J(x0,t0,u)≥J(x0,t0,u * ) (3)

[0084] Subject to the dynamic constraints in system (1), the following generalized non-quadratic function is used to deal with the input constraints:

[0085]

[0086] Among them, r i >0,i=1,2,…,m,is a positive weight factor.

[0087] Equation (4) can be rewritten into the following compact form:

[0088]

[0089] Where R = diag(r1, r2,…, r m ), v=(v1,v2,…,v m ) T , tanh -1 (v / γ)=(tanh -1 (v1 / γ),tanh -1 (v2 / γ),...,tanh -1 (v m / γ)) T .

[0090] The optimal control problem is described by the following optimal cost function:

[0091]

[0092] Among them, t∈[t0,t f ] and u[t,t f ] indicates that the control input u is limited to the time interval [t,t f ]. Assuming that V(x,t) is a first-order continuous differentiable function, the optimal cost function that satisfies the following HJB partial differential equation can be found:

[0093]

[0094] For all Optimal control strategy The control input u can be differentiated by the HJB equation to obtain the following:

[0095]

[0096] The Hamiltonian equation defining the optimal control is:

[0097] H(x,u,λ)=Q(x)+W(u)+λ T (f(x)+g(x)u) (9)

[0098] in, is a vector parameter. We can rewrite the HJB equation (7) as:

[0099]

[0100] In the embodiment of this specification, the partial differential equation is expanded based on the empirical data set of the nonlinear system to obtain a high-order ordinary differential equation, including:

[0101] Establishing the empirical data set based on historical input and output data of the nonlinear system;

[0102] According to the empirical data set and the partial differential equation, a saturated state equation and a saturated partial differential equation based on the state trajectory and the control input are determined as follows:

[0103] in is the initial value The state trajectory of It is defined as unknown optimal control;

[0104] According to the saturated state equation and the saturated partial differential equation, as well as the state equation and the cost function, a high-order ordinary differential equation of the optimal cost function is obtained through a differential dynamic programming algorithm, as shown below: in, S 12 =V xxg , The boundary conditions are

[0105] For example, a test control input is first selected make is the initial value For the initial value x0, we will is defined as the unknown optimal control. Therefore, any saturated input u(t)∈U,t∈[t0,t f The state trajectory x(t) under ] is calculated with parameters express:

[0106]

[0107] in, is the state error of the system, is the error of the control input. The state equation and HJB equation can be written as follows:

[0108]

[0109] Surround the above equation Expanding, we can get the following DDP expansion:

[0110] DDP expansion: Based on the state equation (1) and the cost function (2), let d i is a column vector The i-th element of G=((g1) x ,(g2) x ,...,(g m ) x ). Among them, g i is the i-th column vector of g,i=1,2,…,m. Then, the optimal cost function V and its partial derivative V x 、V xx Satisfy the following formula:

[0111]

[0112] in, S 12 =V xxg , S 22 =W uu , The boundary conditions are

[0113] It should be noted that in equations (14)-(16), the functions V, V x ,V xx ,f,f x ,g,G,Q,Q x ,Q xx ,W,W uu All in Evaluates, omitting parameters for brevity.

[0114] In the embodiment of this specification, the high-order ordinary differential equation is introduced into function approximation to obtain the data-driven model of the nonlinear system, including:

[0115] Determine the basis function that approximates the optimal cost function as follows:

[0116]

[0117] Determining an estimation function for optimal control based on the basis function;

[0118] The estimation function is introduced into the high-order ordinary differential equation to obtain a high-order differential dynamic approximation.

[0119] For example, in order to obtain the optimal cost function V when the system model f is unknown, an independent set of basis functions is used to approximate the unknown function. Define it as follows:

[0120]

[0121] in, and is the set of basis functions; and is the weight set; N a and N b is the number of basis functions in each set of basis functions; and is the approximation error. a and N b When they tend to infinity, the approximate error e ia and e b converges uniformly to zero. Since the exact values of the weights are unknown, the estimation function is defined as in is the set of weight estimates.

[0122] For simplicity, the following compact form is defined: Furthermore, using the symbols defined above, we can get the following compact expression:

[0123] Based on formula (8), we can get The estimation function is:

[0124]

[0125] make The above formula can be rewritten as:

[0126] Substituting the estimated function into the second-order expansion (14)-(15) or the third-order expansion (14)-(16), we can obtain the second-order or third-order DDP approximation.

[0127] In the embodiment of this specification, the estimation function is introduced into the high-order ordinary differential equation to obtain a high-order differential dynamic approximation, and then the following steps are further included:

[0128] Determining an algebraic matrix equation satisfied by the weight approximation based on the high-order differential dynamic approximation;

[0129] When the continuous excitation condition is met, the algebraic matrix equation is optimized to obtain a target algebraic matrix equation for calculating the weight approximation.

[0130] For example, without loss of generality, we provide a second-order DDP approximation as follows:

[0131] Second-order DDP approximation: Based on the second-order DDP expansion (14)-(15), the weight estimate is obtained and At any time interval The following algebraic matrix equation is satisfied:

[0132]

[0133] in,

[0134]

[0135]

[0136]

[0137]

[0138]

[0139]

[0140]

[0141] In addition, if the continuous excitation (PE) condition holds, that is, there is a constant ρ>0 and multiple time intervals Make Then we can get:

[0142]

[0143] Furthermore, the data driven model is iteratively processed according to the set constraints to determine the optimal control of the nonlinear system, including:

[0144] Define the approximate Hamiltonian and control input;

[0145] Based on the definition of the approximately estimated Hamiltonian and the control input, a differential dynamics iterative process is performed on the high-order differential dynamics approximation to determine the optimal control of the nonlinear system.

[0146] For example, the following definitions are given first:

[0147] Definition 2: The approximate Hamiltonian is defined as:

[0148]

[0149] Furthermore, the exact weights of the basis functions that define the system f(x) with unknown model are A* ,ie,f(x)=A * Ψ(x).

[0150] Definition 3: Assume that for any is a constraint operator.

[0151]

[0152] where γ is the constraint on the control input.

[0153] Furthermore, according to the definition of the approximate estimated Hamiltonian and the control input, a differential dynamics iterative process is performed on the high-order differential dynamics approximation to determine the optimal control of the nonlinear system, including:

[0154] Step 1: Set initial parameter values and calculate the initial cost function value, wherein the initial parameter values include at least the initial control input and the initial state variable;

[0155] Step 2: Calculate the iterative control input of the nonlinear system according to the initial parameter value;

[0156] Step 3: determining the iterative state variables of the nonlinear system according to the iterative control input and the state equation;

[0157] Step 4: Calculate the iterative weight approximation based on the target algebraic matrix equation and the number of iterations, and determine whether the first convergence condition is met. If not, return to step 3. If so, proceed to step 5.

[0158] Step 5: Based on the iterative state variable and the optimal cost function, calculate the iterative cost function value and determine whether the iterative cost function value meets the second convergence condition. If not, bring the iterative control input into step 2 to perform control input iteration. If not, determine the iterative control input as the target control input.

[0159] Exemplarily, the iterative process is as follows:

[0160] S301: Select initial value Convergence accuracy ∈> 0. Let x 0 (t),t∈[t0,t f ] is the initial given control input The corresponding state of the system satisfies the following formula:

[0161]

[0162] The initial control is and

[0163] Calculate the initial cost function And set i=0.

[0164] S302: Calculation and λ i , by solving the following formula:

[0165]

[0166] S303: Calculate x i+1 , by solving the following formula:

[0167]

[0168] And calculate the cost function

[0169] S304: Calculation and Based on the following formula:

[0170]

[0171] like Let k i >2k i And go to Step 3. Otherwise, go to Step 5.

[0172] S305: If J i+1 -J i ≥0, k i >2k i and go to Step 2. Otherwise, set And go to Step 2 until

[0173] S306: Setting the value function And set the control input

[0174] For example, in this embodiment, the nonlinear system is selected as follows:

[0175]

[0176] The basis functions of the system equations are chosen as:

[0177] The cost function is defined as:

[0178]

[0179] Where γ = 0.5, R = 1. The optimal cost function of the approximate system is:

[0180]

[0181] Develop a data-driven differential dynamic programming algorithm according to step S204, and set the initial condition of the system to x0 = [2, 1] T , the constraint of the control input is set to |u|<0.5.

[0182] From the optimal cost function At the beginning, after 17 iterations, the value of the weight vector is: Based on formula (24), the 17th optimal control input can be obtained as:

[0183]

[0184] like Figure 3 and Figure 4 As shown in FIG, they are the state trajectory comparison diagrams of the system under initial control and optimal control input in this embodiment, Figure 5 Then the initial cost function V0 and the optimal cost function V of the system in this embodiment are 17 .

[0185] Compared with the prior art, the embodiments of this specification have the following advantages and effects:

[0186] 1. This paper expands the HJB partial differential equation into a higher-order ordinary differential equation based on the differential dynamic programming algorithm, and constructs a new data-driven model;

[0187] 2. This invention effectively solves the "dimensionality curse" problem caused by large computational complexity, and the algorithm converges quickly;

[0188] 3. The present invention can overcome the technical obstacle of the time-varying behavior of the HJB equation caused by the finite time domain cost function.

[0189] On the basis of the above-mentioned method, the embodiment of this specification further provides an optimal control device for a nonlinear system with an unknown model, such as Figure 7 As shown, the square device includes:

[0190] A partial differential equation determination module 701 is used to establish an optimal cost function for a nonlinear system with an unknown model and determine a partial differential equation for solving the optimal cost function;

[0191] A high-order ordinary differential equation determination module 702 is configured to expand the partial differential equation according to the empirical data set of the nonlinear system to obtain a high-order ordinary differential equation;

[0192] A data-driven model determination module 703 is configured to introduce the high-order ordinary differential equation into function approximation to obtain a data-driven model of the nonlinear system;

[0193] The optimal control module 704 is configured to iteratively process the data-driven model according to set constraints to determine the optimal control of the nonlinear system.

[0194] The effects achieved by the above-mentioned device are consistent with the beneficial effects achieved by the above-mentioned method, and are not described in detail in the embodiments of this specification.

[0195] Furthermore, this specification also provides an optimal control device for a nonlinear system with an unknown model, the device comprising:

[0196] an input interface configured to receive a state trajectory of the nonlinear system;

[0197] Memory;

[0198] a processor configured to execute the method described above and generate control instructions;

[0199] An output interface is configured to send the control command to an actuator of the nonlinear system to control the operation of the system.

[0200] For example, its internal structure diagram can be as follows Figure 8 As shown, the device includes a processor, memory, and a network interface connected via a system bus. The processor of the device is used to provide computing and control capabilities. The memory of the device includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operating system and computer program stored in the non-volatile storage medium. The network interface of the device is used to communicate with an external terminal via a network connection. When executed by the processor, the computer program implements a method for identifying a travel surface covering in a computer device.

[0201] Those skilled in the art will understand that Figure 8 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the device to which the solution of the present application is applied. The specific device may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.

[0202] In one embodiment, a computer device is provided, including a memory and a processor. The memory stores a computer program, and the processor implements the steps in the above method embodiments when executing the computer program.

[0203] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps in the above-mentioned method embodiments are implemented.

[0204] In one embodiment, a computer program product is provided, including a computer program, which implements the steps in the above method embodiments when executed by a processor.

[0205] Those skilled in the art will appreciate that all or part of the processes in the above-mentioned embodiment methods can be implemented by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, database or other media used in the embodiments provided in this application may include at least one of non-volatile and volatile memory. Non-volatile memory may include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory may include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM). The database involved in the various embodiments provided herein may include at least one of a relational database and a non-relational database. Non-relational databases may include, but are not limited to, distributed databases based on blockchains. The processor involved in the various embodiments provided herein may be, but are not limited to, a general-purpose processor, a central processing unit, a graphics processing unit, a digital signal processor, a programmable logic unit, a data processing logic unit based on quantum computing, and the like.

[0206] It should also be understood that in the embodiments herein, the term "and / or" merely describes an association between associated objects, indicating that three possible relationships exist. For example, "A and / or B" could represent: A alone, A and B simultaneously, or B alone. Furthermore, the character " / " in this document generally indicates an "or" relationship between the associated objects.

[0207] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the above description has generally described the composition and steps of each example according to function. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this document.

[0208] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0209] In the several embodiments provided herein, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is merely a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interfaces, devices, or units, or can be an electrical, mechanical, or other form of connection.

[0210] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the embodiments herein.

[0211] This article uses specific embodiments to illustrate the principles and implementation methods of this article. The description of the above embodiments is only used to help understand the methods and core ideas of this article. At the same time, for those skilled in the art, based on the ideas of this article, there will be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as a limitation to this article.

Claims

1. An optimal control method for a nonlinear system with an unknown model, wherein the system is a continuously operating dynamic system, characterized in that: The method comprises: For a nonlinear system with an unknown model, an optimal cost function of the system is established, and a partial differential equation for solving the optimal cost function is determined; Expanding the partial differential equation according to the empirical data set of the nonlinear system to obtain a high-order ordinary differential equation; Introducing the high-order ordinary differential equation into function approximation to obtain a data-driven model of the nonlinear system; Iteratively processing the data-driven model according to set constraints to determine optimal control of the nonlinear system; The high-order ordinary differential equation is introduced into function approximation to obtain a data-driven model of the nonlinear system, including: determining a basis function that approximates the optimal cost function; Determining an estimation function for optimal control based on the basis function; Substituting the estimation function into the high-order ordinary differential equation to obtain a high-order differential dynamic approximation; The step of introducing the estimation function into the high-order ordinary differential equation to obtain a high-order differential dynamic approximation further includes: Determining an algebraic matrix equation satisfied by the weight approximation based on the high-order differential dynamic approximation; When the continuous excitation condition is met, the algebraic matrix equation is optimized to obtain a target algebraic matrix equation for calculating the weight approximation; The data-driven model is iteratively processed according to set constraints to determine the optimal control of the nonlinear system, including: Define the approximate Hamiltonian and control input; Based on the definition of the approximately estimated Hamiltonian and the control input, a differential dynamics iterative process is performed on the high-order differential dynamics approximation to determine the optimal control of the nonlinear system.

2. The method according to claim 1, characterized in that For a nonlinear system with an unknown model, an optimal cost function of the system is established, and a partial differential equation for solving the optimal cost function is determined, including: Establish the state equation of a nonlinear system with unknown model: ,in, is the state variable of the system, is the control input of the system, is the system dynamic equation, is the system input state equation; Determine an optimal cost function of the nonlinear system based on the dynamic constraints of the nonlinear system and the state equation of the nonlinear system: ,in, Indicates control input Limited to time intervals Inside; Determine the partial differential equation that solves the optimal cost function: .

3. The method according to claim 1, characterized in that According to the empirical data set of the nonlinear system, the partial differential equation is expanded to obtain a high-order ordinary differential equation, including: Establishing the empirical data set based on historical input and output data of the nonlinear system; According to the empirical data set and the partial differential equation, a saturated state equation and a saturated partial differential equation based on the state trajectory and the control input are determined as follows: ;in is the initial value The state trajectory of It is defined as unknown optimal control; According to the saturated state equation and the saturated partial differential equation, as well as the state equation and the cost function, a high-order ordinary differential equation of the optimal cost function is obtained by a differential dynamic programming algorithm, as shown below: ;in , , , , the boundary conditions are , , is the partial derivative of the optimal cost function V, The partial derivative of , H is the Manhattan value, Q is the positive definite function value, W is the non-negative integrand value, For unknown optimal control, function All in Evaluate.

4. The method according to claim 1, wherein The basis functions are expressed as follows: ,in, is an unknown function, is the cost function, x is the state variable, t is the time, is the set of basis functions; is the weight set; is the number of basis functions in each set of basis functions; is the approximation error.

5. The method according to claim 2, characterized in that Based on the definition of the approximate estimated Hamiltonian and the control input, a differential dynamics iterative process is performed on the high-order differential dynamics approximation to determine the optimal control of the nonlinear system, including: Step 1: Set initial parameter values and calculate the initial cost function value, wherein the initial parameter values include at least the initial control input and the initial state variable; Step 2: Calculate the iterative control input of the nonlinear system according to the initial parameter value; Step 3: determining the iterative state variables of the nonlinear system according to the iterative control input and the state equation; Step 4: Calculate the iterative weight approximation based on the target algebraic matrix equation and the number of iterations, and determine whether the first convergence condition is met. If not, return to step 3. If so, proceed to step 5. Step 5: Based on the iterative state variable and the optimal cost function, calculate the iterative cost function value and determine whether the iterative cost function value meets the second convergence condition. If not, bring the iterative control input into step 2 to perform control input iteration. If not, determine the iterative control input as the target control input.

6. An optimal control device for a nonlinear system with an unknown model, wherein the system is a continuously operating power system, characterized in that: The device comprises: A partial differential equation determination module is used to establish an optimal cost function for a nonlinear system with an unknown model and determine a partial differential equation for solving the optimal cost function; a high-order ordinary differential equation determination module, configured to expand the partial differential equation according to the empirical data set of the nonlinear system to obtain a high-order ordinary differential equation; a data-driven model determination module, configured to introduce the high-order ordinary differential equation into function approximation to obtain a data-driven model of the nonlinear system; an optimal control module, configured to iteratively process the data-driven model according to set constraints to determine an optimal control of the nonlinear system; The high-order ordinary differential equation is introduced into function approximation to obtain a data-driven model of the nonlinear system, including: determining a basis function that approximates the optimal cost function; Determining an estimation function for optimal control based on the basis function; Substituting the estimation function into the high-order ordinary differential equation to obtain a high-order differential dynamic approximation; The step of introducing the estimation function into the high-order ordinary differential equation to obtain a high-order differential dynamic approximation further includes: Determining an algebraic matrix equation satisfied by the weight approximation based on the high-order differential dynamic approximation; When the continuous excitation condition is met, the algebraic matrix equation is optimized to obtain a target algebraic matrix equation for calculating the weight approximation; The data-driven model is iteratively processed according to set constraints to determine the optimal control of the nonlinear system, including: Define the approximate Hamiltonian and control input; Based on the definition of the approximately estimated Hamiltonian and the control input, a differential dynamics iterative process is performed on the high-order differential dynamics approximation to determine the optimal control of the nonlinear system.

7. An optimal control device for a nonlinear system with an unknown model, characterized in that: The device comprises: an input interface configured to receive a state trajectory of the nonlinear system; Memory; A processor configured to execute the method according to any one of claims 1 to 5 and generate a control instruction; An output interface is configured to send the control instruction to an actuator of the nonlinear system to control the operation of the system.

8. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 5 is implemented.