A pseudospectral differential dynamic programming method and related apparatus

By using the pseudospectral differential dynamic programming method, based on the second-order expansion of the Hamilton-Jacobi-Bellman equation and the pseudospectral optimization, the problems of insufficient computational efficiency and iterative convergence in existing methods are solved, and a more efficient solution for nonlinear optimal control is achieved.

CN119689858BActive Publication Date: 2026-04-24BEIHANG UNIV
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
BEIHANG UNIV
Filing Date
2024-12-10
Publication Date
2026-04-24

AI Technical Summary

Technical Problem

Existing differential dynamic programming methods have limitations in terms of computational efficiency and iterative convergence, making it difficult to efficiently solve nonlinear optimal control problems.

Method used

A pseudospectral differential dynamic programming method is adopted. By constructing a second-order expansion based on the Hamilton-Jacobi-Bellman equation, the two-point boundary value problem is determined, and the pseudospectral method and Lagrange interpolation are used for iterative solution to optimize the trajectory update.

Benefits of technology

This improves the computational efficiency and iterative convergence of the differential dynamic programming method, ensuring the reliability and convergence of the iterative solution process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119689858B_ABST
    Figure CN119689858B_ABST
Patent Text Reader

Abstract

The application discloses a pseudospectral differential dynamic programming method and related devices, and relates to the technical field of nonlinear optimal control solution, in which, first, a nonlinear optimal guidance problem is constructed, and an optimal value function thereof satisfies an HJB equation; based on a specified reference trajectory, the HJB equation is second-order expanded to obtain a differential equation about the partial derivative of the optimal value function, then a first two-point boundary value problem is determined and solved to determine an optimal control variable variation, and based on the optimal control variable variation, a second two-point boundary value problem is determined to realize updating of the trajectory; finally, based on the pseudospectral method, the trajectory and a construction matrix are processed by using Lagrange interpolation, the second two-point boundary value problem is transformed and iteratively solved. The above scheme of the application is based on the second-order expansion of the Hamilton-Jacobi-Bellman equation, and the pseudospectral method is embedded into the differential dynamic programming method framework, and has better calculation efficiency and iterative convergence than the prior art.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of nonlinear optimal control solution technology, and in particular to a pseudospectral differential dynamic programming method and related apparatus. Background Technology

[0002] Differential dynamic programming is a commonly used method for solving nonlinear optimal control problems. Based on the Bellman optimality principle and the second-order expansion of the optimal value function, it transforms the nonlinear optimal control problem into a series of differential equation subproblems, and iteratively solves these differential equations to obtain the optimal trajectory and the locally optimal feedback guidance law. Existing differential dynamic programming methods mainly employ sweeping methods and multiple-shot methods to solve the differential equations, but these methods have limitations in terms of computational efficiency and numerical sensitivity. Summary of the Invention

[0003] The purpose of this application is to provide a pseudospectral differential dynamic programming method and related apparatus, which can improve the computational efficiency and iterative convergence of differential dynamic programming methods.

[0004] To achieve the above objectives, this application provides the following solution:

[0005] Firstly, this application provides a pseudospectral differential dynamic programming method, comprising the following steps:

[0006] Based on initial constraints, terminal constraints, and a dynamic model, a nonlinear optimal guidance problem is constructed; the optimal value function of the nonlinear optimal guidance problem satisfies the Hamilton-Jacobi-Bellman equation.

[0007] Based on the specified reference trajectory, the Hamilton-Jacobi-Bellman equation is expanded in second order to obtain a differential equation with respect to the partial derivative of the optimal value function.

[0008] Based on the differential equation of the partial derivative of the optimal value function, the derivative of the construction matrix is ​​obtained, and the first two-point boundary value problem is determined; the first two-point boundary value problem is a homogeneous form of the trajectory two-point boundary value problem.

[0009] Solve the first two-point boundary value problem to determine the optimal control variable variation, and determine the second two-point boundary value problem based on the optimal control variable variation. Solve the second two-point boundary value problem to update the trajectory.

[0010] Based on the pseudospectral method, Lagrange interpolation is used to process the trajectory and construction matrix, and the second two-point boundary value problem is transformed and iteratively solved to obtain the discrete optimal trajectory and discrete construction matrix.

[0011] Secondly, this application provides a pseudospectral differential dynamic programming apparatus, comprising the following modules:

[0012] The nonlinear optimal guidance problem construction module is used to construct nonlinear optimal guidance problems based on initial constraints, terminal constraints, and dynamic models; the optimal value function of the nonlinear optimal guidance problem satisfies the Hamilton-Jacobi-Bellman equation.

[0013] The module for determining the partial derivative differential equation of the value function is used to perform a second-order expansion of the Hamilton-Jacobi-Bellman equation based on a specified reference trajectory, so as to obtain the differential equation of the partial derivative of the optimal value function.

[0014] The module for determining the first two-point boundary value problem is used to obtain the derivative of the construction matrix based on the differential equation of the partial derivative of the optimal value function, and to determine the first two-point boundary value problem; the first two-point boundary value problem is a homogeneous form of the trajectory two-point boundary value problem.

[0015] The second two-point boundary value problem determination module is used to solve the first two-point boundary value problem, determine the optimal control variable variation, and determine the second two-point boundary value problem based on the optimal control variable variation; solving the second two-point boundary value problem realizes the trajectory update.

[0016] The pseudospectral method differential dynamic programming module is used to process the trajectory and construction matrix based on the pseudospectral method and Lagrange interpolation, transform the second two-point boundary value problem and solve iteratively to obtain the discrete optimal trajectory and discrete construction matrix.

[0017] Thirdly, this application provides a computer device, including: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the pseudospectral differential dynamic programming method described above.

[0018] Fourthly, this application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the pseudospectral differential dynamic programming method described above.

[0019] Fifthly, this application provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the pseudospectral differential dynamic programming method described above.

[0020] According to the specific embodiments provided in this application, the following technical effects are disclosed:

[0021] This application provides a pseudospectral differential dynamic programming method and related apparatus. In this method, firstly, a nonlinear optimal guidance problem is constructed based on initial constraints, terminal constraints, and a dynamic model. The optimal value function satisfies the Hamilton-Jacobi-Bellman equation. Based on a specified reference trajectory, the Hamilton-Jacobi-Bellman equation is expanded in second order to obtain a differential equation with partial derivatives of the optimal value function. Then, the first two-point boundary value problem is determined and solved to determine the optimal control variation. Based on the optimal control variation, the second two-point boundary value problem is determined, and solving the second two-point boundary value problem updates the trajectory. Finally, based on the pseudospectral method, Lagrange interpolation is used to process the trajectory and construction matrix. The second two-point boundary value problem is transformed and iteratively solved to obtain a discrete optimal trajectory and a discrete construction matrix. The above-mentioned scheme in this application is based on the second-order expansion of the Hamilton-Jacobi-Bellman equation. Unlike the existing differential dynamic programming method, it constructs two two-point boundary value problems with the same homogeneous solution, which are used to calculate the variational control quantity and perform trajectory update, respectively. Furthermore, it embeds the pseudospectral method into the differential dynamic programming framework, which has better computational efficiency and iterative convergence. Attached Figure Description

[0022] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0023] Figure 1 This is a flowchart of a pseudospectral differential dynamic programming method provided in an embodiment of this application.

[0024] Figure 2 This is a schematic diagram of the x-axis position curve during the digital simulation iteration of a pseudospectral differential dynamic programming method provided in an embodiment of this application.

[0025] Figure 3 This is a schematic diagram of the y-axis position curve during the digital simulation iteration of a pseudospectral differential dynamic programming method provided in an embodiment of this application.

[0026] Figure 4 This is a schematic diagram of the z-axis position curve during the digital simulation iteration of a pseudospectral differential dynamic programming method provided in an embodiment of this application.

[0027] Figure 5 This is a schematic diagram of the x-axis velocity curve during the digital simulation iteration of a pseudospectral differential dynamic programming method provided in an embodiment of this application.

[0028] Figure 6This is a schematic diagram of the y-axis velocity curve during the digital simulation iteration of a pseudospectral differential dynamic programming method provided in an embodiment of this application.

[0029] Figure 7 This is a schematic diagram of the z-axis velocity curve during the digital simulation iteration of a pseudospectral differential dynamic programming method provided in an embodiment of this application.

[0030] Figure 8 This is a schematic diagram of the pitch angle curve during the digital simulation iteration of a pseudospectral differential dynamic programming method provided in an embodiment of this application.

[0031] Figure 9 This is a schematic diagram of the yaw angle curve during the digital simulation iteration of a pseudospectral differential dynamic programming method provided in an embodiment of this application.

[0032] Figure 10 This is a schematic diagram of the pitch rate curve during the digital simulation iteration of a pseudospectral differential dynamic programming method provided in an embodiment of this application.

[0033] Figure 11 This is a schematic diagram of the yaw rate curve during the digital simulation iteration of a pseudospectral differential dynamic programming method provided in an embodiment of this application.

[0034] Figure 12 This is a schematic diagram of the throttling rate curve during the digital simulation iteration of a pseudospectral differential dynamic programming method provided in an embodiment of this application.

[0035] Figure 13 This is a schematic diagram of the quality curve during the digital simulation iteration process of a pseudospectral differential dynamic programming method provided in an embodiment of this application.

[0036] Figure 14 This is a schematic diagram of the functional modules of a pseudospectral differential dynamic programming device provided in an embodiment of this application.

[0037] Figure 15 This is a schematic diagram of the structure of a computer device provided in an embodiment of this application. Detailed Implementation

[0038] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0039] To make the above-mentioned objectives, features and advantages of this application more apparent and understandable, the application will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0040] This application provides a pseudospectral differential dynamic programming method, such as... Figure 1 As shown, it includes the following steps:

[0041] A1. Based on initial constraints, terminal constraints, and a dynamic model, a nonlinear optimal guidance problem is constructed. The optimal value function of the nonlinear optimal guidance problem satisfies the Hamilton-Jacobi-Bellman equation (HJB equation). In this embodiment, the nonlinear optimal guidance problem is shown in the following equation:

[0042]

[0043] Where J is the overall performance index; φ[x(1),a] is the terminal performance index; Here, a represents the integral performance index; τ represents the total flight time; τ represents the normalized time between 0 and 1; x(τ) represents the state variable corresponding to the normalized time τ; u(τ) represents the control variable corresponding to the normalized time τ; x(1) represents the terminal state variable; c f f[τ,x(τ),u(τ),a] represents the terminal constraint value; f[τ,x(τ),u(τ),a] represents the dynamic model; x(0) represents the status bar corresponding to the initial time; x0 represents the given initial time constraint value; and C represents the terminal constraint matrix.

[0044] The optimal value function for the nonlinear optimal guidance problem is shown in the following equation:

[0045]

[0046] Where V[τ,x(τ),a,v] is the optimal value function; Φ[x(1),a,v] is the terminal augmentation performance index containing terminal performance indicators and terminal constraints, Φ[x(1),a,v]=φ[x(1),a]+v T [Cx(1)-c f v is the multiplier vector corresponding to the terminal constraint; s is the normalized time between τ and 1. According to the Bellman optimality principle, V[τ,x(τ),a,v] satisfies the Hamilton-Jacobi-Bellman equation.

[0047] A2. Based on the specified reference trajectory, the Hamilton-Jacobi-Bellman equation is expanded in second order to obtain a differential equation with respect to the partial derivatives of the optimal value function. In this embodiment, the partial derivatives with respect to the optimal value function include the partial derivatives of the optimal value function with respect to the normalized time τ, the partial derivatives of the optimal value function with respect to the state variable x, and the partial derivatives of the optimal value function with respect to the total flight time a.

[0048]

[0049] At time τ = 1, V[1,x(1),a,v] = Φ[x(1),a,v] satisfies the following equation:

[0050]

[0051] Among them, V τ [τ,x(τ),a,v] represents the partial derivative of the optimal value function with respect to normalized time τ; H[τ,x(τ),u(τ),a,v] represents the Hamiltonian function; V x [1,x(1),a,v] represents the partial derivative of the optimal value function with respect to the state variable x; the superscript T indicates the transpose; V a [1,x(1),a,v] represents the partial derivative of the optimal value function with respect to the total flight time a. The Hamiltonian function H[τ,x(τ),u(τ),a,v] can be expressed as follows:

[0052] H[τ,x(τ),u(τ),a,v]=L[τ,x(τ),u(τ),a]+V x [τ,x(τ),a,ν]f[τ,x(τ),u(τ),a].

[0053] At time τ = 0, the initial state is given, so the total flight time a should be chosen to minimize V[0, x(0), a], resulting in the following equation:

[0054]

[0055] Given a fixed x r (τ), u r (τ) and a r Select reference trajectory The Hamiltonian function H can be expanded to second order as follows:

[0056]

[0057] in, The horizontal line above the symbol represents the value of the variable along the reference trajectory; The subscript in the text represents the partial derivative with respect to the corresponding variable.

[0058] V τ The left-hand side of (τ) expands to the second order, as shown in the following equation:

[0059]

[0060] in, The subscript in the text represents the partial derivative with respect to the corresponding variable.

[0061] Substituting the above equation into the HJB equation, we obtain the following differential equation:

[0062]

[0063]

[0064] Based on the above differential equation, the variational form of the optimal control quantity is initially determined as follows:

[0065]

[0066] The parameter with a wavy line is defined as follows:

[0067]

[0068]

[0069] A3. Based on the differential equation of the partial derivative of the optimal value function, obtain the derivative of the constructed matrix and determine the first two-point boundary value problem; the first two-point boundary value problem is the homogeneous form of the trajectory two-point boundary value problem.

[0070] The feedback control law for the aforementioned nonlinear optimal guidance problem can be expressed by the following equation:

[0071]

[0072] To obtain the optimal feedback control law, the Riccati equation needs to be solved. First, based on the terminal conditions, the following matrix is ​​defined:

[0073]

[0074] Matrices X(τ), A, Λ(τ), M(τ), and P are all construction matrices that satisfy the following equation at the terminal time:

[0075]

[0076] Differentiating Λ(τ) and M(τ) respectively, we obtain the following equation:

[0077]

[0078] Based on the differential equation obtained in step A2, the derivatives of matrices X(τ), Λ(τ), and M(τ) can be obtained, expressed as:

[0079]

[0080] Select the current time τ c The boundary conditions are as follows:

[0081] X(τ c )=I,M(τc ) = 0.

[0082] In summary, the first two-point boundary value problem is determined. In this embodiment, the first two-point boundary value problem is as follows:

[0083]

[0084] Where Λ(τ), X(τ), M(τ), P, and A are the construction matrices; the points on the signs represent differentiation; I is the identity matrix; τ c The current normalized time; For φ x The partial derivative of the state quantity along the reference trajectory; For φ x The partial derivative of the total flight time a along the reference trajectory; φ x The partial derivative of the terminal performance index with respect to the state variable x; For φ a The partial derivative of the state quantity along the reference trajectory; For φ a The partial derivative of the total flight time a along the reference trajectory; φ a Let be the partial derivative of the terminal performance index with respect to the total flight time a; Λ(1), X(1), and M(1) are the values ​​of Λ(τ), X(τ), and M(τ) at the terminal time. It can be seen that the above two-point boundary value problem is a homogeneous form of the trajectory two-point boundary value problem.

[0085] A4. Solve the first two-point boundary value problem to determine the optimal control variable variation, and determine the second two-point boundary value problem based on the optimal control variable variation. Solve the second two-point boundary value problem to update the trajectory.

[0086] Solving this first two-point boundary value problem yields the construction matrices X(τ), A, Λ(τ), M(τ), and P, where X(τ) represents the trajectory of the state variable x(τ) with respect to the current state variable x. c The sensitivity of A is the total flight time a relative to the current state variable x. c Sensitivity; Λ(τ) represents the trajectory V x (τ) for the current state variable x c Sensitivity; M(τ) represents the trajectory V a (τ) for the current state variable x c Sensitivity; P represents the multiplier v's sensitivity to the current state variable x. c The sensitivity of the control quantity. Therefore, in this embodiment, the optimal control quantity variation is as follows:

[0087]

[0088] Where δu(τ) is the variational value of the optimal control quantity; This represents the value of the partial derivative of the Hamiltonian function with respect to the control quantity along the reference trajectory; For H u (τ) represents the value of the partial derivative of the control quantity along the reference trajectory; the superscript -1 indicates the inverse of the matrix; This represents the value of the partial derivative of the dynamic model with respect to the control quantity along the reference trajectory; For V x (τ) is the value of the partial derivative of the state quantity along the reference trajectory; V x (τ) represents the partial derivative of the optimal value function with respect to the state quantity; Λ(τ), X(τ), and A are the construction matrices; State variables for the reference trajectory; For V x (τ) is the value of the partial derivative of the total flight time a along the reference trajectory.

[0089] Next, V x (τ) and V a (τ) respectively along the reference Expanding to the first order, we get:

[0090]

[0091] Differentiating the above equation, we get:

[0092]

[0093] In the formula: It can be represented as:

[0094]

[0095] Substituting the optimal control variation into the given values, we get:

[0096]

[0097] Variation of optimal control quantity Substitution and From the formula, we get:

[0098]

[0099]

[0100] Reference Given the trajectory, V x The perturbation value of (τ) can be expressed as:

[0101]

[0102] The terminal conditions will be aligned with the reference. and Expanding to the first order, we get:

[0103]

[0104] The second two-point boundary value problem can be determined based on the above formula, as shown in the following formula:

[0105]

[0106] Where, x c V is the current state variable; x (1) is V x The value of [1,x(1),a,v] at the terminal time; V a (1) is V a The value of [1,x(1),a,v] at the terminal time; G1(τ), G2(τ), G3(τ), G4(τ), G5(τ), G6(τ), g1(τ), g2(τ), g3(τ), p1(τ) and p2(τ) are all intermediate variables, defined by the following formulas:

[0107]

[0108]

[0109] Solving the second two-point boundary value problem obtained above can realize the solution to the trajectory x(τ), V a Updates to (τ), a, and v.

[0110] A5. Based on the pseudospectral method, Lagrange interpolation is used to process the trajectory and construction matrix, and the second two-point boundary value problem is transformed and iteratively solved to obtain the discrete optimal trajectory and discrete construction matrix.

[0111] Introducing normalized time in the interval [-1,1] It satisfies the following formula:

[0112]

[0113] Using Lagrange interpolation, x(τ), V a The following approximations are made for (τ), X(τ), Λ(τ), and M(τ):

[0114]

[0115] in, These are Lagrange interpolation basis functions. The Lagrange interpolation nodes contain LGR integration points. and boundary nodes The two-point boundary value problem can then be transformed into the following equation:

[0116]

[0117]

[0118] Among them, D ki Let be the element in the k-th row and i-th column of the differential matrix D.

[0119] The two-point boundary value problem can be transformed into the following form:

[0120]

[0121] The nonhomogeneous and homogeneous solutions to the two-point boundary value problem are compared using the following formula:

[0122]

[0123] The above equations can then be expressed as follows:

[0124] A s Y s0 =B s .

[0125] A s Y s =B s x c +B s .

[0126] Solving the above equation yields the following expression:

[0127]

[0128] Y is solved iteratively. s To obtain the discrete optimal trajectory a and v, and solve Y in the last iteration. s0 This yields a discrete matrix. A and P. Continuous trajectory x(τ), V a (τ), a and v, as well as matrices X(τ), Λ(τ), M(τ), A and P, can be expressed in analytical form using Lagrange interpolation.

[0129] Existing differential dynamic programming methods focus on iterative problem solving, lacking quantification and analysis of iterative convergence, and therefore cannot guarantee reliable convergence. However, since the numerical computation of the iterative solution is carried out within a pseudospectral framework in the scheme proposed in the above embodiments of this application, an analytical convergence criterion is given as follows:

[0130] Suppose the solution obtained in the i-th iteration is as shown in the following equation:

[0131]

[0132] The (i+1)th guidance calculation process can be represented by the following simple iterative form:

[0133]

[0134] in, Let A be the linear equation solved in the (i+1)th iteration. S B S and b S This is the coefficient matrix that makes up the above linear equation.

[0135] To describe the convergence of the iterative solution process, an iterative convergence index is defined as a matrix. The largest eigenvalue is used to determine the iterative convergence index using the following formula:

[0136]

[0137] Where, ρ c λ is the iterative convergence index; max It is the largest eigenvalue; This is the solution obtained in the i-th iteration.

[0138] Based on the convergence of simple iteration, it can be known that when ρ c When ρ < 1, the iterative solution process converges. Therefore, ρ can be used. c The convergence of the iterative solution process, i.e., ρ c The smaller the value, the stronger the convergence.

[0139] Since iterative convergence is mainly related to the initial conjecture and the dynamic nonlinearity, when ρ c When ρ > 1, it can be reduced by modifying the dynamic model. c In the early stages of iteration, when the initial conjecture deviates significantly from the optimal trajectory, key parameters in the dynamic model are corrected to reduce nonlinearity, thereby reducing ρ. c In the later stages of iteration, the reference trajectory approaches the optimal trajectory, while ensuring ρ c Under the premise of <1, the dynamic model is gradually restored to solve the original problem.

[0140] Next, a digital simulation case will be used to demonstrate the effectiveness of the proposed solution, not an actual flight mission. The simulation scenario will be a flight mission during the descent phase of a reusable rocket of a certain model. Using the pseudospectral differential dynamic programming method of this application, the iterative optimization process of the x-axis position curve after digital simulation is as follows: Figure 2 As shown, the iterative optimization process of the y-axis position curve is as follows: Figure 3 As shown, the iterative optimization process of the z-axis position curve is as follows: Figure 4 As shown, the iterative optimization process of the x-axis velocity curve is as follows: Figure 5 As shown, the iterative optimization process of the y-axis velocity curve is as follows: Figure 6 As shown, the iterative optimization process of the z-axis velocity curve is as follows: Figure 7 As shown, the iterative optimization process of the pitch angle curve is as follows: Figure 8 As shown, the iterative optimization process of the yaw angle curve is as follows: Figure 9 As shown, the iterative optimization process of the pitch rate curve is as follows: Figure 10 As shown, the iterative optimization process of the yaw rate curve is as follows: Figure 11 As shown, the iterative optimization process of the throttling rate curve is as follows: Figure 12 As shown, the iterative optimization process of the quality curve is as follows: Figure 13 As shown.

[0141] The method provided in the above embodiments of this application is based on the second-order expansion of the Hamilton-Jacobi-Bellman equation. Unlike existing differential dynamic programming methods, it constructs two two-point boundary value problems with the same homogeneous solution, which are used to calculate the variational control variable and to update the trajectory, respectively. Furthermore, it embeds the pseudospectral method into the differential dynamic programming framework, resulting in better computational efficiency and iterative convergence. In addition, based on the fixed-point iterative convergence theorem, a quantized convergence criterion for the pseudospectral differential dynamic programming method is proposed, and a model correction method is designed based on this quantized convergence criterion to improve the reliability of algorithm convergence.

[0142] Based on the same inventive concept, this application also provides an apparatus for implementing the pseudospectral differential dynamic programming method described above. The solution provided by this apparatus is similar to the implementation scheme described in the above method; therefore, the specific limitations in one or more pseudospectral differential dynamic programming apparatus embodiments provided below can be found in the limitations of the pseudospectral differential dynamic programming method described above, and will not be repeated here.

[0143] In one exemplary embodiment, such as Figure 14 As shown, a pseudospectral differential dynamic programming device is provided, comprising the following modules:

[0144] The nonlinear optimal guidance problem construction module is used to construct nonlinear optimal guidance problems based on initial constraints, terminal constraints, and dynamic models; the optimal value function of the nonlinear optimal guidance problem satisfies the Hamilton-Jacobi-Bellman equation.

[0145] The module for determining the partial derivative differential equation of the value function is used to perform a second-order expansion of the Hamilton-Jacobi-Bellman equation based on a specified reference trajectory to obtain the differential equation of the partial derivative of the optimal value function.

[0146] The module for determining the first two-point boundary value problem is used to obtain the derivative of the construction matrix based on the differential equation of the partial derivative of the optimal value function, and to determine the first two-point boundary value problem; the first two-point boundary value problem is a homogeneous form of the trajectory two-point boundary value problem.

[0147] The second two-point boundary value problem determination module is used to solve the first two-point boundary value problem, determine the optimal control variable variation, and determine the second two-point boundary value problem based on the optimal control variable variation; solving the second two-point boundary value problem realizes the trajectory update.

[0148] The pseudospectral method differential dynamic programming module is used to process the trajectory and construction matrix based on the pseudospectral method and Lagrange interpolation, transform the second two-point boundary value problem and solve iteratively to obtain the discrete optimal trajectory and discrete construction matrix.

[0149] certainly, Figure 14 The architecture shown is merely exemplary; it can be omitted as needed when implementing different functionalities. Figure 14 One or at least two components of the system shown.

[0150] In one exemplary embodiment, a computer device is provided, which may be a server or a terminal, and its internal structure diagram may be as follows. Figure 15 As shown, the computer device includes a processor, memory, input / output (I / O) interfaces, and a communication interface. The processor, memory, and I / O interfaces are connected via a system bus, and the communication interface is also connected to the system bus via the I / O interfaces. The processor provides computational and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and a database. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The I / O interfaces are used for exchanging information between the processor and external devices. The communication interface is used for communicating with external terminals via a network connection. When the computer program is executed by the processor, it can implement a pseudospectral differential dynamic programming method provided in the above embodiment.

[0151] Those skilled in the art will understand that Figure 15 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0152] In one exemplary embodiment, a computer device is also provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps in the above-described method embodiments.

[0153] In one exemplary embodiment, a computer-readable storage medium is provided storing a computer program that, when executed by a processor, implements the steps in the above-described method embodiments.

[0154] In one exemplary embodiment, a computer program product is provided, including a computer program that, when executed by a processor, implements the steps in the above-described method embodiments.

[0155] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of the relevant data must comply with relevant regulations.

[0156] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM).

[0157] The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, etc., and are not limited to these.

[0158] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0159] This document uses specific examples to illustrate the principles and implementation methods of this application. The descriptions of the above embodiments are only for the purpose of helping to understand the methods and core ideas of this application. Furthermore, those skilled in the art will recognize that, based on the ideas of this application, there will be changes in the specific implementation methods and application scope. Therefore, the content of this specification should not be construed as a limitation of this application.

Claims

1. A pseudospectral differential dynamic programming method, characterized in that, include: Based on initial constraints, terminal constraints, and a dynamic model, a nonlinear optimal guidance problem is constructed; the optimal value function of the nonlinear optimal guidance problem satisfies the Hamilton-Jacobi-Bellman equation. Based on the specified reference trajectory, the Hamilton-Jacobi-Bellman equation is expanded in second order to obtain a differential equation with respect to the partial derivatives of the optimal value function; the partial derivatives of the optimal value function include the partial derivatives of the optimal value function with respect to the normalized time τ, the partial derivatives of the optimal value function with respect to the state variable x, and the partial derivatives of the optimal value function with respect to the total flight time a. Based on the differential equation relating the partial derivatives of the optimal value function, the derivative of the construction matrix is ​​obtained, and the first two-point boundary value problem is determined. The first two-point boundary value problem is a homogeneous form of the trajectory two-point boundary value problem. The first two-point boundary value problem is shown in the following equation: Where Λ(τ), X(τ), M(τ), P, and A are the construction matrices; the points on the signs represent differentiation; I is the identity matrix; τ c The current normalized time; For φ x The partial derivative of the state quantity along the reference trajectory; For φ x The partial derivative of the total flight time a along the reference trajectory; φ x The partial derivative of the terminal performance index with respect to the state variable x; For φ a The partial derivative of the state quantity along the reference trajectory; For φ a The partial derivative of the total flight time a along the reference trajectory; φ a Let be the partial derivative of the terminal performance index with respect to the total flight time a; C is the terminal constraint matrix, and Λ(1), X(1) and M(1) are the values ​​of Λ(τ), X(τ) and N(τ) at the terminal time. The second two-point boundary value problem is shown in the following equation: Where, x c V is the current state variable; x (1) is V x The value of [1,x(1),a,v] at the terminal time; V a (1) is V a The value of [1,x(1),a,v] at the terminal time; G1(τ), G2(τ), G3(τ), G4(τ), G5(τ), G6(τ), g1(τ), g2(τ), g3(τ), p1(τ) and p2(τ) are all intermediate variables; Solve the first two-point boundary value problem to determine the optimal control variable variation, and determine the second two-point boundary value problem based on the optimal control variable variation. Solve the second two-point boundary value problem to update the trajectory; the optimal control variable variation is shown in the following equation: Where δu(τ) is the variational value of the optimal control quantity; This represents the value of the partial derivative of the Hamiltonian function with respect to the control quantity along the reference trajectory; For H u (τ) represents the value of the partial derivative of the control quantity along the reference trajectory; the superscript -1 indicates the inverse of the matrix; This represents the value of the partial derivative of the dynamic model with respect to the control quantity along the reference trajectory; For V x (τ) is the value of the partial derivative of the state quantity along the reference trajectory; V x (τ) represents the partial derivative of the optimal value function with respect to the state variables; Λ(τ), X(τ), and A are the construction matrices; δx(τ) is... State variables for the reference trajectory; For V x (τ) is the value of the partial derivative of the total flight time a along the reference trajectory; Based on the pseudospectral method, Lagrange interpolation is used to process the trajectory and construction matrix, transforming the second two-point boundary value problem and solving iteratively to obtain the discrete optimal trajectory and discrete construction matrix.

2. The pseudospectral differential dynamic programming method according to claim 1, characterized in that, The nonlinear optimal guidance problem is shown in the following equation: Where J is the overall performance index; φ[x(1),a] is the terminal performance index; Here, a represents the integral performance index; τ represents the total flight time; τ represents the normalized time between 0 and 1; x(τ) represents the state variable corresponding to the normalized time τ; u(τ) represents the control variable corresponding to the normalized time τ; x(1) represents the terminal state variable; c f Let f[τ, x(τ), u(τ), a] be the given terminal constraint value; f[τ, x(τ), u(τ), a] be the dynamic model; x(0) be the status bar corresponding to the initial time; x0 be the given initial time constraint value; C be the terminal constraint matrix; the optimal value function of the nonlinear optimal guidance problem is shown in the following equation: Where V[τ,x(τ),a,ν] is the optimal value function; Φ[x(1),a,ν]=φ[x(1),a]+v T [Cx(1)-c f v is the multiplier vector corresponding to the terminal constraint; s is the normalized time between τ and 1.

3. The pseudospectral differential dynamic programming method according to claim 2, characterized in that, The partial derivatives of the optimal value function include the partial derivatives of the optimal value function with respect to the normalized time τ, the partial derivatives of the optimal value function with respect to the state variable x, and the partial derivatives of the optimal value function with respect to the total flight time a: Among them, V τ [τ,x(τ),a,v] represents the partial derivative of the optimal value function with respect to normalized time τ; H[τ,x(τ),u(τ),a,v] represents the Hamiltonian function; V x [1,x(1),a,v] represents the partial derivative of the optimal value function with respect to the state variable x; the superscript T indicates the transpose; V a [1,x(1),a,v] represents the partial derivative of the optimal value function with respect to the total flight time a.

4. The pseudospectral differential dynamic programming method according to claim 1, characterized in that, The iterative convergence index is determined by the following formula: Where, ρ c λ is the iterative convergence index; max It is the largest eigenvalue; This is the solution obtained in the i-th iteration; For the linear equation solved in the (i+1)th iteration, A S B S b S To form the above linear equation The coefficient matrix.

5. A pseudospectral differential dynamic programming device, characterized in that, To implement the pseudospectral differential dynamic programming method as described in any one of claims 1-4, the method includes: A nonlinear optimal guidance problem construction module is used to construct a nonlinear optimal guidance problem based on initial constraints, terminal constraints, and a dynamic model; the optimal value function of the nonlinear optimal guidance problem satisfies the Hamilton-Jacobi-Bellman equation. The module for determining the partial derivative differential equation of the optimal value function is used to perform a second-order expansion of the Hamilton-Jacobi-Bellman equation based on a specified reference trajectory to obtain the differential equation of the partial derivative of the optimal value function. The module for determining the first two-point boundary value problem is used to obtain the derivative of the construction matrix based on the differential equation with respect to the partial derivative of the optimal value function, and to determine the first two-point boundary value problem; the first two-point boundary value problem is a homogeneous form of the trajectory two-point boundary value problem. The second two-point boundary value problem determination module is used to solve the first two-point boundary value problem, determine the optimal control variable variation, and determine the second two-point boundary value problem based on the optimal control variable variation; solving the second two-point boundary value problem realizes the trajectory update; The pseudospectral method differential dynamic programming module is used to process the trajectory and construction matrix based on the pseudospectral method and Lagrange interpolation. It transforms the second two-point boundary value problem and performs iterative solution to obtain the discrete optimal trajectory and discrete construction matrix.

6. A computer device, comprising: A memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that the processor executes the computer program to implement the pseudospectral differential dynamic programming method according to any one of claims 1-4.

7. A computer-readable storage medium having a computer program stored thereon, characterized in that, When executed by a processor, the computer program implements the pseudospectral differential dynamic programming method as described in any one of claims 1-4.

8. A computer program product, comprising a computer program, characterized in that, When executed by a processor, the computer program implements the pseudospectral differential dynamic programming method as described in any one of claims 1-4.

Citation Information

Patent Citations

  • Online trajectory optimization method for power descending section of reusable rocket

    CN118586181A