Spacecraft trajectory optimization method, system, medium, and apparatus

By integrating Lie group SE(3) modeling and reinforcement learning predictive controller with dynamic event triggering strategy, the problem of attitude and orbit coupling effect in spacecraft trajectory optimization is solved, the control accuracy and computational efficiency are improved, and it is suitable for rendezvous and docking, on-orbit maintenance services and other tasks.

CN116853523BActive Publication Date: 2025-11-28SHANGHAI SATELLITE ENG INST
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202310715373.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-06-15
Publication Date
2025-11-28
Estimated Expiration
2043-06-15

AI Technical Summary

Technical Problem

Existing technologies fail to effectively consider the coupling effects of attitude and orbit in spacecraft trajectory optimization, resulting in insufficient control accuracy and an inability to effectively handle attitude line-of-sight constraints and system computational burden, making it difficult to meet the needs of complex space missions.

Method used

A spacecraft attitude and orbit integrated dynamic model was established using the Lie group SE(3), and discretized by the Lie group variational integral method. A finite-time domain reinforcement learning predictive controller was constructed, and a dynamic event triggering strategy was introduced to optimize the control problem and reduce the computational burden.

Benefits of technology

It achieves integrated attitude and orbit modeling, improves control accuracy, avoids the unwinding problem of quaternion representation, improves the accuracy and computational efficiency of HJB equation solutions, reduces the system's computational burden, and is suitable for trajectory optimization of complex space missions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116853523B_ABST
    Figure CN116853523B_ABST
Patent Text Reader

Abstract

The application provides a spacecraft trajectory optimization method, system, medium and equipment, including the following steps: establishing a spacecraft attitude and orbit integrated dynamics model based on Lie group SE(3); adopting a Lie group variational integral method and based on group characteristics, performing discretization processing on the dynamics model to obtain a discretized dynamics model; establishing a performance index function, converting the optimization control problem into an extreme value problem of solving a target function under constraints; constructing a finite time domain reinforcement learning predictive controller to solve the optimization problem; and constructing a dynamic event triggering strategy to correct the predictive controller. The application reduces the system calculation burden by introducing a dynamic event triggering control strategy, and effectively solves the six-degree-of-freedom spacecraft trajectory optimization control problem under constraints.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of spacecraft control, in particular to a spacecraft trajectory optimization method, system, medium and equipment, and especially to an SE(3) spacecraft trajectory optimization method based on reinforcement learning prediction. BACKGROUND

[0002] Spacecraft integrated attitude and orbit trajectory optimization control is a key technology to solve the problems of rendezvous and docking, on-orbit servicing, planetary soft landing, space station on-orbit assembly and other space missions. The traditional modeling and control method considers the attitude and orbit motion separately and ignores the coupling effect, only realizing the formal integrated attitude and orbit control, which affects the control accuracy. With the rapid development of space technology, the task demand becomes more and more complex, and in order to meet these task demands, the spacecraft attitude and orbit motion needs to have higher control accuracy. Compared with the traditional modeling method, the modeling method using Lie group SE(3) or dual quaternion not only realizes the integrated modeling in a true sense, but also guarantees higher modeling and control accuracy due to fully considering the coupling effect. However, it should be pointed out that the modeling method based on dual quaternion is inevitable to have the unwinding problem due to using quaternion to describe the attitude. Therefore, exploring the application of the integrated modeling and optimization control method on Lie group SE(3) in the field of spacecraft trajectory optimization control has very important significance to promote the development of spacecraft integrated modeling and control technology.

[0003] The article [Wang Yulin, Shang Wei, Hong Haichao, “Sub-optimal fixed-finite-horizon spacecraft configuration control on SE(3)”, Chinese Journal of Aeronautics, 2022] is about spacecraft attitude and orbit integrated modeling and optimal control method based on Lie group SE(3). Based on the continuous time domain attitude and orbit integrated dynamic model established based on Lie group SE(3) in the article [Ye Dong, Zhang Jianqiao, Sun Zhaowei, “Extended state observer-based finite-time controller design for coupled spacecraft formation with actuators saturation”, Advances in Mechanical Engineering, 2017, 9(4): 1-17], the Lie group variational integral method is adopted and the discrete dynamic model is obtained based on the group characteristics. The model predictive static programming method is used, and the constraint of spacecraft force and torque output is considered. The problem of six-degree-of-freedom spacecraft trajectory optimization control in finite time domain is solved. For on-orbit spacecraft, optical instruments such as cameras and infrared interferometers are usually carried. These instruments need to avoid direct alignment with strong light during use to protect the light-sensitive devices in the instruments, which requires considering the attitude line-of-sight angle constraint in the spacecraft trajectory planning process.

[0004] Model predictive control has been widely used in spacecraft trajectory optimization due to its ability to handle constraints and achieve high-performance optimization goals. The patent document with publication number CN108536014A discloses a model predictive control method for spacecraft attitude avoidance considering the dynamic characteristics of flywheels. Based on the established spacecraft attitude dynamics prediction model, the dynamic characteristics of flywheels and the line-of-sight angle constraints of instruments are considered. Performance index functions for different tasks are designed, and the optimal trajectory of the attitude is obtained by solving the extreme value of the objective function. The patent document with publication number CN113859589A discloses a spacecraft attitude control method based on model predictive control and sliding mode control. A compound incremental controller is designed based on the spacecraft attitude dynamics model to complete spacecraft attitude control. The target trajectory of attitude tracking is obtained through model predictive control, which ensures that the three-axis attitude has superior tracking performance in the presence of system uncertainties. However, the above patent documents are not suitable for solving the spacecraft trajectory optimization problem solved by the present application.

[0005] For traditional nonlinear optimal control problems, it is generally necessary to solve the Hamilton-Jacobi-Bellman (HJB) equation of the objective function, but it is generally difficult to obtain an analytical solution of the equation for complex systems. Reinforcement learning can effectively solve the above problems and obtain an approximate solution of the HJB equation with high precision. In recent years, reinforcement learning theory has been widely used in the field of optimal control. The patent document with publication number CN114036631A discloses a spacecraft autonomous rendezvous and docking guidance strategy generation method based on reinforcement learning, which mainly uses Markov decision process to model the spacecraft rendezvous and docking process, and based on neural network learning training data, constructs a decision table, to generate an optimal guidance strategy and complete the spacecraft autonomous rendezvous and docking task. The patent document with publication number CN112357120A discloses a reinforcement learning attitude constraint control method considering actuator installation deviation, which mainly uses reinforcement learning to learn and train controller parameters in real time, so that the controller evolves from a simple control strategy to a suboptimal controller, thereby improving the execution efficiency of spacecraft on-orbit tasks. However, the above patent documents are not suitable for solving the spacecraft trajectory optimization problem solved by the present application.

[0006] Event-driven control can effectively reduce the computational burden on the satellite. The patent document with publication number CN113722821A discloses a convex method for spacecraft rendezvous and docking trajectory planning event constraints, which mainly equivalently converts event constraints into convex form without using sequence approximation method, and further improves the convergence and efficiency of spacecraft rendezvous and docking trajectory planning with event constraints without changing the problem solution space, thereby solving the problem of low computational efficiency and poor convergence when existing methods are used to convex event constraints for trajectory planning. However, the above patent document is not suitable for solving the spacecraft trajectory optimization problem solved by the present application.

[0007] The patent document with the publication number CN112084581A discloses a small-thrust perturbation rendezvous trajectory optimization method and system, which comprises the following steps: calculating four-pulse velocity increments based on given orbit elements of a space vehicle and a target at starting and rendezvous time; assuming that the small-thrust switching strategy is on-off-on, estimating two small-thrust on-time; taking the small-thrust midpoint time as the equivalent pulse time, and recalculating the new four-pulse velocity increments according to the two small-thrust on-time; until the size change of the pulse increments calculated in the previous two times is less than a preset value, the small-thrust on-time is output; inputting the estimated small-thrust on-time into an indirect optimization model to solve the optimal control rate, the transfer trajectory and the mass change; calculating the increment percentage δ of the optimal control rate corresponding velocity increment and the pulse velocity increment, if δ is greater than a threshold, the step 5 is solved again, if δ is less than the threshold, the optimal control rate and the transfer trajectory are output. However, the patent document does not consider the attitude line-of-sight angle constraint in the spacecraft trajectory planning process. SUMMARY

[0008] In view of the defects in the prior art, the purpose of the present application is to provide a spacecraft trajectory optimization method, system, medium and equipment.

[0009] According to the spacecraft trajectory optimization method provided by the present application, the following steps are included:

[0010] Step 1: establishing a spacecraft attitude-orbit integrated dynamics model based on Lie group SE(3);

[0011] Step 2: discretizing the dynamics model in step 1 based on group characteristics by using Lie group variational integral method to obtain a discretized dynamics model;

[0012] Step 3: establishing a performance index function through the discretized dynamics model in step 2, and converting the optimization control problem into an extreme value problem of solving a target function under constraints;

[0013] Step 4: constructing a finite-time domain reinforcement learning predictive controller through the discretized dynamics model in step 2, and solving the converted optimization control problem in step 3 through the predictive controller;

[0014] Step 5: constructing a dynamic event triggering strategy through the discretized dynamics model in step 2, and correcting the predictive controller in step 4.

[0015] The present application also provides a spacecraft trajectory optimization system, which comprises the following modules:

[0016] Module M1: establishing a spacecraft attitude-orbit integrated dynamics model based on Lie group SE(3);

[0017] Module M2: the Lie group variational integral method is adopted, and based on group characteristics, the dynamic model in module M1 is discretized to obtain a discretized dynamic model;

[0018] Module M3: through the discretized dynamic model in module M2, a performance index function is established, and the optimization control problem is converted into an extreme value problem of solving a target function under constraints;

[0019] Module M4: through the discretized dynamic model in module M2, a finite time domain reinforcement learning predictive controller is constructed, and the converted optimization control problem in module M3 is solved through the predictive controller;

[0020] Module M5: through the discretized dynamic model in module M2, a dynamic event triggering strategy is constructed to correct the predictive controller in module M4.

[0021] The application also provides a computer readable storage medium storing a computer program, the computer program being executed by a processor to implement the steps of the spacecraft trajectory optimization method.

[0022] The embodiment also provides an electronic device, including a memory, a processor and a computer program stored in the memory and executable on the processor, the computer program being executed by the processor to implement the steps of the spacecraft trajectory optimization method.

[0023] Compared with the prior art, the application has the following beneficial effects:

[0024] 1. The application mainly explores a spacecraft attitude and orbit integrated dynamic modeling method based on Lie group SE(3), an optimal control method combining reinforcement learning and model predictive control, and solves a six-degree-of-freedom spacecraft trajectory optimization control problem, and through the introduction of a dynamic event triggering control strategy, the system calculation burden can be effectively reduced, and the engineering application potential of the designed control method is greatly increased.

[0025] 2. The application is based on Lie group and Lie algebra knowledge, and a Lie group SE(3) spacecraft attitude and orbit integrated dynamic model is derived and established, the obtained model realizes true integrated modeling compared with a traditional attitude and orbit integrated model, the modeling precision is improved due to the full consideration of attitude and orbit coupling effects, and compared with a modeling method based on dual quaternions, the unwinding problem existing in the quaternions representing attitude motion is effectively avoided.

[0026] 3. The application obtains an optimization strategy for each prediction interval through reinforcement learning, and compared with a traditional model predictive control method, the HJB equation solution precision is improved, and the calculation efficiency is effectively improved.

[0027] 4. This invention modifies the obtained optimal control sequence by designing a dynamic event triggering strategy. Without affecting system stability, it effectively solves the problem of the spacecraft bus communication resources being occupied by control signal updates and transmissions. In addition, by introducing dynamic variables into the event triggering conditions, the triggering conditions can be quickly adjusted when the system state changes, further improving the system's control performance.

[0028] 5. The spacecraft attitude and orbit integrated trajectory optimization method proposed in this invention is applicable to on-orbit intelligent mission planning for aerospace missions such as rendezvous and docking, on-orbit maintenance services, planetary soft landing, and on-orbit assembly of space stations. The application of event-triggered control theory greatly reduces the on-orbit computational pressure of the system. These two advantages increase the engineering potential of the trajectory optimization method proposed in this invention for future space-based intelligent systems. Attached Figure Description

[0029] Other features, objects, and advantages of the present invention will become more apparent from the following detailed description of non-limiting embodiments with reference to the accompanying drawings:

[0030] Figure 1 This is a flowchart of the spacecraft trajectory optimization method of the present invention;

[0031] Figure 2 A schematic diagram of the Earth's equatorial inertial coordinate system and the spacecraft's orbital coordinate system as defined in this invention;

[0032] Figure 3 The solution process involves discretizing the continuous time-domain attitude and orbit integrated dynamic model described by the Lie group SE(3).

[0033] Figure 4 This is a flowchart illustrating how reinforcement learning can be used to solve the HJB equation and obtain the optimal solution to the objective function. Detailed Implementation

[0034] The present invention will now be described in detail with reference to specific embodiments. These embodiments will help those skilled in the art to further understand the present invention, but do not limit the invention in any way. It should be noted that those skilled in the art can make several changes and improvements without departing from the concept of the present invention. These all fall within the protection scope of the present invention.

[0035] Example 1

[0036] like Figures 1 to 4 As shown, this embodiment provides a spacecraft trajectory optimization method, including the following steps:

[0037] Step 1: Establish an integrated attitude and orbit dynamic model of the spacecraft based on the Lie group SE(3);

[0038] Step 2: Discretize the dynamic model in Step 1 using the Lie group variational integration method and based on group properties to obtain a discretized dynamic model;

[0039] Step 3: Establish a performance index function through the discretized dynamic model in Step 2, and convert the optimization control problem into solving the extreme value problem of the objective function under constraints;

[0040] Step 4: Construct a finite time domain reinforcement learning predictive controller through the discretized dynamic model in Step 2, and solve the converted optimization control problem in Step 3 through the predictive controller;

[0041] Step 5: Construct a dynamic event triggering strategy through the discretized dynamic model in Step 2, and modify the predictive controller in Step 4.

[0042] Step 1 is to establish a basic six-degree-of-freedom spacecraft dynamic model based on the Lie group SE(3), but the need for trajectory optimization algorithm design is a discrete dynamic equation, so Step 2 uses the Lie group variational integration method to discretize the model, and Steps 3, 4 and 5 are specific optimization algorithm design, and the models used are all based on the discretized model in Step 2. Step 3 is to design an optimization control algorithm, Step 4 is to solve the optimization problem in Step 3 based on a reinforcement learning predictive controller, and Step 5 is to introduce a dynamic event triggering strategy to reduce system computation and modify the controller.

[0043] The embodiment also provides a computer readable storage medium storing a computer program, the computer program being executed by a processor to implement the steps of the spacecraft trajectory optimization method.

[0044] The embodiment also provides an electronic device including a memory, a processor, and a computer program stored on the memory and executable on the processor, the computer program being executed by the processor to implement the steps of the spacecraft trajectory optimization method.

[0045] Example 2

[0046] The embodiment provides a spacecraft trajectory optimization system, including the following modules:

[0047] Module M1: Establish a spacecraft attitude and orbit integrated dynamic model based on the Lie group SE(3);

[0048] Module M2: Discretize the dynamic model in Module M1 using the Lie group variational integration method and based on group properties to obtain a discretized dynamic model;

[0049] Module M3: Establish the performance index function through the discrete dynamic model in module M2, and convert the optimization control problem into solving the extreme value problem of the objective function under constraints;

[0050] Module M4: Construct the finite time horizon reinforcement learning predictive controller through the discrete dynamic model in module M2, and solve the converted optimization control problem in module M3 through the predictive controller;

[0051] Module M5: Construct the dynamic event-triggered strategy through the discrete dynamic model in module M2, and correct the predictive controller in module M4.

[0052] Example 3

[0053] Those skilled in the art can understand this embodiment as a more specific description of Embodiment 1 and Embodiment 2.

[0054] This embodiment provides an SE(3) spacecraft trajectory optimization method based on reinforcement learning prediction, comprising:

[0055] Step 1: Establish a spacecraft attitude-orbit integrated dynamic model based on Lie group SE(3);

[0056] Step 2: Obtain a discrete prediction model by using Lie group variational integral method and based on group characteristics;

[0057] Step 3: Design a performance index function, and convert the optimization control problem into solving the extreme value problem of the objective function under constraints;

[0058] Step 4: Design a finite time horizon reinforcement learning predictive controller to solve the above optimization problem;

[0059] Step 5: Design a dynamic event-triggered strategy to correct the predictive controller, and ensure the steady-state performance of the closed-loop system.

[0060] Further, the step 1 comprises:

[0061] Define two coordinate systems: the Earth equatorial inertial coordinate system F I (x I ,y I ,z I ) and the spacecraft body coordinate system F b (x b ,y b ,z b ), C is the direction cosine matrix of the spacecraft from the coordinate system F b to the coordinate system F I , that is, the spacecraft attitude, which is an element of Lie group SO(3), and SO(3) is a special orthogonal set, satisfying: SO(3)={C∈R 3×3C T C = I 3×3 , det(C) = 1}, R is a real number set, R 3×3 is a space composed of 3 × 3 real number matrix, different upper indices represent corresponding matrix or vector dimension, () T is to find the transpose of a matrix, I 3×3 is a 3 × 3 unit matrix, det() is to find the determinant of a matrix, then the spacecraft attitude kinematics equation can be expressed as:

[0062]

[0063] Where, ω = [ω1, ω2, ω3] T ∈ R 3×1 Indicates the attitude angular velocity of the spacecraft coordinate system F b relative to the coordinate system F I , the subscripts 1, 2, 3 represent the angular velocity components of ω in the three inertial principal axis directions of the spacecraft, Indicates the skew-symmetric matrix composed of three-dimensional vectors, is the Lie algebra of SO(3).

[0064] Let the position coordinates of the spacecraft in the inertial system F I be R, and the orbital velocity in the body coordinate system F b be v, then the spacecraft orbital kinematics equation can be expressed as:

[0065]

[0066] In order to integrate the description of the spacecraft attitude and orbit motion, the Lie group SE(3) is introduced as a mathematical tool. SE(3) is a group space composed of semi-direct product SO(3) × R 3 in 4 × 4 homogeneous form, and its element g has the following form

[0067]

[0068] Where, g is the pose configuration of the spacecraft, 0 1×3 is a 1 × 3 zero vector, then the spacecraft pose-orbit integrated kinematics equation described by the Lie group SE(3) is expressed as:

[0069]

[0070] Where, is the first derivative of the spacecraft configuration, ξ = [ω T , v T ] T ∈ R 6×1 is a six-dimensional velocity vector composed of attitude angular velocity and orbital velocity, is defined as:

[0071]

[0072] where, is the Lie algebra of SE(3), and R 6×1 isomorphism;

[0073] The system F b The spacecraft attitude and orbit dynamics equations are:

[0074]

[0075] where, J e R 3×3 is the spacecraft's moment of inertia; is the first derivative of the angular velocity, τ c is the control torque generated by the spacecraft actuators, d τ is the external disturbance torque experienced by the spacecraft; m is the mass of the spacecraft, f c is the control force generated by the spacecraft actuators, d f is the external disturbance force experienced by the spacecraft.

[0076] For the Lie group SE(3), there exists an adjoint map:

[0077] ad X Y = [X, Y] = XY - YX (7)

[0078] where X and Y are matrices of corresponding dimensions. Through algebraic operations, the above adjoint operator can be expressed in the form of a matrix:

[0079]

[0080] Two elements The left-invariant inner product defined on SE(3) is defined as:

[0081]

[0082] The Lie bracket on SE(3) is:

[0083]

[0084] The operator ad defined in equation (10) represents a linear operation between the Lie algebra and SE(3), and its adjoint operator can be obtained through the dual operation of the Lie algebra, which has the following form:

[0085]

[0086] Based on formula (6) and (11), the spacecraft integrated attitude and orbit dynamics equation can be expressed as:

[0087]

[0088] Wherein, Ξ = diag(J, mI 3×3 ) ∈ R 6×6 is the spacecraft mass characteristic matrix composed of the moment of inertia J and the mass m, diag() is the diagonal matrix operation, is the first derivative of the velocity, is the control vector composed of the control force and the moment, is the disturbance vector composed of the disturbance force and the moment.

[0089] Further, the step 2 comprises:

[0090] For spacecraft motion, the attitude angle change range is generally within positive and negative π, but the change amplitude of the orbit motion is relatively large. In order to ensure that the trajectory changes of the orbit and the attitude motion are in the same order of magnitude, the orbit motion is normalized, and is defined as:

[0091] R′ = R / R m ,v′ = v / v m (13)

[0092] Wherein, R′ ∈ R 3×1 ,v′ ∈ R 3×1 indicate the normalized position and velocity, R m ,v m is a constant, indicating the normalization parameter. Then, the normalized spacecraft kinematics and dynamics equation expression is obtained:

[0093]

[0094] The Lie group variational integral method is used to discretize formula (14), and the discrete dynamics model is obtained, and its expression is as follows:

[0095]

[0096] Wherein, k is the time sequence, the parameter subscript k, k+1 indicates the corresponding parameter value at the corresponding discrete time, h is the discrete time interval, f k ∈ SO(3) is the group variation of C k , J d = 0.5tr(J)I 3×3 -J, tr() is the trace of the matrix. In the discretization process, the external disturbance force and moment suffered by the system are ignored. f′ ck = f ck / vm The system output, the system state and the input and output are defined as:

[0097]

[0098] Since C k ,f k ∈SO(3), each element in the group SO(3) includes 9 sub-elements and satisfies 6 constraints, and the Jacobian matrix of formula (15) is difficult to solve, so many optimization methods are difficult to apply directly. Next, the variational method and Lie group characteristics are used to process the integrated dynamics model of pose and orbit, and a discrete dynamics model suitable for reinforcement learning predictive controller design is obtained.

[0099] For predictive control, at time T0, the initial orbit and attitude motion state of the spacecraft is X0, and after the action of control sequence U k (k=0,1,2,…), at time T0+kt p (t p is the prediction time interval), the system output state is Y k (g k ,ξ k ). Denote the output error of the system (i.e. the error of Y k and Y dk , Y dk (g dk ,ξ dk ) is the target trajectory of the spacecraft at this time) as dY k =[dC k dR′ k dω k dv′ k ], where, are the actual pose configuration and velocity of the spacecraft at this time, are the target pose configuration and velocity of the spacecraft at this time, C dk and R′ dk are the target attitude and target orbit position of the spacecraft, d() is the variation of the corresponding parameter, and the continuous time domain expression of each element in dY k is:

[0100]

[0101] where φ∈R 3×1 is an intermediate variable used to represent dC, and its derivative satisfies the right side of formula (17). Since dC can be equivalently represented by φ, the discrete variation form of the system state, input and output in formula (16) is defined as:

[0102] The following discrete-time dynamics model is obtained for the design of the easy-to-reinforce learning predictive controller:

[0103]

[0104] where f is a Lipschitz function, T0 is the initial time, and x0 is the initial state. k ,B k are matrices with corresponding dimensions, and their specific expressions are as follows:

[0105]

[0106] Further, the step 3 includes that the to-be-optimized objective function is designed as:

[0107]

[0108] where N is the prediction time length, Q, H, and P are in R 12×12 are positive definite symmetric weight matrices. The objective function includes two parts: the first part is the cost function after the predictive control in the finite time domain, when Q > H, the optimization objective focuses on the stabilization time; when Q < H, the optimization objective focuses on the energy consumption; the second part is used to evaluate the last trajectory optimization tracking accuracy of the spacecraft after the predictive control in the finite time domain, and satisfies:

[0109]

[0110] where K is in R 6×12 is a linear feedback control gain, P is a prediction end penalty matrix, and Ω α is a region containing the prediction end, α > 0, and u k = Kx k can guarantee the asymptotic stability of the closed-loop system, P, Ω α are calculated offline.

[0111] The constraint problems faced by the spacecraft trajectory optimization mainly consider (but are not limited to): the maximum output force and torque constraints of the spacecraft actuators, and the line-of-sight angle constraints of the spacecraft attitude pointing. The mathematical expressions of the constraints are shown in formulas (27) and (28):

[0112]

[0113] where τ max ,f max are the upper limits of the control torque and the control force, respectively.

[0114] The line-of-sight constraint of spacecraft attitude pointing is taken as an example that the optical camera field of view cannot point to the sun. The line-of-sight constraint can be defined as: in the body coordinate system F b of the spacecraft, a line-of-sight axis whose unit vector is l b , the unit position vector of the spacecraft and the sun is r b , and the line-of-sight constraint angle is θ, then the following relationship is satisfied:

[0115] (l b ,r b )≤cosθ (28)

[0116] The six-degree-of-freedom spacecraft trajectory optimization control problem is converted into the following mathematical problem, that is, in the prediction interval [T0, T0+Nt p ], the following optimization problem is solved:

[0117]

[0118] Wherein, E{} is the expectation operator.

[0119] Further, the step 4 comprises:

[0120] The cost function in the step length [kt p , (k+1)t p ] is defined as

[0121]

[0122] The following is abbreviated as r(k). The optimization problem is rewritten as:

[0123]

[0124] Wherein,

[0125]

[0126] According to the Bellman optimization theory, there is an optimal objective function satisfying the following Hamilton-Jacobi-Bellman (HJB) equation:

[0127]

[0128] Solving the above equation can obtain the optimal control The specific form is:

[0129]

[0130] Since it is difficult to obtain the analytical solution of the HJB equation, the finite horizon reinforcement learning method is used to train the numerical solution of the equation. In the prediction control interval [k, N-1], the initial target function value is defined as J i=0 () = 0, and then for i = 0, 1, … and τ ∈ [k, N-1], the formula (35) is used to iteratively learn and calculate

[0131]

[0132] Objective function is updated by the following formula:

[0133]

[0134] Further, the step 5 comprises:

[0135] Based on the dynamic event-triggered predictive control, it is defined that at time t i , the predictive control is performed according to the system state , and a control vector is obtained. i+1 Until the next triggering time t i comes, no new predictive control update is performed, that is, in the interval [t i+1 , t ∞ ], the control vector remains unchanged, and the two triggering times satisfy the following relationship:

[0136] It can be seen that the dynamic change of n can greatly save the computing resources of the spacecraft system.

[0137] The state error is defined as:

[0138]

[0139] Wherein, Then, the following dynamic event-triggering rule is designed:

[0140]

[0141] Wherein, β ≥ 0, σ ∈ (0, 1), ρ, γ are K ∞ Functions, δ satisfies:

[0142]

[0143] Wherein, μ ∈ (0, 1), δ0≥ 0, ‖e|| is the two-norm of the calculation vector.

[0144] When formula (39) is not satisfied, the Values ​​and assign values. t i+1 =t i +nt p , k = k + n, and then a new predictive control sequence is updated through reinforcement learning predictive control.

[0145] Example 4

[0146] Those skilled in the art can understand this embodiment as a more specific description of Embodiment 1 and Embodiment 2.

[0147] This embodiment provides an SE(3) spacecraft trajectory optimization system based on reinforcement learning prediction, including the following modules:

[0148] Module M1: Establishing an integrated dynamic model of spacecraft attitude and orbit based on Lie group SE(3);

[0149] Module M2: Employs the Lie group variational integral method and obtains a discretized prediction model based on group characteristics;

[0150] Module M3: Design performance index functions to transform the optimization control problem into a problem of finding the extremum of the objective function under constraints;

[0151] Module M4: Design a finite-time reinforcement learning predictive controller to solve the above optimization problem;

[0152] Module M5: Design a predictive controller with dynamic event-triggered strategy correction and ensure the steady-state performance of the closed-loop system.

[0153] Furthermore, module M1 performs the following process:

[0154] Define two coordinate systems: the Earth's equatorial inertial coordinate system F. I (x I ,y I ,z I ) and spacecraft body coordinate system F b (x b ,y b ,z b C represents the spacecraft's position from coordinate system F. b Switch to coordinate system F I The direction cosine matrix, i.e., the spacecraft attitude, is an element of the Lie group SO(3), where SO(3) is a special orthogonal set satisfying: SO(3)={C∈R 3×3 :C T C = I 3×3 Let det(C) = 1, and R be the set of real numbers. 3×3 Let be a space consisting of 3×3 real matrices, where different superscripts represent the corresponding matrix or vector dimensions. TI 3×3 is the 3x3 identity matrix, det() is the determinant of a matrix, and the spacecraft's attitude kinematics equation can be expressed as:

[0155]

[0156] where ω = [ω1, ω2, ω3] T ∈ R 3×1 represents the attitude angular velocity of the spacecraft coordinate system F b relative to the coordinate system F I , and subscripts 1, 2, 3 represent the angular velocity components of ω in the three inertial principal axis directions of the spacecraft, represents an antisymmetric matrix composed of three-dimensional vectors, is the Lie algebra of SO(3).

[0157] Let the position coordinates of the spacecraft in the inertial system F I be R, and the orbital velocity in the body system F b be v, then the spacecraft's orbital kinematics equation can be expressed as:

[0158]

[0159] In order to integrate the spacecraft's attitude and orbital motion, the Lie group SE(3) is introduced as a mathematical tool. SE(3) is a group space composed of semi-direct product SO(3) x R 3 in 4x4 homogeneous form, and its element g has the following form

[0160]

[0161] where g is the pose configuration of the spacecraft, 0 1×3 is a 1x3 zero vector, and the spacecraft's integrated attitude and orbital kinematics equation described by the Lie group SE(3) is expressed as:

[0162]

[0163] where is the first derivative of the spacecraft configuration, ξ = [ω T , v T ] T ∈ R 6×1 is a six-dimensional velocity vector composed of attitude angular velocity and orbital velocity, is defined as:

[0164]

[0165] where is the Lie algebra of SE(3), and R6×1 isomorphism;

[0166] The system F b The spacecraft's attitude and orbit dynamics equations are:

[0167]

[0168] where J e R 3×3 is the spacecraft's moment of inertia; is the first derivative of the angular velocity, τ c is the control torque generated by the spacecraft's actuators, d τ is the external disturbance torque acting on the spacecraft; m is the spacecraft's mass, f c is the control force generated by the spacecraft's actuators, d f is the external disturbance force acting on the spacecraft.

[0169] For the Lie group SE(3), there exists an adjoint mapping:

[0170] ad X Y = [X, Y] = XY - YX (47)

[0171] where X and Y are matrices of corresponding dimensions. Through algebraic operations, the above adjoint operator can be expressed in the form of a matrix:

[0172]

[0173] Two elements The left-invariant inner product defined on SE(3) is defined as:

[0174]

[0175] The Lie bracket on is:

[0176]

[0177] The operator ad defined by equation (50) represents the linear operation between the Lie algebra and SE(3), and its adjoint operator can be obtained through the dual operation of the Lie algebra, which has the following form:

[0178]

[0179] Based on equations (46) and (51), the spacecraft's integrated attitude and orbit dynamics equations can be expressed as:

[0180]

[0181] where Ξ = diag(J, mI3×3 )∈R 6×6 is the spacecraft mass property matrix composed of the moment of inertia J and mass m, diag() is the diagonal matrix operation, is the first derivative of velocity, is the control vector composed of control force and torque, is the disturbance vector composed of disturbance force and torque.

[0182] Further, the module M2 performs the following process:

[0183] For spacecraft motion, the attitude angle change range is generally within plus or minus π, but the change amplitude of the orbit motion is relatively large. In order to ensure that the trajectory changes of the orbit and the attitude motion are in the same order of magnitude, the orbit motion is normalized, and is defined as:

[0184] R′=R / R m ,v′=v / v m (53)

[0185] wherein R′∈R 3×1 ,v′∈R 3×1 denote the normalized position and velocity, R m ,v m are constants, and denote the normalization parameters. Then, the normalized kinematics and dynamics equation expressions of the spacecraft are obtained as follows:

[0186]

[0187] The Lie group variational integral method is used to discretize formula (54), and the discretized dynamics model is obtained, and its expression is as follows:

[0188]

[0189] wherein k is a time sequence, the parameter subscripts k and k+1 denote the corresponding parameter values at the corresponding discrete time, h is a discretization time interval, f k ∈SO(3) is the group variation of C k , J d =0.5tr(J)I 3×3 -J, and tr() is the trace of a matrix. In the discretization process, the external disturbance force and torque suffered by the system are ignored. f′ ck =v ck / v m is the orbit part output of the system after normalization, and the state and input and output of the system are defined as:

[0190]

[0191] Since Ck ,f k ∈SO(3), each element in the group SO(3) includes 9 sub-elements and satisfies 6 constraints, and the Jacobian matrix of formula (55) is difficult to solve, so many optimization methods are difficult to apply directly. Next, the variational method and Lie group characteristics are used to process the integrated dynamics model of pose and orbit, and a discrete dynamics model suitable for reinforcement learning predictive controller design is obtained.

[0192] For predictive control, at time T0, the initial orbit and attitude motion state of the spacecraft is X0, and after the control sequence U k (k = 0, 1, 2, …) acts, at time T0+kt p (t p is the prediction time interval), the system output state is Y k (g k , ξ k ). Denote the system output error (i.e. the error of Y k and Y dk , Y dk (g dk , ξ dk ) is the target trajectory of the spacecraft at this time) as dY k = [dC k dR′ k dω k dv′ k ], where, are the actual pose configuration and velocity of the spacecraft at this time, are the target pose configuration and velocity of the spacecraft at this time, C dk and R′ dk are the target attitude and target orbit position of the spacecraft, d() is the variation of the corresponding parameter, and the continuous time domain expression of each element in dY k is:

[0193]

[0194] where φ ∈ R 3×1 is an intermediate variable used to represent dC, and its derivative satisfies the right side of formula (57). Since dC can be equivalently represented by φ, the following discrete variation form is defined for the system state, input and output in formula (56):

[0195] The following discrete dynamics model suitable for reinforcement learning predictive controller design is obtained:

[0196]

[0197] where f is a Lipschitz function, T0 is the initial time, and x0 is the initial state. k ,B k are matrices with corresponding dimensions, and their specific expressions are as follows:

[0198]

[0199] Further, the module M3 performs the following process:

[0200] The to-be-optimized objective function is designed as:

[0201]

[0202] where N is the prediction time length, Q, H, and P are in R 12×12 is a positive definite symmetric weight matrix. The objective function includes two parts: the first part is the cost function after the prediction control in the finite time domain, when Q > H, the optimization objective focuses on the stabilization time; when Q < H, the optimization objective focuses on energy consumption; the second part is used to evaluate the last trajectory optimization tracking accuracy of the spacecraft after the prediction control in the finite time domain, and satisfies:

[0203]

[0204] where K is in R 6×12 is a linear feedback control gain, P is a prediction end penalty matrix, and Ω α is a region containing the prediction end, and α > 0, and u k = Kx k can guarantee the asymptotic stability of the closed-loop system, P, Ω α are calculated offline.

[0205] The constraint problems faced by spacecraft trajectory optimization mainly consider (but are not limited to): the maximum output force and torque constraints of spacecraft actuators, and the line-of-sight angle constraints of spacecraft attitude pointing. The mathematical expressions of the constraints are shown in equations (67) and (68):

[0206]

[0207] where τ max , f max are the upper limits of the output of the control torque and the control force, respectively.

[0208] The line-of-sight angle constraint of spacecraft attitude pointing takes the case that the field of view of an optical camera cannot point to the sun as an example, and the line-of-sight constraint can be defined as: a line-of-sight axis in the spacecraft body coordinate system F b , whose unit vector is l b , and the unit position vector of the spacecraft and the sun is r b, the line-of-sight constraint angle is θ, and the following relationship is satisfied among the three:

[0209] (l b ,r b )≤cosθ (68)

[0210] The six-degree-of-freedom spacecraft trajectory optimization control problem is converted into the following mathematical problem for description, that is, in the prediction interval [T0, T0+Nt p ], the following optimization problem is solved:

[0211]

[0212] Wherein, E{} is the expectation operator.

[0213] Further, the module M4 performs the following process:

[0214] The cost function in the step [kt p ,(k+1)t p ] is defined as

[0215]

[0216] The following is abbreviated as r(k). The optimization problem is rewritten as:

[0217]

[0218] Wherein,

[0219]

[0220] According to the Bellman optimization theory, there is an optimal objective function Satisfying the following Hamilton-Jacobi-Bellman (HJB) equation:

[0221]

[0222] Solving the above equation can obtain the optimal control The specific form is:

[0223]

[0224] Since it is difficult to obtain the analytical solution of the HJB equation, the finite time domain reinforcement learning method is used to train the numerical solution of the equation. In the prediction control interval [k, N-1], the initial objective function value is defined as J i=0 ()=0, then for i=0,1,… and τ∈[k,N-1], the formula (75) is used to iteratively learn and calculate

[0225]

[0226] Objective function The update is performed by the following equation:

[0227]

[0228] Further, the module M5 performs the following process:

[0229] Based on the dynamic event-triggered predictive control, defined as at time t i , the predictive control is performed according to the system state to obtain a control vector Until the next trigger time t i+1 comes, no new predictive control update will be performed, that is, the control vector i remains unchanged within the interval [t i+1 , t The two trigger times satisfy the following relationship:

[0230] It can be seen that the dynamic change of n can greatly save the computing resources of the spacecraft system.

[0231] Define the state error:

[0232]

[0233] Wherein, Then, the following dynamic event-triggering rule is designed:

[0234]

[0235] Wherein, β≥0, σ∈(0,1), ρ, γ are K ∞ Functions, δ satisfies:

[0236]

[0237] Wherein, μ∈(0,1), δ0≥0, ||e|| is the two-norm of the calculation vector.

[0238] When formula (79) is not satisfied, the value of at this time is calculated, and the assignment t i+1 = t i + nt p , k=k+n is performed, and then the new predictive control sequence update is performed by reinforcement learning predictive control.

[0239] Example 5

[0240] Those skilled in the art can understand this embodiment as a more specific description of Embodiment 1 and Embodiment 2.

[0241] This embodiment provides a reinforcement learning-based prediction method for SE(3) spacecraft trajectory optimization, including:

[0242] Step 1: Establish an integrated attitude and orbit dynamic model of the spacecraft based on the Lie group SE(3);

[0243] Step 2: Employ the Lie group variational integral method and obtain a discretized prediction model based on group characteristics;

[0244] Step 3: Design the performance index function, transforming the optimization control problem into a problem of finding the extremum of the objective function under constraints;

[0245] Step 4: Design a finite-time reinforcement learning prediction controller to solve the above optimization problem;

[0246] Step 5: Design a dynamic event-triggered strategy to correct the predictive controller and ensure the steady-state performance of the closed-loop system.

[0247] This embodiment adopts a controller design concept that combines reinforcement learning and model predictive control. It rationally designs the objective function to achieve the comprehensive optimization of time and energy. By introducing a dynamic event-triggered control strategy, it reduces the computational burden of the system and effectively solves the problem of constrained six-degree-of-freedom spacecraft trajectory optimization control.

[0248] like Figure 1 As shown, the specific implementation steps of this embodiment are as follows:

[0249] Step 1: Based on the Lie group SE(3), derive and establish a six-degree-of-freedom spacecraft attitude and orbit integrated dynamic model.

[0250] Define two coordinate systems: the Earth's equatorial inertial coordinate system F. I (x I ,y I ,z I ) and spacecraft body coordinate system F b (x b ,y b ,z b ),like Figure 2 As shown. The origin of the inertial coordinate system is the Earth's center of mass, O. I x I Pointing to the vernal equinox, O I z I Aligned with the Earth's axis of rotation, pointing towards the North Pole, O I y I This is obtained using the right-hand rule; the origin of the spacecraft's body coordinate system is its center of mass, and its three coordinate axes coincide with its principal inertial axes, forming a right-hand coordinate system. C represents the spacecraft's coordinate system F.b Go to F I The direction cosine matrix, i.e., the spacecraft attitude, is an element of the Lie group SO(3), where SO(3) is a special orthogonal set satisfying: SO(3)={C∈R 3×3 :C T C = I 3×3 Let det(C) = 1, and R be the set of real numbers. 3×3 Let be a space consisting of 3×3 real matrices, where different superscripts represent the corresponding matrix or vector dimensions. T For the transpose of the matrix, I 3×3 Given a 3×3 identity matrix, and det() which calculates the determinant of a matrix, the spacecraft's attitude kinematics equations can be expressed as:

[0251]

[0252] Where ω = [ω1, ω2, ω3] T ∈R 3×1 In the spacecraft's intrinsic coordinate system, F represents the coordinate system. b Relative to coordinate system F I The attitude angular velocity, with subscripts 1, 2, 3 indicating the angular velocity components of ω along the three principal inertial axes of the spacecraft. This represents an antisymmetric matrix composed of three-dimensional vectors. Let SO(3) be the Lie algebra.

[0253] Record the spacecraft in the inertial frame F I The position coordinates below are R, and in this system F b If the orbital velocity is v, then the orbital kinematic equation of the spacecraft can be expressed as:

[0254]

[0255] To provide a unified description of the attitude and orbital motion of a spacecraft, the mathematical tool of the Lie group SE(3) is introduced. SE(3) is a semi-direct product of SO(3) × R. 3 A group space spanned in a 4×4 homogeneous form has an element g with the following form.

[0256]

[0257] Where g represents the spacecraft's attitude configuration, 0 1×3 If the zero vector is 1×3, then the spacecraft attitude-orbit integrated kinematic equations described by the Lie group SE(3) are expressed as:

[0258]

[0259] in, The first derivative of the spacecraft configuration, ξ = [ω T ,v T ] T ∈ R 6×1 is a six-dimensional velocity vector consisting of the attitude angular velocity and the orbital velocity, defined as:

[0260]

[0261] where, is the Lie algebra of SE(3) and is isomorphic to R 6×1 ;

[0262] The equations of spacecraft attitude and orbit dynamics in the system F b are:

[0263]

[0264] where, J ∈ R 3×3 is the spacecraft's moment of inertia; is the first derivative of the angular velocity, τ c is the control torque generated by the spacecraft actuators, d τ is the external disturbance torque experienced by the spacecraft; m is the mass of the spacecraft, f c is the control force generated by the spacecraft actuators, d f is the external disturbance force experienced by the spacecraft.

[0265] For the Lie group SE(3), there exists an adjoint map:

[0266] ad X Y = [X, Y] = XY - YX (87)

[0267] where X and Y are matrices of corresponding dimensions. Through algebraic manipulation, the above adjoint operator can be expressed in the form of a matrix:

[0268]

[0269] Two elements are defined on SE(3) as:

[0270]

[0271] The Lie bracket on SE(3) is:

[0272]

[0273] The operator ad defined in equation (90) represents the Lie algebra The inverse adjoint operator of the linear operation between SE(3) can be obtained by the dual operation of Lie algebra, which has the following form:

[0274]

[0275] Based on the formula (86) and (91), the spacecraft integrated attitude and orbit dynamics equation can be integrated and expressed as:

[0276]

[0277] Where, Ξ = diag(J, mI 3×3 ) ∈ R 6×6 is the spacecraft mass characteristic matrix composed of the moment of inertia J and the mass m, diag() is the diagonal matrix operation, is the first derivative of the velocity, is the control vector composed of the control force and the moment, is the disturbance vector composed of the disturbance force and the moment.

[0278] Step 2: Discretize the integrated attitude and orbit dynamics model using Lie group variational integral method, and convert the integrated attitude and orbit dynamics model into a form that is easy to design a reinforcement learning predictive controller based on group characteristics.

[0279] For spacecraft motion, the attitude angle change range is generally within plus or minus π, but the change amplitude of the orbit motion is relatively large. In order to ensure that the trajectory changes of the orbit and the attitude motion are in the same order of magnitude, the orbit motion is normalized, and is defined as:

[0280] R′ = R / R m ,v′ = v / v m (93)

[0281] Where, R′ ∈ R 3×1 ,v′ ∈ R 3×1 denote the normalized position and velocity, R m ,v m are constants, which represent the normalization parameters. Then, the normalized spacecraft kinematics and dynamics equation expression is obtained:

[0282]

[0283] Discretize formula (94) using Lie group variational integral method to obtain the discrete dynamics model, which is expressed as:

[0284]

[0285] where k is the time sequence, the parameter subscript k, k+1 represents the corresponding parameter value at the discrete time, h is the discretization time interval, f k ∈SO(3) is the group variation of C k d =0.5tr(J)I 3×3 -J, tr() is the trace of the matrix. In the discretization process, the external disturbance force and torque suffered by the system are ignored. The flow chart of the model discretization solving process is shown in Figure 3 .

[0286] Define f′ ck = f ck / v m as the orbit part output of the system after normalization processing, the state of the system and the input and output are defined as:

[0287]

[0288] Since C k ,f k ∈SO(3), each element in the group SO(3) includes 9 sub-elements and satisfies 6 constraints, and the Jacobian matrix of formula (95) is difficult to solve, so many optimization methods are difficult to apply directly. Next, the variational method and Lie group characteristics are used to process the integrated dynamics model of pose and orbit, and a discretized dynamics model easy to design a reinforcement learning predictive controller is obtained.

[0289] For predictive control, at T0time, the initial orbit and attitude motion state of the spacecraft is X0, after the action of the control sequence U k (k=0, 1, 2, …), at T0+kt p (time t p is the prediction time interval), the output state of the system is Y k (g k , ξ k ). Denote the output error of the system (i.e. the error of Y k and Y dk , Y dk (g dk , ξ dk ) is the target trajectory of the spacecraft at this time) as dY k =[dC k dR′ k dω k dv′ k ], where, are the actual pose configuration and velocity of the spacecraft at this time, are the target pose configuration and velocity of the spacecraft at this time, C dk and R′ dk ​are the target attitude and the target orbit position of the spacecraft, respectively, d() is the variation of the corresponding parameter, dY k The continuous-time expression of each element in is:

[0290]

[0291] where φ∈R 3×1 is an intermediate variable used to represent dC, whose derivative satisfies the right side of equation (97). Since dC can be equivalently represented by φ, the discrete variations of the system states, inputs, and outputs in equation (96) are defined as follows:

[0292] The following discrete-time dynamics model is obtained, which is easy to use for the design of reinforcement learning model predictive controllers:

[0293]

[0294] where f is a Lipschitz function, T0 is the initial time, and x0 is the initial state. k ,B k are matrices with corresponding dimensions, and their specific expressions are as follows:

[0295]

[0296] Step 3: Establish a mathematical model of the constraints, design an optimization index function, and convert the optimization problem into a problem of finding the extreme value of the objective function under constraints.

[0297] The to-be-optimized objective function is designed as:

[0298]

[0299] where N is the prediction time length, Q, H, P∈R 12×12 are positive definite symmetric weight matrices. The objective function includes two parts: the first part is the cost function after the prediction control within a finite time domain. When Q>H, the optimization objective focuses on the stabilization time; when Q is used to evaluate the optimization tracking accuracy of the final trajectory of the spacecraft after the prediction control within a finite time domain, and satisfies:

[0300]

[0301] where K∈R 6×12 is the linear feedback control gain, P is the prediction end penalty matrix, Ω α is a region containing the prediction end, α>0, and u k =Kx kP, Ω α are calculated offline. The detailed calculation is as follows:

[0302] After the closed-loop system is controlled by the finite horizon predictive controller for Nt p time, it moves into the region Ω α . At this time, the discretized system error dynamics equation is in the form of equation (102), x k = 0 is the equilibrium point of the system. Considering that the disturbance suffered by the system is f(d k ), when the controller is designed as u k = Kx k , the closed-loop system can be obtained as:

[0303]

[0304] The prediction end point penalty matrix P is designed as:

[0305]

[0306] where Q * = Q + K T HK∈R 12×12 is a positive definite symmetric matrix, λ max () is the maximum eigenvalue of the matrix.

[0307] The Lyapunov function V(x) is selected as:

[0308]

[0309] Taking the derivative of the above equation and substituting equations (107) and (108) into it, we can obtain:

[0310]

[0311] where, Therefore, according to equation (110), the closed-loop system is asymptotically stable in the region Ω α .

[0312] The constraint problems faced by spacecraft trajectory optimization mainly consider (but are not limited to): the maximum output force and torque constraints of spacecraft actuators, and the line-of-sight angle constraints of spacecraft attitude pointing. The mathematical expressions of the constraints are shown in equations (111) and (112):

[0313]

[0314] where τ max , f max are the upper limits of the output of the control torque and the control force, respectively.

[0315] The line-of-sight angle constraint of spacecraft attitude pointing, for example, the optical camera field of view cannot point to the sun, the line-of-sight constraint can be defined as: in the spacecraft body coordinate system F b , a line-of-sight axis whose unit vector is l b , the unit position vector of the spacecraft and the sun is r b , and the line-of-sight constraint angle is θ, then the following relationship is satisfied:

[0316] (l b ,r b )≤cosθ (112)

[0317] The six-degree-of-freedom spacecraft trajectory optimization control problem is converted into the following mathematical problem, that is, in the prediction interval [T0, T0+Nt p ], the following optimization problem is solved:

[0318]

[0319] Where E{} is the expectation operator.

[0320] Step 4: Design a finite time domain reinforcement learning predictive controller to solve the above optimization problem and ensure the stability of the closed-loop system.

[0321] Define the cost function in the step [kt p , (k+1)t p ] as

[0322]

[0323] The optimization problem is rewritten as:

[0324]

[0325] Where,

[0326]

[0327] According to the Bellman optimization theory, there is an optimal objective function Satisfying the following Hamilton-Jacobi-Bellman (HJB) equation:

[0328]

[0329] Solving the above equation can obtain the optimal control The specific form is:

[0330]

[0331] Since it is difficult to obtain the analytical solution of the HJB equation, the finite horizon reinforcement learning method is used to train the numerical solution of the equation, and the solving process is shown in Figure 4

[0332] First, initialize the parameter value, set i = 0, select a small positive number ε > 0, and define the target function value for

[0333] Second, for i = 0, 1, … and τ ∈ [k, N-1], use formula (119) to iteratively learn and calculate

[0334]

[0335] Third, according to the discretized dynamic model (95), calculate the system state x τ+1 at the next time.

[0336] Fourth, update the target function by

[0337]

[0338] Fifth, calculate If , it means that the target function trained by reinforcement learning has converged to the optimal state, and the iteration is ended. Otherwise, go back to the second step to continue reinforcement learning training until the solution of equation (117) is obtained.

[0339] Step 5: Design a dynamic event-triggered strategy to modify the predictive controller, which effectively reduces the computational burden of the system while ensuring the steady-state performance of the closed-loop system.

[0340] Based on dynamic event-triggered predictive control, at time t i , according to the system state , the predictive control is performed to obtain a control vector Until the next trigger time t i+1 comes, no new predictive control update will be performed, that is, in the interval [t i , t i+1 ], the control vector remains unchanged, and the two trigger times satisfy the following relationship:

[0341] From this, it can be seen that the dynamic change of n can greatly save the computational resources of the spacecraft system.

[0342] Define the state error as

[0343] ​​

[0344] wherein, Then, the dynamic event-triggered rule is designed as follows:

[0345]

[0346] wherein, β≥0, σ∈(0,1), ρ,γ are K ∞ functions, δ satisfies:

[0347]

[0348] wherein, μ∈(0,1), δ0≥0, ||e|| is the two-norm of the calculation vector.

[0349] For the closed-loop system, if the system is input-to-state stable, then there is a selected Lyapunov function satisfying:

[0350]

[0351] wherein, is a positive number greater than zero.

[0352] After introducing the event-triggered strategy to complete the modification of the reinforcement learning predictive controller, the following Lyapunov function is considered to prove that the introduction of the event-triggered strategy does not affect the stability of the closed-loop system.

[0353]

[0354] If β=0, then according to formula (123), δ≥0; if β≠0, according to formula (123), it can be obtained that:

[0355]

[0356] Substituting formula (127) into formula (124), it can be obtained that

[0357]

[0358] Therefore, it can be obtained that, for all satisfy δ≥0. Therefore, the Lyapunov function V2≥0.

[0359] Differentiating formula (127) and substituting formula (123), (124), (126) into it, it can be obtained that:

[0360]

[0361] Therefore, after introducing the dynamic event trigger strategy, the closed loop system is still stable.

[0362] When formula (123) is not satisfied, the value of t at this time is calculated And the assignment is made t i+1 = t i + nt p , k = k + n Then a new prediction control sequence is updated by reinforcement learning prediction control.

[0363] The application reduces the system computing burden by introducing a dynamic event trigger control strategy, and effectively solves the problem of trajectory optimization control of a six-degree-of-freedom spacecraft under constraints.

[0364] Those skilled in the art know that, in addition to implementing the system provided by the present application and each device, module, unit thereof in the form of pure computer readable program code, the same function can also be achieved by logically programming the method steps in the form of logic gates, switches, application specific integrated circuits, programmable logic controllers and embedded microcontrollers. Therefore, the system provided by the present application and each device, module, unit thereof can be considered as a hardware component, and the devices, modules, units included therein for achieving various functions can also be considered as structures within the hardware component. The devices, modules, units for achieving various functions can also be considered as both software modules for implementing methods and structures within hardware components.

[0365] The specific embodiments of the present application are described above. It should be understood that the present application is not limited to the above specific embodiments, and those skilled in the art can make various changes or modifications within the scope of the claims, which does not affect the essential content of the present application. In the case of no conflict, the embodiments of the present application and the features in the embodiments can be combined with each other arbitrarily.

Claims

1. A spacecraft trajectory optimization method, characterized in that, Includes the following steps: Step 1: Establish an integrated attitude and orbit dynamic model of the spacecraft based on the Lie group SE(3); Step 2: Using the Lie group variational integral method and based on the group properties, the dynamic model in Step 1 is discretized to obtain a discretized dynamic model; Step 3: Using the discretized dynamic model from Step 2, establish the performance index function, transforming the optimization control problem into a problem of finding the extremum of the objective function under constraints; Step 4: Construct a finite-time domain reinforcement learning predictive controller using the discretized dynamic model from Step 2, and solve the transformed optimization control problem from Step 3 using the predictive controller; Step 5: Construct a dynamic event triggering strategy using the discretized dynamic model from Step 2, and modify the predictive controller from Step 4; Step 1 specifically involves: Define two coordinate systems: the Earth's equatorial inertial coordinate system F. I (x I ,y I ,z I ) and spacecraft body coordinate system F b (x b ,y b ,z b C represents the spacecraft's position from coordinate system F. b Switch to coordinate system F I The direction cosine matrix is ​​an element of the Lie group SO(3), where SO(3) is a special orthogonal set satisfying: SO(3)={C∈R 3×3 :C T C = I 3×3 Let det(C) = 1, and R be the set of real numbers. 3×3 Let be a space consisting of 3×3 real matrices, where different superscripts represent the corresponding matrix or vector dimensions. T To find the transpose of a matrix, I 3×3 Given a 3×3 identity matrix, and det() which calculates the determinant of a matrix, the spacecraft's attitude kinematics equations can be expressed as: Where ω = [ω1, ω2, ω3] T ∈R 3×1 In the spacecraft's intrinsic coordinate system, F represents the coordinate system. b Relative to coordinate system F I The attitude angular velocity, with subscripts 1, 2, and 3 indicating the angular velocity components of ω along the three principal inertial axes of the spacecraft. This represents an antisymmetric matrix composed of three-dimensional vectors. Let SO(3) be the Lie algebra; Record the spacecraft in the inertial frame F I The position coordinates below are R, and in this system F b If the orbital velocity is v, then the orbital kinematic equations of the spacecraft are expressed as: The attitude and orbital motion of the spacecraft are described by the Lie group SE(3), where SE(3) is a semi-direct product of SO(3) × R. 3 A group space spanned in a 4×4 homogeneous form has an element g of the following form: Where g represents the spacecraft's attitude configuration, 0 1×3 If the zero vector is 1×3, then the spacecraft attitude-orbit integrated kinematic equations described by the Lie group SE(3) are expressed as: in, The first derivative of the spacecraft configuration, ξ=[ω T ,v T ] T ∈R 6×1 The velocity vector is composed of attitude angular velocity and orbital velocity. Defined as: in, Let be the Lie algebra of SE(3), and R 6×1 isomorphism; This system F b The spacecraft's attitude and orbital dynamics equations are as follows: Where, J∈R 3×3 The moment of inertia of the spacecraft; τ is the first derivative of the angular velocity. c The control torque generated by the spacecraft actuators, d τ f is the external disturbance torque acting on the spacecraft; m is the mass of the spacecraft, f c The control force generated by the spacecraft actuators, d f External disturbances experienced by a spacecraft; The adjoint mapping for the Lie group SE(3) is: ad X Y=[X,Y]=XY-YX (7) Where X and Y are matrices with corresponding dimensions, the above adjoint operator can be expressed in matrix form through algebraic operations: Two elements (3) The left-invariant inner product defined on SE(3) is defined as: The brackets above are: The operator ad defined by formula (10) represents the Lie algebra. The linear operation between SE(3) and its inverse adjoint operator is obtained through the dual operation of Lie algebras and has the following form: Based on formulas (6) and (11), the integrated dynamic equations of spacecraft attitude and orbit are expressed as follows: Where Ξ=diag(J,mI) 3×3 )∈R 6×6 Let be the spacecraft mass characteristic matrix consisting of the moment of inertia J and the mass m. `diag()` performs diagonal matrix operations. The first derivative of velocity, The control vector is composed of control force and torque. This is the interference vector composed of the interfering force and torque; Step 2 specifically involves: To normalize the trajectory changes of both orbital and attitude motions to the same order of magnitude, the orbital motion is defined as follows: R′=R / R m ,v′=v / v m (13) Where R′∈R 3×1 ,v′∈R 3×1 R represents the normalized position and velocity. m ,v m Let be a constant, representing the normalization parameter, to obtain the normalized kinematic and dynamic equations of the spacecraft: The Lie group variational integral method is used to discretize equation (14) to obtain the discretized dynamic model, whose expression is: Where k is the time series, the subscripts k and k+1 of each parameter represent the corresponding parameter values ​​at the discrete time points, h is the discretization time interval, and f is the time series. k ∈SO(3) is C k The group change, J d =0.5tr(J)I 3×3 -J, tr() is used to find the trace of a matrix, defined as f′ ck =f ck / v m The normalized orbital output of the system is given. The system state and input / output are defined as follows: The attitude-orbit integrated dynamic model is processed using variational methods and Lie group properties to obtain a discretized dynamic model; In predictive control, at time T0, the spacecraft's initial orbital and attitude motion is X0, and after the control sequence U... k (k=0,1,2,…) acts on T0+kt p At time t p To predict the time interval, the system output state is Y. k (g k ξ k Let the system output error be: dY k =[dC k dR k ′dω k dv′ k ] in, These represent the actual spacecraft attitude configuration and velocity at that moment. These represent the target attitude configuration and velocity of the spacecraft at that moment, respectively, C dk and R d ′ k Let represent the target attitude and target orbital position of the spacecraft, respectively, and d() be the variation of the corresponding parameters, dY k The continuous-time domain representation of each element is as follows: Where, φ∈R 3×1 As an intermediate variable used to represent dC, its derivative satisfies the formula on the right side of equation (17). dC is equivalently represented by φ. The discrete variational forms of the system state, input, and output in equation (16) are defined as follows: The following discretized dynamic model is obtained, which is easy to design for reinforcement learning predictive controllers: Where f is the Lipschitz function, T0 is the initial time, x0 is the initial state, and A k B k For a matrix with the corresponding dimensions, its specific expression is as follows: Step 3 specifically involves: The objective function to be optimized is designed as follows: Where N is the prediction time length, Q,H,P∈R 12×12 The objective function consists of two parts: (The weight matrix is ​​positive definite and symmetric.) The first part is the cost function after predictive control within a finite time domain. When Q > H, the optimization objective focuses on the settling time; when Q < H, the optimization objective focuses on energy consumption. Part Two Used to evaluate the accuracy of spacecraft's final trajectory optimization tracking after finite-time predictive control, and satisfying the following: Where K∈R 6×12 Ω represents the linear feedback control gain, P is the prediction endpoint penalty matrix, and Ω is the linear feedback control gain. α For a region containing the predicted endpoint, α > 0, and u k =Kx k ; The constraints faced by spacecraft trajectory optimization include: the maximum output force of the spacecraft actuators, the torque constraints of the spacecraft actuators, and the line-of-sight angle constraints of the spacecraft attitude pointing. The mathematical expressions of the constraints are shown in Equations (27) and (28): Where, τ max ,f max These are the upper limits of the control torque and control force outputs, respectively. The line-of-sight constraint is: in the spacecraft's body coordinate system F b One of the line-of-sight axes, whose unit vector is l b The unit position vector of the line connecting the spacecraft and the sun is r. b If the line-of-sight constraint angle is θ, then the following relationship exists between the three: (l b ,r b )≤cosθ (28) The spacecraft trajectory optimization control problem is transformed into the following mathematical problem, which is described in the prediction interval [T0, T0+Nt]. p Within this scope, address the following optimization problem: Where E{} is the expectation operator; Step 4 specifically involves: Define step size [kt] p ,(k+1)t p The cost function within [ ] is: Simplifying formula (30) to r(k), the optimization problem is rewritten as: in, According to Bellman optimization theory, there exists an optimal objective function. Satisfying the following Hamilton-Jacobi-Bellman equations: Solving the above equations yields the optimal control. The specific form is as follows: The numerical solution of the equation is obtained by training using a finite-time reinforcement learning method. Within the prediction control interval [k, N-1], the initial objective function value is defined as J. i=0 ()=0, for i=0,1,… and τ∈[k,N-1], use formula (35) for iterative learning calculation. objective function Update using the following formula: Step 5 specifically involves: Predictive control based on dynamic event triggering is defined as... i According to the system status Predictive control is performed to obtain a control vector. Until the next trigger time t i+1 No new predictive control updates will be performed before the arrival of [t]. i ,t i+1 Within ) control vector It remains unchanged, and the two trigger times satisfy the following relationship: Define state error: in, Design the following dynamic event triggering rules: Where β≥0, σ∈(0,1), and ρ,γ are K ∞ The function δ satisfies: Where μ∈(0,1), δ0≥0, and ||e|| is the second norm of the vector; When formula (39) is not satisfied, calculate the value at this time. Values ​​and assign values. t i+1 =t i +nt p The new predictive control sequence is updated by reinforcement learning predictive control, where k = k + n.

2. A spacecraft trajectory optimization system, characterized in that, Includes the following modules: Module M1: Establishing an integrated dynamic model of spacecraft attitude and orbit based on Lie group SE(3); Module M2: The Lie group variational integral method is used, and based on the group properties, the dynamic model in Module M1 is discretized to obtain a discretized dynamic model. Module M3: Using the discretized dynamic model in Module M2, a performance index function is established, transforming the optimization control problem into a problem of finding the extremum of the objective function under constraints; Module M4: Constructs a finite-time domain reinforcement learning predictive controller based on the discretized dynamics model in Module M2, and solves the transformed optimization control problem in Module M3 through the predictive controller; Module M5: Constructs a dynamic event triggering strategy based on the discretized dynamic model in Module M2, and corrects the predictive controller in Module M4; The module M1 performs the following process: Define two coordinate systems: the Earth's equatorial inertial coordinate system F. I (x I ,y I ,z I ) and spacecraft body coordinate system F b (x b ,y b ,z b C represents the spacecraft's position from coordinate system F. b Switch to coordinate system F I The direction cosine matrix is ​​an element of the Lie group SO(3), where SO(3) is a special orthogonal set satisfying: SO(3)={C∈R 3×3 :C T C = I 3×3 Let det(C) = 1, and R be the set of real numbers. 3×3 Let be a space consisting of 3×3 real matrices, where different superscripts represent the corresponding matrix or vector dimensions. T To find the transpose of a matrix, I 3×3 Given a 3×3 identity matrix, and det() which calculates the determinant of a matrix, the spacecraft's attitude kinematics equations can be expressed as: Where ω = [ω1, ω2, ω3] T ∈R 3×1 In the spacecraft's intrinsic coordinate system, F represents the coordinate system. b Relative to coordinate system F I The attitude angular velocity, with subscripts 1, 2, and 3 indicating the angular velocity components of ω along the three principal inertial axes of the spacecraft. This represents an antisymmetric matrix composed of three-dimensional vectors. Let SO(3) be the Lie algebra; Record the spacecraft in the inertial frame F I The position coordinates below are R, and in this system F b If the orbital velocity is v, then the orbital kinematic equation of the spacecraft can be expressed as: The attitude and orbital motion of the spacecraft are described by the Lie group SE(3), where SE(3) is a semi-direct product of SO(3) × R. 3 A group space spanned in a 4×4 homogeneous form has an element g of the following form: Where g represents the spacecraft's attitude configuration, 0 1×3 If the zero vector is 1×3, then the spacecraft attitude-orbit integrated kinematic equations described by the Lie group SE(3) are expressed as: in, The first derivative of the spacecraft configuration, ξ=[ω T ,v T ] T ∈R 6×1 The velocity vector is composed of attitude angular velocity and orbital velocity. Defined as: in, Let be the Lie algebra of SE(3), and R 6×1 isomorphism; This system F b The spacecraft's attitude and orbital dynamics equations are as follows: Where, J∈R 3×3 The moment of inertia of the spacecraft; τ is the first derivative of the angular velocity. c The control torque generated by the spacecraft actuators, d τ f is the external disturbance torque acting on the spacecraft; m is the mass of the spacecraft, f c The control force generated by the spacecraft actuators, d f External disturbances experienced by a spacecraft; The adjoint mapping for the Lie group SE(3) is: ad X Y=[X,Y]=XY-YX (47) Where X and Y are matrices with corresponding dimensions, the above adjoint operator can be expressed in matrix form through algebraic operations: Two elements (3) The left-invariant inner product defined on SE(3) is defined as: The brackets above are: The operator ad defined by formula (50) represents the Lie algebra. The linear operation between SE(3) and its inverse adjoint operator is obtained through the dual operation of Lie algebras and has the following form: Based on formulas (46) and (51), the integrated dynamic equations of spacecraft attitude and orbit are expressed as follows: Where Ξ=diag(J,mI) 3×3 )∈R 6×6 Let be the spacecraft mass characteristic matrix consisting of the moment of inertia J and the mass m. `diag()` performs diagonal matrix operations. The first derivative of velocity, The control vector is composed of control force and torque. This is the interference vector composed of the interfering force and torque; The module M2 performs the following process: To normalize the trajectory changes of both orbital and attitude motions to the same order of magnitude, the orbital motion is defined as follows: R′=R / R m ,v′=v / v m (53) Where R′∈R 3×1 ,v′∈R 3×1 R represents the normalized position and velocity. m ,v m Let be a constant, representing the normalization parameter, to obtain the normalized kinematic and dynamic equations of the spacecraft: The Lie group variational integral method is used to discretize formula (54) to obtain the discretized dynamic model, whose expression is: Where k is the time series, the subscripts k and k+1 of each parameter represent the corresponding parameter values ​​at the discrete time points, h is the discretization time interval, and f is the time series. k ∈SO(3) is C k The group change, J d =0.5tr(J)I 3×3 -J, tr() is used to find the trace of a matrix, defined as f′ ck =f ck / v m The normalized orbital output of the system is given. The system state and input / output are defined as follows: The attitude-orbit integrated dynamic model is processed using variational methods and Lie group properties to obtain a discretized dynamic model; In predictive control, at time T0, the spacecraft's initial orbital and attitude motion is X0, and after the control sequence U... k (k=0,1,2,…) acts on T0+kt p At time t p To predict the time interval, the system output state is Y. k (g k ξ k Let the system output error be: dY k =[dC k dR k ′dω k dv′ k ] in, These represent the actual spacecraft attitude configuration and velocity at that moment. These represent the target attitude configuration and velocity of the spacecraft at that moment, respectively, C dk and R′ dk Let represent the target attitude and target orbital position of the spacecraft, respectively, and d() be the variation of the corresponding parameters, dY k The continuous-time domain representation of each element is as follows: Where, φ∈R 3×1 As an intermediate variable used to represent dC, its derivative satisfies the formula on the right side of equation (57). dC is equivalently represented by φ. The discrete variational forms of the system state, input, and output in equation (56) are defined as follows: The following discretized dynamic model is obtained, which is easy to design for reinforcement learning predictive controllers: Where f is the Lipschitz function, T0 is the initial time, x0 is the initial state, and A k B k For a matrix with the corresponding dimensions, its specific expression is as follows: The module M3 performs the following process: The objective function to be optimized is designed as follows: Where N is the prediction time length, Q,H,P∈R 12×12 The objective function consists of two parts: (The weight matrix is ​​positive definite and symmetric.) The first part is the cost function after predictive control within a finite time domain. When Q > H, the optimization objective focuses on the settling time; when Q < H, the optimization objective focuses on energy consumption. Part Two Used to evaluate the accuracy of spacecraft's final trajectory optimization tracking after finite-time predictive control, and satisfying the following: Where K∈R 6×12 Ω represents the linear feedback control gain, P is the prediction endpoint penalty matrix, and Ω is the linear feedback control gain. α For a region containing the predicted endpoint, α > 0, and u k =Kx k ; The constraints faced by spacecraft trajectory optimization include: the maximum output force of the spacecraft actuators, the torque constraints of the spacecraft actuators, and the line-of-sight angle constraints of the spacecraft attitude pointing. The mathematical expressions of the constraints are shown in formulas (67) and (68): Where, τ max ,f max These are the upper limits of the control torque and control force outputs, respectively. The line-of-sight constraint is: in the spacecraft's body coordinate system F b One of the line-of-sight axes, whose unit vector is l b The unit position vector of the line connecting the spacecraft and the sun is r. b If the line-of-sight constraint angle is θ, then the following relationship exists between the three: (l b ,r b )≤cosθ (68) The spacecraft trajectory optimization control problem is transformed into the following mathematical problem, which is described in the prediction interval [T0, T0+Nt]. p Within this scope, address the following optimization problem: Where E{} is the expectation operator; The module M4 performs the following process: Define step size [kt] p ,(k+1)t p The cost function within [ ] is: Simplifying formula (70) to r(k), the optimization problem is rewritten as: in, According to Bellman optimization theory, there exists an optimal objective function. Satisfying the following Hamilton-Jacobi-Bellman equations: Solving the above equations yields the optimal control. The specific form is as follows: The numerical solution of the equation is obtained by training using a finite-time reinforcement learning method. Within the prediction control interval [k, N-1], the initial objective function value is defined as J. i=0 ()=0, for i=0,1,… and τ∈[k,N-1], use formula (75) for iterative learning calculation. objective function Update using the following formula: The module M5 performs the following process: Predictive control based on dynamic event triggering is defined as... i According to the system status Predictive control is performed to obtain a control vector. Until the next trigger time t i+1 No new predictive control updates will be performed before the arrival of [t]. i ,t i+1 Within ) control vector It remains unchanged, and the two trigger times satisfy the following relationship: Define state error: in, Design the following dynamic event triggering rules: Where β≥0, σ∈(0,1), and ρ,γ are K ∞ The function δ satisfies: Where μ∈(0,1), δ0≥0, and ||e|| is the second norm of the vector; When formula (79) is not satisfied, calculate the value at this time. Values ​​and assign values. t i+1 =t i +nt p The new predictive control sequence is updated by reinforcement learning predictive control, where k = k + n.

3. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the steps of the spacecraft trajectory optimization method of claim 1.

4. An electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the computer program is executed by the processor, it implements the steps of the spacecraft trajectory optimization method of claim 1.

Citation Information

Patent Citations

  • Spacecraft low-thrust perturbation rendezvous trajectory optimization method and system

    CN112084581A

  • Reinforcement learning attitude constraint control method considering installation deviation of actuating mechanism

    CN112357120A

  • Spacecraft rendezvous and docking trajectory planning event constraint convexity method

    CN113722821A

  • Spacecraft attitude control method based on model predictive control and sliding mode control

    CN113859589A

  • Spacecraft autonomous rendezvous and docking guidance strategy generation method based on reinforcement learning

    CN114036631A