A method for optimal control of attitude of a space non-cooperative target manipulation platform through reinforcement learning
By designing a robust near-optimal controller, the attitude control problem of a space non-cooperative target manipulation platform under external disturbances was solved, achieving low-energy, high-precision, and highly robust attitude control, thus improving the system's stability and control performance.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- NORTHWESTERN POLYTECHNICAL UNIV
- Filing Date
- 2022-11-21
- Publication Date
- 2026-04-17
AI Technical Summary
Existing technologies struggle to achieve optimal control with low energy consumption, high precision, and high robustness under external time-varying disturbances in the attitude control of space non-cooperative target manipulation platforms, especially the control performance of critic neural networks is limited under continuous excitation conditions.
A reinforcement learning-based optimal control method for attitude is designed. By relaxing the continuous excitation conditions, adopting the weight update law of the critical neural network, and adding a robust term to suppress external disturbances, a robust near-optimal controller is constructed to achieve the stability and optimality of the attitude control system.
Ensuring the optimality and robustness of the attitude control system in the presence of external disturbances improves the attitude control performance of the space non-cooperative target manipulation platform.
Smart Images

Figure CN116482969B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of attitude control of space non-cooperative target manipulation platforms based on reinforcement learning, and particularly relates to an optimal attitude control method for space non-cooperative target manipulation platforms based on reinforcement learning. Background Technology
[0002] Non-cooperative space targets refer to space objects that cannot provide valid information about themselves or interact with other space targets, such as malfunctioning spacecraft, space debris, and junk. In recent years, with the continuous development of space technology and the rapid increase in space exploration activities, the number of non-cooperative space targets has increased. These targets have lost ground control over them, posing a potential threat to spacecraft in orbit. On the other hand, the increase in non-cooperative space targets also squeezes space resources and hinders subsequent launch missions. Therefore, in the space domain, how to safely and effectively capture and deal with non-cooperative space targets has become an important issue. To accomplish operations such as capturing, clearing, and attacking / defending against non-cooperative space targets, high-precision, low-energy attitude control is required for a space non-cooperative target control platform equipped with optical payloads, robotic arms, and lasers.
[0003] High-performance attitude maneuvering control of space non-cooperative target control platforms enables high-precision orientation of camera optical axes, robotic arms, lasers, and other components mounted on the platform. This allows for tasks such as photographic capture, space monitoring, space debris capture or removal, and space offensive and defensive operations. In particular, with recent technological advancements and evolving international dynamics, space non-cooperative target control platforms have gained significant strategic importance in both offensive and defensive capabilities. Therefore, space non-cooperative target control platforms have become a crucial development direction for various countries in the aerospace field. Their attitude control is fundamental to achieving the aforementioned space missions, prompting researchers to study their low-energy, high-precision, and robust attitude control systems.
[0004] However, achieving low-energy-consumption optimal control methods is hampered by the difficulty in solving the Hamilton-Jacobi-Bellman equations. To address this issue, invention patent CN113219842B designed an approximate optimal control law based on reinforcement learning, utilizing a critic neural network approximator to approximate the optimal control law. However, this method designs the weight update law of the critic neural network under the strict assumption of continuous excitation, and this optimal control strategy has a weak ability to suppress time-varying disturbances. When the controlled system is affected by external time-varying disturbances, its control performance will be greatly weakened. Summary of the Invention
[0005] To achieve optimal attitude control for a space non-cooperative target manipulation platform, the present invention aims to provide an attitude reinforcement learning-based optimal control method for such a platform, thereby balancing the relationship between control cost and control performance. Furthermore, the designed critic neural network weight update law can relax the continuous excitation conditions in traditional reinforcement learning-based optimal control methods.
[0006] To achieve the above-mentioned objectives, the technical solution adopted by this invention is as follows:
[0007] An optimal control method for attitude reinforcement learning of a space non-cooperative target manipulation platform includes the following steps:
[0008] Step 1: The actuator provides effective torque to the space non-cooperative target control platform, which serves as the input to the attitude control system. The output of the attitude control system is the attitude correction Rodrigues parameter. and attitude angular velocity All data are obtained through sensitive elements, thereby establishing a mathematical model of the attitude kinematics and dynamics of a space non-cooperative target control platform subjected to external time-varying disturbances;
[0009] Step 2, introduce augmenting variables Includes attitude correction Rodrigue parameters and attitude angular velocity Then, the nominal system of the attitude kinematics and dynamics mathematical model in step 1 is represented as a compact vector system, and the objective function and optimal control law are designed.
[0010] Step 3: Design a weight update law for the critic neural network with relaxed continuous excitation conditions to obtain an approximately optimal control law that minimizes the objective function; the resulting critic neural network weights... The update law is designed as follows:
[0011]
[0012] Where β, η0, and ζ all represent positive constants. Let be a positive definite diagonal matrix, where variable μ is a small positive constant, and tanh(·) denotes the hyperbolic tangent function. Let σ(x) represent the partial derivative of the basis function σ(x) with respect to x. u n For an approximate optimal control law, e ε This is the Bellman error term. Represents a positive definite constant matrix The inverse matrix, g(x) = [0, 3, J -1 ] TLet f(x) be the augmented distribution matrix, f(x) be the systematic term, and 03 denote a 3×3 all-zero matrix. The positive definite rotational inertia matrix of the non-cooperative target control platform in space. The inverse matrix;
[0013] Step 4: Based on the near-optimal control law obtained in Step 3, a robust term is added to suppress external disturbances and improve the robustness of the attitude control system; the robust near-optimal controller makes the attitude control system both robust and optimal; the resulting stable robust near-optimal controller u c for:
[0014]
[0015] in, This represents an approximately optimal control law. Both k and c are positive constants, and k is greater than the upper bound of the magnitude of the external disturbance torque. The optimal weights W for the critic neural network * The estimated value.
[0016] Furthermore, in step 1, a mathematical model of the attitude kinematics and dynamics of the space non-cooperative target control platform subjected to external time-varying disturbances is established. The specific steps are as follows:
[0017]
[0018]
[0019] Where, G(ρ)=(I3-ρ × +ρρ T -(1+ρ T ρ)I3 / 2) / 2, This represents the positive definite moment of inertia matrix of the non-cooperative target control platform in space. This represents the disturbance torque experienced by the control platform for a non-cooperative target in space. Representing robust approximations Optimal Controller , and These represent the corrected Rodrigues attitude and attitude angular velocity of the space non-cooperative target control platform, respectively, with the symbol (·). × This represents an antisymmetric matrix.
[0020] Furthermore, step 2 represents the nominal system of the attitude kinematics and dynamics mathematical model from step 1 as a compact vector system, and designs the objective function and optimal control law. The specific steps are as follows:
[0021] Based on augmented variables The compact vector system representation of the nominal system of attitude kinematics and dynamics mathematical model is then:
[0022]
[0023] f(x) = [ω, -J -1 ω × Jω] T
[0024] g(x) = [0, 3, J -1 ] T
[0025] in, u represents the numerical derivative of the augmented variable x. n This represents an approximately optimal control law;
[0026] For a compact vector system, the objective function is expressed as:
[0027]
[0028] in, and All represent positive definite constant matrices, and τ represents the integration variable;
[0029] The optimal control law for:
[0030]
[0031] Among them, V * (x) is the objective function to be minimized. V represents * The partial derivative of (x) with respect to x, i.e.
[0032] Further, step 3 includes: using a critic neural network to describe the minimized objective function as follows:
[0033] V * (x)=W *T σ(x)+ε(x)
[0034] in, It is the set of basis function vectors, and satisfies σ(0)=0; Let ε(x) represent the optimal weight vector of the basis functions, and let ε(x) be the approximation error of the critic neural network.
[0035] The optimal control law can be rewritten as:
[0036]
[0037] Wherein, the partial derivatives of the basis function σ(x) with respect to x This represents the partial derivative of the critique neural network approximation error ε(x) with respect to x;
[0038] The objective function approximated using a critic neural network structure is expressed as:
[0039]
[0040] The optimal control law approximated using a critic neural network, i.e., the approximate optimal control law, is expressed as:
[0041]
[0042] The attitude reinforcement learning optimal control method for a space non-cooperative target manipulation platform proposed in this invention adopts the above technical solution, and its advantages compared with the prior art are as follows:
[0043] (1) The critical neural network weight update law proposed in this invention can relax the continuous excitation conditions in the traditional optimal control method based on reinforcement learning.
[0044] (2) A robust approximate optimal controller based on reinforcement learning can enable the attitude control system of a space non-cooperative target manipulation platform to achieve consistent final bounded stability under external disturbances.
[0045] (3) A robust term is added to the traditional reinforcement learning-based optimal control law to suppress time-varying external disturbance torque, thus proposing a new stable reinforcement learning-based robust approximate optimal controller, which improves the robustness of the attitude control system while ensuring the optimality of the system. Attached Figure Description
[0046] Figure 1 This is a flowchart of an attitude reinforcement learning optimal control method for a space non-cooperative target manipulation platform according to the present invention.
[0047] Figure 2 This is a block diagram of the optimal attitude control structure for a space non-cooperative target manipulation platform based on reinforcement learning.
[0048] Figure 3 This is the weight update effect of the critic neural network in this embodiment of the invention;
[0049] Figure 4 This refers to the control torque of the space non-cooperative target manipulation platform in this embodiment of the invention.
[0050] Figure 5 The convergence effect of the attitude correction Rodrigues parameters of the space non-cooperative target control platform in this embodiment of the invention. Detailed Implementation
[0051] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention. Furthermore, the technical features involved in the various embodiments of this invention described below can be combined with each other as long as they do not conflict with each other.
[0052] like Figure 1 As shown, the attitude reinforcement learning optimal control method for a space non-cooperative target manipulation platform of the present invention includes the following steps:
[0053] Step 1: Establish a mathematical model of the attitude kinematics and dynamics of a space non-cooperative target control platform that considers external disturbances;
[0054] Step 2: Rewrite the nominal system of the attitude kinematics and dynamics mathematical model into a compact vector system, and design the objective function and optimal control law;
[0055] Step 3: Design a critic neural network structure to approximate the objective function, and propose a new neural network weight law that can relax the continuous excitation conditions;
[0056] Step 4: Considering the influence of external time-varying disturbances, based on the output of the critic neural network, an approximate optimal control law based on reinforcement learning is obtained. A robust optimal controller is designed to suppress external disturbances, so that the attitude control system of the space non-cooperative target control platform has both optimality and strong robustness.
[0057] Specifically, the steps for establishing the attitude kinematics and dynamics models of the space non-cooperative target control platform considering external disturbances in step 1 are as follows:
[0058]
[0059]
[0060] in, This represents the positive definite moment of inertia matrix of the non-cooperative target control platform in space. This represents the disturbance torque experienced by the control platform for a non-cooperative target in space. Representing robust approximations Optimal Controller . and These represent the corrected Rodrigues attitude and attitude angular velocity of the space non-cooperative target control platform, respectively. (·) × This represents an antisymmetric matrix.
[0061] In this embodiment, the positive definite moment of inertia matrix of the space non-cooperative target control platform is:
[0062]
[0063] The disturbance torque experienced by the space non-cooperative target control platform is expressed as:
[0064]
[0065] Where · represents the Euclidean norm of the vector, t represents time, sin(·) is the sine function, and cos(·) is the cosine function.
[0066] Step 2, which represents the nominal system of attitude kinematics and dynamics mathematical models as a compact vector system, is as follows:
[0067] Define augmented variables The compact vector system is then represented as:
[0068]
[0069] f(x) = [ω, -J -1 ω × Jω] T
[0070] g(x) = [0, 3, J -1 ] T
[0071] in, Let f(x) represent the numerical derivative of the augmented variable x, f(x) represent the systematic term, and g(x) = [0, 3, J]. -1 ] T Represents the augmented allocation matrix. The positive definite rotational inertia matrix of the non-cooperative target control platform in space. The inverse matrix, 03 represents a 3×3 all-zero matrix, u n This represents an approximately optimal control law.
[0072] For a compact vector system, the objective function is expressed as:
[0073]
[0074] Where Q = diag([2 2 2 2 2 2]) and R = diag([0.21 0.21 0.21]) are both positive definite constant matrices. τ is the integration variable. The objective function to be minimized is expressed as:
[0075]
[0076] For the minimized objective function V * Taking the derivative of both sides with respect to time, we obtain the HJB equation:
[0077]
[0078] in, This represents the optimal control law. V represents * The partial derivative of (x) with respect to x, i.e.
[0079] Furthermore, solving the HJB equation yields the optimal control law as follows:
[0080]
[0081] Among them, R -1 =diag([1 / 0.21 1 / 0.21 1 / 0.21]).
[0082] Step 3 includes: designing an objective function V that is minimized using a critical neural network. * (x) can be written as:
[0083] V * (x)=W *T σ(x)+ε(x)
[0084] in, It is a basis function vector, which is chosen in this embodiment. Let ε(x) represent the unknown optimal weight vector corresponding to the basis function, and let ε(x) be the approximation error of the critic neural network.
[0085] Then, the optimal control law is rewritten as:
[0086]
[0087] Where, the partial derivative of the basis function σ(x) with respect to x is This represents the partial derivative of the critique neural network approximation error ε(x) with respect to x.
[0088] Furthermore, the objective function approximated using the critic neural network structure is expressed as:
[0089]
[0090] in, The optimal weights W for the critic neural network * The estimated value.
[0091] The optimal control law approximated using a critic neural network is expressed as follows:
[0092]
[0093] in, The optimal weights W for the critic neural network * The estimated value can be obtained at this point:
[0094]
[0095] in, Let ε0 represent the weight estimation error of the critic neural network, and ε1 represent the residual term of the approximation error of the critic neural network. Therefore, the weight update law of the critic neural network is designed as follows:
[0096]
[0097] in, Let represent a positive definite diagonal matrix. variable tanh(·) denotes the hyperbolic tangent function. β = 0.001, η0 = 0.00002, ζ = 0.5. In this embodiment, the initial weights of the critic neural network are W0 = [25 40 30 20 12 25]. T .
[0098] In step 4, the robust near-optimal controller u of the attitude stabilization control system of the space non-cooperative target manipulation platform c Designed as follows:
[0099]
[0100] in, Both k = 0.8 and c = 0.002 are positive values, and k = 0.02 is greater than the upper limit of the external disturbance torque magnitude.
[0101] Using the robust approximate optimal controller given above, the attitude control system of the entire space non-cooperative target control platform can be guaranteed to be uniformly and ultimately boundedly stable under the condition of being subjected to external disturbance torque.
[0102] like Figure 2 As shown, the optimal attitude control method for a space non-cooperative target control platform of the present invention needs to be implemented by several parts, including a critical neural network, an approximate optimal control law, a robust term, a robust approximate optimal controller, and a mathematical model of the attitude of the space non-cooperative target control platform.
[0103] In this invention, a critical neural network approximates the objective function using the system's state information to obtain an approximate optimal control law. Considering the impact of external disturbances on the on-orbit non-cooperative target control platform, a robust approximate optimal controller is proposed using the concept of robust control. The control signal is transmitted to the actuator to provide actual control torque to the non-cooperative target control platform, achieving attitude control. The angular velocity of the non-cooperative target control platform, output by the platform's dynamic model, can be measured by a gyroscope, while the attitude, output by the kinematic model, can be measured by an attitude sensor. Ultimately, this invention proposes a reinforcement learning-based optimal control method for the attitude of a non-cooperative target control platform, ensuring the optimality and robustness of the attitude control system even under disturbances.
[0104] The invention is further illustrated below with specific embodiments, and numerical simulations are performed to verify the robust approximate optimal controller proposed in this invention. Simulation results are as follows: Figure 3-5 As shown, Figure 3 This invention demonstrates the estimation of the optimal weights for a critic neural network using a novel weight update law. Figure 4 The required control torque for the attitude maneuvering of a space non-cooperative target control platform is demonstrated. Figure 5 The convergence effect of attitude correction Rodrigues parameters for a space non-cooperative target control platform is demonstrated. The effectiveness of this embodiment and the robust near-optimal controller designed in this invention is verified using digital simulation software.
[0105] Those skilled in the art will readily understand that the above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
Claims
1. A posture reinforcement learning optimal control method for a space non-cooperative target manipulation platform, characterized in that: Includes the following steps: Step 1: The actuator provides effective torque to the space non-cooperative target control platform, which serves as the input to the attitude control system. The output of the attitude control system is the attitude correction Rodrigues parameter. and attitude angular velocity All data are obtained through sensitive elements, thereby establishing a mathematical model of the attitude kinematics and dynamics of a space non-cooperative target control platform subject to external time-varying disturbances. Step 2, introduce augmenting variables Includes attitude correction Rodrigues parameters and attitude angular velocity Then, the nominal system of the attitude kinematics and dynamics mathematical model in step 1 is represented as a compact vector system, and the objective function and optimal control law are designed. Step 3: Design a weight update law for the critic neural network with relaxed continuous excitation conditions, thereby obtaining an approximately optimal control law that minimizes the objective function; the resulting critic neural network weights... The update law is designed as follows: ; in, , and All represent positive numbers. Let be a positive definite diagonal matrix, where ,variable , It is a small positive number. Represents the hyperbolic tangent function. Describing basis functions right The partial derivative, , This is an approximately optimal control law. This is the Bellman error term. Represents a positive definite constant matrix The inverse matrix, For the augmented allocation matrix, For system items, express A matrix of all zeros The positive definite rotational inertia matrix of the non-cooperative target control platform in space. The inverse matrix; Step 4: Based on the near-optimal control law obtained in Step 3, a robust term is added to suppress external disturbances and improve the robustness of the attitude control system; the robust near-optimal controller makes the attitude control system both robust and optimal; the resulting stable robust near-optimal controller for: ; in, This represents an approximately optimal control law. , and All are positive numbers and satisfy the following conditions: Greater than the upper limit of the external disturbance torque magnitude, Optimal weights for the critic neural network The estimated value.
2. The attitude reinforcement learning optimal control method for a space non-cooperative target manipulation platform according to claim 1, characterized in that: In step 1, a mathematical model of the attitude kinematics and dynamics of a space non-cooperative target control platform subjected to external time-varying disturbances is established. The specific steps are as follows: ; in, , , This represents the positive definite moment of inertia matrix of the non-cooperative target control platform in space. This represents the disturbance torque experienced by the control platform for a non-cooperative target in space. This indicates a robust, approximately optimal controller. and These represent the corrected Rodrigues attitude and attitude angular velocity of the space non-cooperative target control platform, respectively, with symbols... This represents an antisymmetric matrix.
3. The attitude reinforcement learning optimal control method for a space non-cooperative target manipulation platform according to claim 2, characterized in that: Step 2 represents the nominal system of the attitude kinematics and dynamics mathematical model from Step 1 as a compact vector system, and designs the objective function and optimal control law. The specific steps are as follows: Based on augmented variables Then, the compact vector system of the nominal system of the attitude kinematics and dynamics mathematical model is represented as: ; in, Represents augmented variables The numerical derivative, This represents an approximately optimal control law; For a compact vector system, the objective function is expressed as: ; in, and Both represent positive definite constant matrices. Represents the integral variable; The optimal control law for: ; in, To minimize the objective function, express right The partial derivatives of, i.e. .
4. The attitude reinforcement learning optimal control method for a space non-cooperative target manipulation platform according to claim 3, characterized in that: Step 3 includes: using a critic neural network to describe the minimized objective function as follows: ; in, It is the set of basis function vectors, and satisfies ; This represents the optimal weight vector of the basis functions. It is the approximation error of the critic neural network; The optimal control law can be rewritten as: ; Wherein, basis functions right partial derivatives , Indicates the approximation error of the critic neural network. right The partial derivative; The objective function approximated using a critic neural network structure is expressed as: ; The optimal control law approximated using a critic neural network, i.e., the approximate optimal control law, is expressed as: 。
Citation Information
Patent Citations
Secondary capture control strategy for space instability non-cooperative targets
CN107102548A
Reinforcement learning based adaptive control method for small unmanned helicopter
CN109696830A