Spacecraft attitude control method based on high-order nonlinear all-wheel-drive system
By establishing a high-order nonlinear full drive system model and evaluating reward weights using inverse optimal control method and designing an optimal controller, the spacecraft attitude control problem of high-order nonlinear systems is solved, and the system stability and performance optimization is achieved.
Patent Information
- Application Number
- CN202510491929.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-18
- Publication Date
- 2025-07-29
AI Technical Summary
The prior art is difficult to effectively solve the spacecraft attitude control problem of high-order nonlinear full drive systems. Especially under factors such as multi-body coupling and flexible component vibration, the nonlinear characteristics of the system are complex, resulting in increased control difficulty. The existing methods often only ensure local stability, which limits practical applications.
By establishing a high-order nonlinear full drive system model, using the inverse optimal control method to evaluate unknown reward weights, design the optimal controller, and adjust the weights through iterative optimization algorithm to minimize performance indicators, and combine the observation data of the expert system to track the trajectory, simplify the controller design process.
The spacecraft attitude controller design with a high-order nonlinear full drive system is realized, the control performance is optimized, the controller design process is simplified, the system is stable and the expert system trajectory is tracked, and the control effect is improved.
Smart Images

Figure BDA0005365837390000025 
Figure BDA0005365837390000044 
Figure BDA0005365837390000052
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of spacecraft attitude control, and particularly relates to a spacecraft attitude control method based on a high-order nonlinear fully actuated system. Background Art
[0002] With the diversification and complexity of space missions, the spacecraft attitude control system needs to consider many factors, such as multi-body coupling, flexible component vibration, etc. This makes the nonlinear characteristics of the system more complex, facing the optimal control problem of a high-order nonlinear fully actuated system, increasing the difficulty of control. How to design a spacecraft controller for a high-order nonlinear fully actuated system and optimize the system performance index is a challenge.
[0003] Nonlinear control theory has attracted the attention of the industrial and academic circles due to its complexity and wide applicability. Among many past methods, the state-space method has played a key role, and many classical nonlinear control theories and methods have thus developed. However, the state-space method takes state variables as the core, and the obtained controller often only guarantees local stability, which limits their application in actual engineering.
[0004] Current research is mainly based on first-order nonlinear systems. High-order nonlinear fully actuated systems have challenges such as high dimension and complex calculation. The inverse optimal control of high-order nonlinear fully actuated systems has not been explored. Due to high-order nonlinear fully actuated systems having more state variables and complex dynamic couplings, and a larger dimension of the state matrix, the computational difficulty of solving the optimal control problem of such systems has increased significantly. Summary of the Invention
[0005] In view of the deficiencies of the prior art, the present invention proposes a spacecraft attitude control method based on a high-order nonlinear fully actuated system, which is realized through the following steps:
[0006] A spacecraft attitude control method based on a high-order nonlinear system, comprising the following steps:
[0007] Step 1, establish a high-order nonlinear fully actuated system model for spacecraft attitude control;
[0008] Step 2, through the inverse optimal control method, use the observation data of the expert system to evaluate the unknown reward weight in the performance index, and design an optimal controller based on the reward weight;
[0009] Step 3, obtain the control gain error between the learning system and the expert system;
[0010] Step 4, adjust the reward weight through an iterative optimization algorithm to minimize the performance index of the controller and track the trajectory of the expert system.
[0011] Preferably, the high-order nonlinear fully actuated system model for spacecraft attitude control described in step 1 is:
[0012]
[0013] where the state vector is the attitude angle of the spacecraft, φ is the roll angle, θ is the pitch angle, is the yaw angle, u is the control input, is the nonlinear term including centrifugal force and other external disturbances, is the control input matrix of the fully actuated system.
[0014] Preferably, the performance index described in step 2 is:
[0015]
[0016] where δ e is the expert state vector, v e is the expert control input, Q e ≥0 is the symmetric positive semi-definite weight matrix of the expert system, R e >0 is the symmetric positive definite weight matrix of the expert system.
[0017] Preferably, the optimal controller described in step 2 is:
[0018]
[0019] where, is the optimal expert control input, is the control gain of the optimal controller of the expert system, is the positive definite symmetric matrix of the expert system, and B is the control input matrix of the linear system.
[0020] Preferably, step 3 specifically includes:
[0021] Based on the observation data of the expert system, approximately obtain the estimated control gain of the expert system by batch least squares method, and obtain the error between the control gain of the learning system and the estimated control gain of the expert system.
[0022] Preferably, the formula for approximately obtaining the estimated control gain of the expert system by batch least squares method is:
[0023]
[0024] where, is the estimated control gain of the expert system, δ e is the expert state vector, and v e is the expert control input.
[0025] Preferably, the error between the obtained learning system control gain and the estimated expert system control gain is given by the formula:
[0026]
[0027] where K is the control gain of the optimal controller for the high-order nonlinear fully actuated system, is the control gain error between the control gain and the estimated expert system, is the control gain of the estimated expert system, v e is the expert control input, is the norm of the control gain error between the control gain and the estimated expert system, which is a scalar.
[0028] Preferably, the iterative optimization algorithm described in step 4 specifically includes:
[0029] Set the initial weight matrix, learning rate, and penalty coefficient;
[0030] Calculate the control gain of the current iteration through the Bellman equation;
[0031] Update the weight matrix based on error minimization until the convergence condition is met.
[0032] Preferably, the formula for updating the weight matrix is:
[0033] Q i+1 =-A T g(P i )-g(P i )A+g(P i )BR -1 B T g(P i )+λ(Q e -Q i ) (39)
[0034] where A is the state matrix of the linear system, B is the control input matrix of the linear system, λ(Q e -Q i ) is the regularization term, g(P i ) is the reward weight correction factor, and R is a symmetric positive definite weight matrix.
[0035] Preferably, the convergence condition is that the norm of the error between the learning system control gain and the estimated expert system control gain is less than the set threshold.
[0036] Advantages of the present invention:
[0037] 1) By applying the full - drive system method, the design of the spacecraft attitude high - order nonlinear full - drive system controller is simplified, and at the same time, the control performance of the nonlinear system is optimized;
[0038] 2) The inverse optimal control algorithm is proposed. By observing the data of the expert system to estimate the reward weight, the controller design process is significantly simplified, and the optimization of the performance index under unknown weights is achieved. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] Figure 1 It is a flowchart of the spacecraft attitude control method based on the high - order nonlinear full - drive system according to an embodiment of the present invention;
[0040] Figure 2 It is a curve graph of the reward weight and the control strategy learning process according to an embodiment of the present invention;
[0041] Figure 3 It is a trajectory graph of the learning system state during the operation of the spacecraft according to an embodiment of the present invention;
[0042] Figure 4 It is a trajectory graph of the expert system state during the operation of the spacecraft according to an embodiment of the present invention;
[0043] Figure 5 It is a curve graph of the learning process of the high - order nonlinear full - drive system control input during the operation of the spacecraft according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0044] In order to enable those skilled in the art to better understand the technical solutions of the present invention, the present invention will be further described in detail below with reference to the accompanying drawings.
[0045] In the embodiments of the present invention, for the spacecraft attitude control method based on the high - order nonlinear full - drive system, the flowchart of the method is as Figure 1 shown, and includes the following steps:
[0046] Step 1: Based on the full - drive system method, establish a high - order nonlinear full - drive system model for spacecraft attitude control, and transform the nonlinear optimal control problem into a linear optimal control problem. Specifically:
[0047] The dynamic model of the spacecraft attitude control system can be written in the following second - order matrix form:
[0048]
[0049] Among them, the state vector is the attitude angle of the spacecraft, φ is the roll angle, θ is the pitch angle, is the yaw angle, u is the control input, is the inertia matrix of the system, I x$I_x$ is the moment of inertia of the spacecraft about the x-axis, $I$ y $I_y$ is the moment of inertia of the spacecraft about the y-axis, $I$ z $I_z$ is the moment of inertia of the spacecraft about the z-axis, $T$ is the gravity gradient or other external force term, $\omega_0$ is the angular velocity of the Earth's rotation, $D$ is the damping matrix;
[0050] where the $\pi$ function is:
[0051]
[0052] The spacecraft attitude control system is written in the form of a standard high-order fully actuated system as:
[0053]
[0054] where,
[0055]
[0056] where, $N$ is the nonlinear term containing centrifugal force and other external disturbances, $B$ is the control input matrix of the fully actuated system, and the value range of $\varphi$ is
[0057] Generalize Equation (2) to the general form of high-order nonlinearity:
[0058]
[0059] where $u\in R$ r is the control input, $\gamma\in R$ r is the state vector, $R$ r is the r-dimensional real space, $f(\cdot)$ is the nonlinear term vector function of the spacecraft attitude control system, and $G(\cdot)$ is an invertible matrix function. For Equation (2), $n = 2$ at this time.
[0060] Design a linear quadratic optimal controller:
[0061] For Equation (2), the linear quadratic optimal controller is designed as:
[0062]
[0063] $v=-K\gamma$ (0~1) (8)
[0064] where $v$ is the linear term control input of the high-order nonlinear fully actuated system, $K$ is the control gain of the optimal controller of the high-order nonlinear fully actuated system, and $\gamma$ (0~1) is the state trajectory of the learner.
[0065] Substitute Formulas (3) to (7) into Formula (2) to obtain the closed-loop system:
[0066]
[0067] Problem 1: Design the optimal controller u for the nonlinear system * to stabilize the spacecraft attitude control system of Formula (1) and minimize the performance index:
[0068]
[0069] where J1 is the performance index of the high-order nonlinear fully actuated system, is the transpose of the 0th to (n - 1)th order vector of state γ, δ is the 0th to (n - 1)th order vector of state γ, Q is an unknown symmetric semi-positive definite weight matrix, and R is a symmetric positive definite weight matrix.
[0070] The feedback controller is designed as:
[0071] u * = G -1 (δ T )[v - f(δ T )] (11)
[0072] And
[0073] v = γ (n) (12)
[0074] It can be seen from (12) that:
[0075]
[0076] where, is the state matrix of the linear system, is the control input matrix of the linear system;
[0077] Problem 2: Find the optimal controller v for the linear system * to ensure the stability of the linear spacecraft attitude control system of Formula (13) and minimize the performance index J2 of the following linear system:
[0078]
[0079] For the linear system of Formula (13), assuming v * = -Kδ, the value function V(δ) following the quadratic form can be established:
[0080] V(δ) = δ T Pδ (15)
[0081] where P is a positive definite symmetric matrix of the control strategy of the optimal controller.
[0082] Based on the principle of dynamic programming, the Bellman equation H(δ, v) of the student system is derived as follows:
[0083] H(δ, v) = (Aδ + Bv) T Pδ + δ T P(Aδ + Bv) + δ T Qδ + v T Rv = 0 (16)
[0084] Let The optimal controller can be derived as:
[0085] v * = -Kδ = -R -1 B T Pδ (17)
[0086] Substitute equation (17) into equation (16), and the algebraic Riccati equation is generated:
[0087] A T P + PA + Q - PBR -1 B T P = 0 (18);
[0088] The optimal controller v * = -Kδ and the control gain K in equations (17) and (18) and the positive definite symmetric matrix P as the control strategy of the optimal controller are the optimal solutions of problem 2 under arbitrarily given symmetric semi - positive definite weight matrix Q and symmetric positive definite weight matrix R.
[0089] Substitute equation (17) into equation (11) to get
[0090] u * = -B -1 (δ T )[Kδ + f(δ T )] (19)
[0091] Since v = γ (n) , then we get
[0092] J1 = J2 (20)
[0093] The fully - actuated system method transforms the non - linear optimal control problem 1 into a linear optimal control problem 2, making the design of the optimal controller v * become simple, which is equivalent to the design of u * in problem 1.
[0094] Step 2: Define an expert system. Through the inverse optimal control method, use its observation data to evaluate the unknown reward weights in the performance index, and design an optimal controller to optimize the performance index. Specifically:
[0095] According to the formula (13) of the linear spacecraft attitude control system, the expert system is defined as follows:
[0096]
[0097] where δ e ∈R n is the expert state vector, v e ∈R n is the expert control input, and R n is the n-dimensional real number space.
[0098] The performance index J of the expert system e is as follows:
[0099]
[0100] where Q e ≥0 is the symmetric positive semi-definite weight matrix of the expert system, and R e >0 is the symmetric positive definite weight matrix of the expert system.
[0101] At this time, it is assumed that Q e is unknown and R e is known. The control strategy v e =-K e δ e is adopted. The quadratic form of the cost function of the expert system formula (21) is:
[0102]
[0103] where P e is the positive definite symmetric matrix, and K e is the control gain of the expert system.
[0104] According to the optimal control theory, the Bellman equation H e (δ e , v e ) of the expert system is:
[0105] H e (δ e , v e )=(Aδ e +Bv e ) T P e δ e +δ e T P e (Aδ e +Bv e )+δ e T Q e δ e+v e T R e v e =0 (24)
[0106] Satisfied The obtained optimal controller is as follows:
[0107]
[0108] Wherein, is the optimal expert control input, is the control gain of the optimal controller of the expert system, is the positive definite symmetric matrix of the expert system;
[0109] At this time, P e Satisfies the expert algebraic Riccati equation:
[0110]
[0111] Assume that the control inputs v e and δ e of the expert system are optimal, and the corresponding data are measurable. The v e and δ e data of the expert system formula (21) and the symmetric positive definite weight matrix R e are known, and the corresponding Q e 、P e and K e are unknown. Use the v e and δ e observation data of the expert system to find the weights of the algebraic Riccati equation (18) to make it equivalent to Q e and R e , and design an optimal controller to not only stabilize the spacecraft attitude control system formula (1), but also be able to track the trajectory generated by the expert system.
[0112] Step 3. Based on the observation data of the expert system, approximately obtain the predicted control gain of the expert system through batch least squares method, and obtain the error between the control gain of the learning system and the predicted control gain of the expert system, specifically:
[0113] To ensure that the formula (18) of the learner and the formula (26) of the expert system produce the same control gain, an error is introduced to represent the difference between the control gain K obtained from formula (17) and the K e * obtained from formula (25). Minimize this error to make K gradually approach K e * .
[0114] First, the data v of the expert system trajectory e and δ e are measurable, and the collected data can be used to solve for K in formula (25) e * .
[0115] The expansion of formula (25) is
[0116] [v e (t - (k - 1)τ),..., v e (t - τ), v e (t)] = -K e * [δ e (t - (k - 1)τ),..., δ e (t - τ), δ e (t)](27)
[0117] where v e (t) is the control input data collected at time t, δ e (t) is the state trajectory data collected at time t, τ is the time interval for data acquisition, the integer k is greater than or equal to the matrix dimension n, and represents the number of data of the measured v e and δ e .
[0118] K e can be approximated by the batch least squares method as
[0119]
[0120] where is the control gain of the estimated expert system, and define
[0121] Define the error between K and as:
[0122]
[0123] where is the control gain error between the control gain and the estimated expert system, is the norm of the control gain error between the control gain and the estimated expert system, which is a scalar. When ,
[0124] Step 4: Based on the iterative model-based inverse optimal control algorithm, generate an optimal controller through multiple iterations, and use the controller for spacecraft attitude control to ensure system stability and track the expert system trajectory. Specifically:
[0125] Step 4.1. Set initial conditions: Set R = R e , initial weight Q0 ≥ 0, initial learning rate α0, penalty coefficient λ, i = 0;
[0126] Step 4.2. Calculate the control gain H(δ, v) of the learning system for the current iteration through the Bellman equation formula (16):
[0127] Step 4.3. Based on error minimization, use the gradient descent method to update the weight matrix until the convergence condition is met, specifically:
[0128] To minimize Regarding P as a variable, optimize P using the gradient descent method.
[0129] Define the reward weight correction factor g(P) to repeatedly improve P:
[0130]
[0131] where the learning rate α > 0 controls the speed and update amplitude during the gradient descent process.
[0132] Update α at each iteration:
[0133]
[0134] where i is the current iteration number and 0.001 is the decay factor.
[0135] This iterative process can converge quickly in the early stage and be able to adjust more precise parameters in the later stage.
[0136] Using the inverse optimal control theory to obtain
[0137] Q = -(A T g(P) + g(P)A - f(P)BR -1 B T g(P)) + λ(Q e - Q i ) (32)
[0138] where λ(Q e - Q i ) is the regularization term, and λ is used to control the adjustment speed to make the updated Q closer to the target value. These improvements effectively solve the problem of the algorithm falling into the local optimum.
[0139] Iterate multiple times until the optimal Q * and P * are obtained, such that they satisfy the following conditions:
[0140] Q * + P* A + A T P * -P * BR -1 B T P * = 0 (33)
[0141] K * = R -1 B T P * = K e (34);
[0142] Calculate using formula (28) Further use formula (31) to obtain the dynamic decay learning rate;
[0143] Calculate the optimal strategy:
[0144] Calculate P using formula (18) i :
[0145] Q i +P i A + A T P i -P i BR -1 B T P i = 0 (35)
[0146] Update v using formula (17) i and K i :
[0147] v i = -K i δ = -R -1 B T P i δ (36)
[0148] Error minimization:
[0149] Measure in formula (30) and
[0150]
[0151] Calculate the correction factor g(P):
[0152]
[0153] Update the weight matrix:
[0154] Q i+1 = -A T g(P i) - g(P i ) A + g(P i ) BR -1 B T g(P i ) + λ(Q e - Q i ) (39)
[0155] Let \(i = i + 1\), repeat steps 4.2 to step 4.3, iterate until and \(\|Q\ i+1 - Q i \| \lt \varepsilon q , where \(\varepsilon d =\varepsilon q = 0.0001 is the set convergence condition threshold, obtain the controller for spacecraft attitude control, and use the controller for spacecraft attitude control to ensure system stability and track the expert system trajectory.
[0156] The full - drive system method proposed by the present invention simplifies the controller design and is applicable to high - order nonlinear full - drive systems; at the same time, the inverse - optimal method is used to solve the problem of how to select the weights in the performance index. The algorithm learns the weights of the expert system to ensure the optimality of the system performance index.
[0157] Embodiment:
[0158] In the embodiment of the present invention, as Figures 2 - 5 shown, in order to more intuitively demonstrate the effectiveness of the spacecraft attitude control method based on high - order nonlinear full - drive systems proposed by the present invention, the MATLAB software is used and experimental examples are used to verify the effectiveness of the method proposed by the present invention;
[0159] Set \(I x = 18\ kg\cdot m 2 , I y = 21\ kg\cdot m 2 , I z = 24\ kg\cdot m 2 , the earth's rotation angular velocity \(\omega_0 = 7.292115\times10 -5 rad / s; adopting this method, \(v e and \(\delta e for tracking the expert system trajectory, obtain \(Q e = diag(1.5, 1.5, 1.5, 1, 1, 1)\), \(P e = blockdiag(p1, p2, p3)\), \(K e = blockdiag(k1, k2, k3);
[0160] Where k1 = [1.2247 1.9873], k2 = [1.2247 1.8573], k3 = [1.0000 1.7321].
[0161] In the embodiment of the present invention, based on the proposed full - drive system method and inverse optimal control method, the optimal control problem of the high - order nonlinear full - drive system of the spacecraft attitude is solved, and the experimental results are reflected in the simulation process. From Figure 3 it can be clearly seen that the initial attitude of the spacecraft is unstable. By using the method of the present invention for control, it can be restored to a stable state within 5 seconds, and is very close to Figure 4 the trajectory of the operation of the expert system in [reference], which proves the effectiveness of the proposed control algorithm and realizes the optimal control of the high - order nonlinear full - drive system.
[0162] The method provided by the present invention has been introduced in detail above. Specific examples are used in this article to elaborate on the principle and implementation manner of the present invention. The description of the above embodiments is only used to help understand the core idea of the present invention. It should be noted that for those of ordinary skill in the art of this technology, without departing from the principle of the present invention, several improvements and modifications can be made to the present invention, and these improvements and modifications also fall within the protection scope of the claims of the present invention.
Claims
1. A spacecraft attitude control method based on a high-order nonlinear system, characterized in that, It includes the following steps: Step 1, establish a high-order nonlinear fully actuated system model for spacecraft attitude control; Step 2, through the inverse optimal control method, use the observation data of the expert system to evaluate the unknown reward weights in the performance index, and design an optimal controller based on the reward weights; Step 3, obtain the control gain error between the learning system and the expert system; Step 4, adjust the reward weights through an iterative optimization algorithm to minimize the performance index of the controller and track the trajectory of the expert system.
2. The spacecraft attitude control method based on a high-order nonlinear system according to claim 1, wherein: The high-order nonlinear fully actuated system model for spacecraft attitude control described in Step 1 is: Among them, the state vector is the attitude angle of the spacecraft, φ is the roll angle, θ is the pitch angle, is the yaw angle, u is the control input, is the nonlinear term including centrifugal force and other external disturbances, is the control input matrix of the fully actuated system.
3. The spacecraft attitude control method based on a high-order nonlinear system according to claim 1, characterized in that: The performance index described in Step 2 is: Among them, δ e is the expert state vector, v e is the expert control input, Q e ≥ 0 is the symmetric positive semi - definite weight matrix of the expert system, and R e > 0 is the symmetric positive definite weight matrix of the expert system.
4. The spacecraft attitude control method based on a high-order nonlinear system according to claim 1, characterized in that: The optimal controller described in Step 2 is: wherein, is the optimal expert control input, is the control gain of the optimal controller of the expert system, is the positive definite symmetric matrix of the expert system, and B is the control input matrix of the linear system.
5. The spacecraft attitude control method based on a high-order nonlinear system according to claim 1, characterized in that: The specific content of Step 3 includes: Based on the observation data of the expert system, approximately obtain the estimated control gain of the expert system through batch least squares method, and obtain the error between the control gain of the learning system and the estimated control gain of the expert system.
6. The spacecraft attitude control method based on a high-order nonlinear system according to claim 5, characterized in that: The formula for approximately obtaining the estimated control gain of the expert system through batch least squares method is: Among them, is the control gain of the estimated expert system, δ e is the expert state vector, v e is the expert control input.
7. The spacecraft attitude control method based on a high-order nonlinear system according to claim 5, characterized in that: The formula for obtaining the error between the control gain of the learning system and the estimated control gain of the expert system is: where K is the control gain of the optimal controller for the high-order nonlinear fully actuated system, is the control gain error between the control gain and the predicted expert system, is the control gain of the predicted expert system, v e is the expert control input, is the control gain error norm between the control gain and the estimated expert system, which is a scalar.
8. The spacecraft attitude control method based on a high-order nonlinear system according to claim 1, characterized in that: The specific content of the iterative optimization algorithm described in Step 4 includes: Set the initial weight matrix, learning rate and penalty coefficient; Calculate the control gain of the current iteration through the Bellman equation; Update the weight matrix based on error minimization until the convergence condition is satisfied.
9. The spacecraft attitude control method based on a high-order nonlinear system according to claim 8, characterized in that: The formula for updating the weight matrix is: Q i+1 = -A T g(P i ) - g(P i )A + g(P i )BR -1 B T g(P i ) + λ(Q e -Q i ) (39) where A is the state matrix of the linear system, B is the control input matrix of the linear system, λ(Q e -Q i ) is the regularization term, g(P i ) is the reward weight correction factor, and R is a symmetric positive definite weight matrix.
10. The spacecraft attitude control method based on a high-order nonlinear system according to claim 8, wherein: The convergence condition is: the error norm between the control gain of the learning system and the estimated control gain of the expert system is less than the set threshold.