Control gain adjustment method and spacecraft attitude controller

By establishing a spacecraft attitude nominal dynamic model and strategy optimization algorithm, the control gain is automatically adjusted, and the dependence on precise moment of inertia in traditional methods is solved, and the spacecraft attitude control effect and adjustment efficiency are improved.

CN120276478APending Publication Date: 2025-07-08TSINGHUA UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510432147.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-08
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

Traditional spacecraft attitude control methods rely on the acquisition of precise moment of inertia, resulting in poor control gain adjustment effect and affecting attitude control effect.

Method used

By establishing a target dynamic model, using the representative value or average value of the spindle moment of inertia design range as the nominal value, the spacecraft attitude nominal dynamic model is determined, and the initial linear and nonlinear control gains are determined under the cost function, and the control gains are automatically adjusted in combination with the strategy optimization algorithm.

Benefits of technology

Reduces dependence on precise moment of inertia, improves control gain adjustment efficiency and spacecraft attitude control effect, simplifies the control design process, and reduces dependence on expert experience and labor costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120276478A_ABST
    Figure CN120276478A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a control gain adjustment method and a spacecraft attitude controller. The dependence of control gain adjustment on precise rotational inertia can be reduced. According to the method, a representative value or an average value of a design range of the rotational inertia of a spacecraft spindle is determined as a spindle nominal value, a spacecraft attitude nominal dynamical model is determined accordingly, and then a linear quadratic adjustment gain of a linear part of the spacecraft attitude nominal dynamical model under a cost function is determined as an initial linear control gain; and the zero matrix is determined as an initial nonlinear control gain, so that the initial control gain relatively close to the target control gain is reasonably determined on the premise of not depending on the precise rotational inertia, and gain adjustment is carried out according to the initial control gain, so that the spacecraft attitude controller is adjusted to the target control gain.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of spacecraft control, and particularly to a control gain adjustment method and a spacecraft attitude controller. Background Art

[0002] Spacecraft attitude control, that is, adjusting the spacecraft from an initial attitude to a target attitude, is usually achieved by a model-based control method. Specifically, this control objective is realized by adjusting the control gain of the spacecraft attitude controller.

[0003] However, traditional control gain adjustment methods usually assume that the moment of inertia is accurately known. However, in practical applications, it is difficult to accurately obtain the moment of inertia of the spacecraft, resulting in poor adjustment effect of the control gain, and further affecting the attitude control effect of the spacecraft. Summary of the Invention

[0004] The objective of the embodiments of the present application is to provide a control gain adjustment method and a spacecraft attitude controller to reduce the dependence of control gain adjustment on accurate moment of inertia, thereby improving the control gain adjustment effect of the spacecraft attitude controller.

[0005] In a first aspect, the embodiments of the present application provide a control gain adjustment method, and the method includes:

[0006] Establish a target dynamic model, which is used to characterize the correlation between closed-loop trajectory data, multiple system matrices, a non-linear basis determined based on spacecraft physical information, an attitude control law, and state variables. The state variables are determined according to the error quaternion vector part in the spacecraft kinematics and dynamics model and the spacecraft angular velocity. The attitude control law is determined according to a sign function matrix gain, the state variables, a non-linear part determined based on the non-linear basis, a linear control gain, and a non-linear control gain;

[0007] Determine the representative value or average value within the design range of the spacecraft's principal axis moment of inertia as the nominal value of the principal axis, and based on the nominal value of the principal axis, convert the system matrices corresponding to the non-linear basis and the attitude control law in the target dynamic model into nominal system matrices to obtain the nominal dynamic model of the spacecraft attitude;

[0008] Determine the linear quadratic regulation gain of the linear part of the nominal dynamic model of the spacecraft attitude under a cost function as the initial linear control gain, and determine a zero matrix as the initial non-linear control gain. The cost function is used to characterize the correlation between the cost function value, the state weight matrix, the control input weight matrix, the state variables, and the attitude control law;

[0009] Take the initial linear control gain and the initial non - linear control gain as the initial control gains of the spacecraft attitude controller, and adjust the spacecraft attitude controller from the initial control gains to the target control gains.

[0010] In the second aspect of the embodiments of the present application, a control gain adjustment device is provided. The device includes:

[0011] A model - building module, configured to build a target dynamic model, which is used to characterize the correlation relationship between closed - loop trajectory data, multiple system matrices, a non - linear basis determined based on spacecraft physical information, an attitude control law, and state variables. The state variables are determined according to the error quaternion vector part and the spacecraft angular velocity in the spacecraft kinematics and dynamics models, and the attitude control law is determined according to a sign - function matrix gain, the state variables, a non - linear part determined based on the non - linear basis, a linear control gain, and a non - linear control gain;

[0012] A first processing module, configured to determine a representative value or an average value within the design range of the spacecraft principal - axis moment of inertia as the principal - axis nominal value, and convert the system matrices corresponding to the non - linear basis and the attitude control law in the target dynamic model into nominal system matrices based on the principal - axis nominal value, so as to obtain a spacecraft attitude nominal dynamic model;

[0013] A second processing module, configured to determine the linear - quadratic regulation gain of the linear part of the spacecraft attitude nominal dynamic model under a cost function as the initial linear control gain, and determine a zero matrix as the initial non - linear control gain. The cost function is used to characterize the correlation relationship between a cost - function value, a state - weight matrix, a control - input weight matrix, state variables, and an attitude control law;

[0014] A gain - adjustment module, configured to take the initial linear control gain and the initial non - linear control gain as the initial control gains of the spacecraft attitude controller, and adjust the spacecraft attitude controller from the initial control gains to the target control gains.

[0015] In the third aspect of the embodiments of the present application, a computer program product is provided, including computer programs / instructions, and when the computer programs / instructions are executed by a processor, the steps of the control - gain adjustment method described in the first aspect are implemented.

[0016] In the fourth aspect of the embodiments of the present application, a computer - readable storage medium is provided, on which computer programs / instructions are stored, and when the computer programs / instructions are executed by a processor, the steps of the control - gain adjustment method described in the first aspect are implemented.

[0017] In a fifth aspect of the embodiments of the present application, a spacecraft attitude controller is provided, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the steps of the control gain adjustment method described in the first aspect are implemented.

[0018] It can be seen from the above technical solution that the present application determines the representative value or average value within the designed range of the spacecraft's principal axis moment of inertia as the nominal value of the principal axis, and accordingly determines the nominal dynamic model of the spacecraft attitude. Furthermore, the linear quadratic regulation gain of the linear part of the nominal dynamic model of the spacecraft attitude under the cost function is determined as the initial linear control gain, and the zero matrix is determined as the initial non-linear control gain. Thus, without relying on the precise moment of inertia, a relatively close initial control gain to the target control gain is reasonably determined. Thereby, while ensuring the control gain adjustment efficiency, the dependence on the precise moment of inertia for control gain adjustment is reduced, thereby improving the control gain adjustment effect of the spacecraft attitude controller, and further improving the attitude control effect of the spacecraft. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] To more clearly illustrate the embodiments of the present application, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0020] Figure 1 It is a flowchart of the implementation of a control gain adjustment method provided by an embodiment of the present application;

[0021] Figure 2 It is a flowchart of a gain automatic adjustment method for a spacecraft data-driven attitude controller based on a policy optimization algorithm provided by an embodiment of the present application;

[0022] Figure 3 It is a trajectory diagram of the change of the spacecraft error quaternion provided by an embodiment of the present application;

[0023] Figure 4 It is a trajectory diagram of the change of the actual angular velocity of the spacecraft provided by an embodiment of the present application;

[0024] Figure 5 It is a trajectory diagram of the change of the actual control torque of the spacecraft provided by an embodiment of the present application;

[0025] Figure 6 It is a schematic structural diagram of a control gain adjustment device provided by an embodiment of the present application;

[0026] Figure 7 It is a schematic diagram of a spacecraft attitude controller provided by an embodiment of the present application. Detailed implementation manners

[0027] The following will clearly and completely describe the technical solutions in the embodiments of the present application with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the protection scope of the present application.

[0028] Spacecraft attitude control is crucial for many space missions and can serve scientific observation tasks such as aligning antennas, cameras, and scientific instruments. Currently, model-based control methods such as proportional-integral-differential (PID) control, linear quadratic regulator, and model predictive control based on model regulation are often used to achieve spacecraft attitude control. Specifically, this control goal is achieved by adjusting the control gain of the spacecraft attitude controller.

[0029] These methods usually assume that the spacecraft dynamics model is accurately known, especially assuming that the moment of inertia is accurately known. However, in practical applications, it is difficult to accurately obtain the moment of inertia of the spacecraft. Taking ground test methods such as the swing method (which is currently the main method for obtaining the moment of inertia) as an example, although such methods can obtain the moment of inertia relatively accurately, they are costly, time-consuming, and may damage the spacecraft, which makes it difficult to accurately obtain the moment of inertia of the spacecraft by means of such methods in practical applications. As a result, the adjustment effect of the control gain in practical applications is not good, which in turn affects the spacecraft attitude control effect.

[0030] Based on the above analysis, in view of the problem that the control gain adjustment in the related technology highly depends on the accurate moment of inertia, the embodiments of the present application provide a control gain adjustment method and a spacecraft attitude controller, which can reduce the dependence of the control gain adjustment on the accurate moment of inertia, thereby improving the control gain adjustment effect of the spacecraft attitude controller, and further improving the spacecraft attitude control effect.

[0031] First, to facilitate the understanding of the technical solutions provided by the present application, the main technical concepts involved in the embodiments of the present application will be briefly described below.

[0032] The control gain of the spacecraft attitude controller: refers to the magnitude of the control action of the spacecraft attitude controller on the system output; taking the spacecraft attitude controller designed based on the PID controller as an example, its control gain can be the proportional gain Kp in the PID controller.

[0033] The representative value: refers to the lower confidence limit of the arithmetic mean of all observed values within a set range.

[0034] See Figure 1 shown in the following is the implementation flowchart of a control gain adjustment method provided by an embodiment of the present application. The method may include the following steps:

[0035] Step S101: Establish a target dynamics model, where the target dynamics model is used to characterize the correlation relationship between closed-loop trajectory data, multiple system matrices, a non-linear basis determined based on spacecraft physical information, an attitude control law, and state variables. The state variables are determined according to the error quaternion vector part in the spacecraft kinematics and dynamics model and the spacecraft angular velocity. The attitude control law is determined according to a sign function matrix gain, the state variables, a non-linear part determined based on the non-linear basis, a linear control gain, and a non-linear control gain.

[0036] In specific implementation, analyze the pre-established spacecraft kinematics and dynamics model described by quaternion, design state variables and an attitude control law, and determine the correlation relationship between closed-loop trajectory data, multiple system matrices, a non-linear basis determined based on spacecraft physical information, an attitude control law, and state variables according to the spacecraft kinematics and dynamics model, the state variables, and the attitude control law, so as to establish a target dynamics model. Among them, the multiple system matrices include: system matrices corresponding to the attitude control law, closed-loop trajectory data, and non-linear basis respectively.

[0037] Step S102: Determine the representative value or average value within the design range of the spacecraft's principal axis moment of inertia as the nominal value of the principal axis, and based on the nominal value of the principal axis, convert the system matrices corresponding to the non-linear basis and the attitude control law in the target dynamics model into nominal system matrices to obtain the nominal dynamics model of the spacecraft attitude.

[0038] In specific implementation, considering that it is not easy to obtain accurate moment of inertia, but the design range of the principal axis moment of inertia is usually provided by the overall design department. Therefore, the present application selects the representative value or average value within this range as the nominal value of the principal axis, and based on this nominal value of the principal axis, converts the parts related to the moment of inertia in the target dynamics model (i.e., the system matrices corresponding to the non-linear basis and the attitude control law respectively) into parts independent of the moment of inertia (i.e., nominal system matrices) to obtain the nominal dynamics model of the spacecraft attitude. Subsequently, using this nominal dynamics model of the spacecraft attitude to determine the initial linear control gain can reduce the dependence on accurate moment of inertia in the process of determining the initial control gain.

[0039] Exemplarily, the design range of the spacecraft's principal axis moment of inertia can be expressed as follows:

[0040]

[0041] where J iRepresents the principal moment of inertia about the i-axis among the three axes of the spacecraft, J , respectively represent the design lower bound and design upper bound of the principal moment of inertia; taking J = 30 kg·m 2 , as an example, the nominal value of the principal axis J0 = 60 kg·m 2 can be set.

[0042] Step S103: Determine the linear quadratic regulation gain of the linear part of the nominal attitude dynamics model of the spacecraft under the cost function as the initial linear control gain, and determine the zero matrix as the initial non-linear control gain. The cost function is used to characterize the correlation relationship between the cost function value, the state weight matrix, the control input weight matrix, the state variables, and the attitude control law.

[0043] In specific implementation, the linear part of the nominal attitude dynamics model of the spacecraft (which is independent of the moment of inertia, for example, it can include the system matrix corresponding to the state variables and the nominal system matrix corresponding to the attitude control law) is determined as the initial linear control gain under the cost function, and the zero matrix is determined as the initial non-linear control gain. In this way, the determined initial control gains (i.e., the initial linear control gain and non-linear control gain) can be relatively close to the target control gain (i.e., the optimal control gain, used to adjust the spacecraft to the target attitude).

[0044] Step S104: Use the initial linear control gain and the initial non-linear control gain as the initial control gains of the spacecraft attitude controller, and adjust the spacecraft attitude controller from the initial control gains to the target control gains.

[0045] In specific implementation, through repeated adjustment and trial-and-error based on expert experience, the spacecraft attitude controller is adjusted from the initial control gains to the target control gains.

[0046] As a possible implementation method, introduce a policy optimization algorithm into the adjustment process of the control gains. Specifically, use the initial control gains as the warm start of the policy optimization algorithm, and perform iterative optimization of the control gains based on policy gradient estimation until the iteration end condition (such as reaching the set number of iteration steps) is reached. Then, it can be realized that the spacecraft attitude controller is automatically adjusted from the initial control gains to the target control gains. This can eliminate the manual repeated adjustment and trial-and-error link of the control gains in the traditional model-based control method, thereby improving the adjustment efficiency and reducing the dependence on expert experience and labor costs for the control gain adjustment.

[0047] Optionally, the spacecraft attitude controller is designed based on a combined controller similar to the classical Proportion-Differentiation (PD) control and feedforward compensation.

[0048] It can be seen from the above technical solution that in this application, the representative value or average value within the designed range of the spacecraft's principal axis moment of inertia is determined as the nominal value of the principal axis, and based on this, the nominal dynamic model of the spacecraft attitude is determined. Furthermore, the linear quadratic regulation gain of the linear part of the nominal dynamic model of the spacecraft attitude under the cost function is determined as the initial linear control gain, and the zero matrix is determined as the initial non-linear control gain. Thus, without relying on the exact moment of inertia, a relatively close initial control gain to the target control gain is reasonably determined. Thereby, while ensuring the control gain adjustment efficiency, the dependence on the exact moment of inertia for control gain adjustment is reduced, thus improving the control gain adjustment effect of the spacecraft attitude controller, and further improving the attitude control effect of the spacecraft.

[0049] As a possible implementation manner, the target dynamic model is established through the following steps:

[0050] Use predefined state variables Decompose the spacecraft kinematic and dynamic models into linear and non-linear parts, where q ev represents the error quaternion vector part in the spacecraft kinematic and dynamic models, ω represents the spacecraft angular velocity, and the superscript represents vector transpose;

[0051] Discretize each part obtained from the decomposition through the sampling time h to establish the target dynamic model, and the target dynamic model is expressed as follows:

[0052] x t+1 = A 1t x t + C t φ1(x t ) + B t u t , q e0 (0) ≥ 0

[0053] x t+1 = A 2t x t + C t φ2(x t ) + B t u t , q e0 (0) < 0

[0054] Among them, in practical applications, h can be set to 0.1 s, x t+1 represents the closed-loop trajectory data, φ1(xt ), φ2(x t ) both represent non - linear bases determined based on spacecraft physical information, u t represents the attitude control law, q e0 represents the scalar part of the error quaternion in the spacecraft kinematics and dynamics model, A 1t , A 2t , B t , C t respectively represent the system matrices, and their definitions can be expressed as follows:

[0055]

[0056] Among them, the constant matrices C J1 and C J2 are respectively related to the moment of inertia of the spacecraft and are unknown. R m×n represents the set of m×n real matrices. The matrices I3 and 03 respectively represent the 3 - row and 3 - column identity matrix and zero matrix, that is

[0057] Optionally, the nominal dynamic model of the spacecraft attitude is expressed as follows:

[0058]

[0059] Among them, x t+1 represents the closed - loop trajectory data, x t represents the state variable, φ1(x t ) represents the non - linear basis, u t represents the attitude control law, A it represents the system matrix, represents the nominal system matrix, and h represents the sampling time. The matrices I3 and 03 respectively represent the 3 - row and 3 - column identity matrix and zero matrix, and J0 represents the nominal value of the principal axis.

[0060] Optionally, the cost function is expressed as follows:

[0061]

[0062] Among them, C(K s ) represents the cost function value, x t represents the state variable at time step t, Q and R respectively represent the state weight matrix and the control input weight matrix, u t represents the attitude control law at time step t. The superscript represents vector transpose; in practical applications, the values of the state weight matrix and the control input weight matrix can be set and adjusted according to the requirements of the spacecraft attitude control task. For example, Q can be set to 10 -2I6, R = I3, where matrix I3 and I6 represent the 3x3 and 6x6 identity matrices respectively.

[0063] Optionally, determining the linear quadratic regulation gain of the linear part of the nominal spacecraft attitude dynamics model under the cost function as the initial linear control gain includes:

[0064] Taking the linear part of the nominal spacecraft attitude dynamics model and determining the linear quadratic regulation gain under the cost function as the initial linear control gain

[0065] In specific implementation, taking the initial linear control gain and determining it as the linear part of the above-mentioned nominal spacecraft attitude dynamics model under the above cost function. For example, based on the expressions of the above-mentioned nominal spacecraft attitude dynamics model and the above cost function, the linear quadratic regulation gain can be determined as [0.1I3, 2.45I3], that is, the initial linear control gain can be determined and the initial non-linear control gain can be correspondingly determined Thus, for the initial control gain it can be specifically expressed as K ini = [0.1I3, 2.45I3, 0 3×12 .

[0066] Optionally, to describe the reorientation of the spacecraft attitude, a spacecraft dynamics and kinematics model is established, and the spacecraft kinematics and dynamics model can be expressed as follows:

[0067]

[0068] where u ∈ R 3 is the control torque, is the spacecraft angular velocity, is the derivative of the spacecraft angular velocity, and the superscript represents vector transpose, J s ∈ R 3 is the spacecraft moment of inertia, is the error quaternion describing the spacecraft attitude, q e0 is the scalar part of the error quaternion, is the vector part of the error quaternion, and it satisfies the constraint and are the derivatives of the scalar and vector of the error quaternion; the error quaternion can be specifically defined using the quaternion and the desired quaternion as In practical applications, the initial value of the quaternion can be set and the desired quaternion can be set and ω × are respectively the skew-symmetric matrices of q ev and ω, and their specific forms are as follows:

[0069]

[0070] The moment of inertia J of the spacecraft s can be specifically defined as:

[0071]

[0072] where J x , J y , J z are the principal axis moments of inertia, and J xy , J xz , J yz are the products of inertia, which are used to represent the inertia coupling of the three axes of the spacecraft, and satisfy:

[0073] (1) J x + J y ≥ J z , J x + J z ≥ J y , J z + J y ≥ J x ;

[0074] (2) J x >> J xy , J x >> J xz , J y >> J yx , J y >> J yz , J z >> J zx , J y >> J zy

[0075] where >> means much greater than.

[0076] Optionally, based on the above spacecraft kinematics and dynamics models, the specific form of the nonlinear basis is as follows:

[0077]

[0078] Optionally, based on the above nonlinear basis, the nonlinear part φ(ω t ) in the attitude control law is expressed as follows:

[0079]

[0080] As a possible implementation manner, the attitude control law is determined by the following formula:

[0081] u t =-diag[sgn(q e0 (0))I3, I3]K x x t -K J φ(ω t )

[0082] where u t represents the attitude control law; diag[sgn(q e0 (0))I3, I3] represents the sign function matrix gain, which can be used to avoid the attitude unwinding phenomenon of the spacecraft; K x ∈R 3×6 and K J ∈R 3×12 respectively represent the linear control gain and the nonlinear control gain to be determined, which are respectively used to stabilize the linear part of the target dynamic model and cancel the nonlinear part of the target dynamic model; x t represents the state variable, and φ(ω t ) represents the nonlinear part.

[0083] As a possible implementation manner, the adjustment of the spacecraft attitude controller from the initial control gain to the target control gain includes:

[0084] Automatically adjusting the spacecraft attitude controller from the initial control gain to the target control gain based on a policy optimization algorithm;

[0085] where the iterative form of the control gain in the policy optimization algorithm is expressed as follows:

[0086]

[0087] where respectively represent the (n + 1)-th and n-th iterative values of the control gain, η represents the iteration step size, represents the gradient estimation.

[0088] In this implementation manner, the present application uses a policy optimization algorithm to achieve the automatic adjustment from the initial control gain to the optimal control gain , and designs the iterative form of the control gain in the policy optimization algorithm to ensure the adjustment effect of the policy optimization algorithm on the control gain.

[0089] Thus, based on the target gain automatically adjusted by the policy optimization algorithm, the adaptability of the spacecraft attitude control system can be improved, enabling it to achieve high-precision, fast convergence, and low-energy consumption attitude control even when the moment of inertia is unknown, external disturbances exist, and high maneuverability requirements for the attitude are present.

[0090] Optionally, the policy optimization algorithm automatically adjusts the spacecraft attitude controller from the initial control gain to the target control gain, including:

[0091] When the set conditions are met, the policy optimization algorithm automatically adjusts the spacecraft attitude controller from the initial control gain to the target control gain, and the set conditions are expressed as follows:

[0092]

[0093] where ΔJ represents the maximum deviation between the true value and the nominal value of the principal axis moment of inertia, h represents the sampling time, J0 represents the principal axis nominal value, C B ,C C represents a constant, and:

[0094]

[0095] where the constant ρ1 satisfies ρ1 ∈ [(ρ ini +1) / 2,1), c ini ,ρ ini respectively represent the linear convergence coefficient and the exponential convergence coefficient of the linear part of the nominal dynamic model of the spacecraft attitude at the initial linear control gain under the condition that the linear part is the above in this case, it satisfies The constant c2 satisfies The constant

[0096] In this embodiment, the present application explicitly defines the allowable boundary of the spacecraft moment of inertia (i.e., the set conditions) to reasonably determine the moment of inertia boundary and accordingly restrict the use of the policy optimization algorithm. Thereby, the optimization performance and convergence performance of the policy optimization algorithm can be guaranteed, so as to provide theoretical and performance guarantees for the embodiment of automatically converging the initial control gain to the optimal control gain In this way, the reliability and stability of the spacecraft attitude control system can be effectively improved. At the same time, the need for accurately identifying the spacecraft moment of inertia in traditional model-based control methods can be avoided, and the empirical dependence and labor costs of manual repeated trial-and-error adjustment of control gains can be reduced.

[0097] Optionally, the gradient estimation in the policy optimization algorithm is determined through the following steps:

[0098] For each trajectory, determine the control gain through the first formula and sample the initial state variable x0;

[0099] For each trajectory, when the number of time steps t has not reached the set iteration length T, iteratively execute the following steps:

[0100] Determine the attitude control law through the second formula;

[0101] Collect the closed-loop trajectory data x t+1 and the instantaneous cost function value c t ;

[0102] For each trajectory, when the number of time steps t reaches the set iteration length, determine the cost function estimate through the third formula

[0103] Determine the gradient estimate through the fourth formula;

[0104] The first formula is expressed as follows:

[0105]

[0106] The second formula is expressed as follows:

[0107] u t =-K x x t -K J φ(ω t )

[0108] The third formula is expressed as follows:

[0109]

[0110] The fourth formula is expressed as follows:

[0111]

[0112] where represents the (i - 1)-th iteration value of the control gain K s , U (i) represents a matrix randomly drawn from the set of set matrices, K x and K J represent the linear control gain and the non-linear control gain respectively, x t represents the state variable, φ(ω t ) represents the non-linear part, represents the gradient estimate, N represents the number of trajectories, and p and n represent the control gain K sThe row dimension and column dimension, where r represents the smoothing parameter.

[0113] Optionally, the closed-loop trajectory data x t+1 is determined by the following formula:

[0114] x t+1 = A 1t x t + C t φ(x t ) + B t u t

[0115] where x t represents the state variable, φ(x t ) represents the non-linear basis in the target dynamics model, u t represents the attitude control law, and A 1t , B t , C t represent different system matrices in the target dynamics model;

[0116] The instantaneous cost function value c t is determined by the following formula:

[0117]

[0118] where x t represents the state variable, u t represents the attitude control law, Q and R respectively represent the state weight matrix and the control input weight matrix, and the superscript represents vector transpose.

[0119] Exemplarily, referring to Figure 2 the flowchart of the gain automatic adjustment method of the spacecraft data-driven attitude controller based on the policy optimization algorithm shown, the gain automatic adjustment method provided by this application mainly includes the following steps:

[0120] S1: Design a data-driven anti-windup attitude control law u t = -diag[sgn(q e0 (0))I3, I3]K x x t - K J φ(ω t ).

[0121] In this step, the linear and non-linear parts of the spacecraft kinematics and dynamics models described by quaternions are decomposed, and a data-driven anti-windup attitude control law is designed. The linear control gain and non-linear control gain in the anti-windup attitude control law are to be determined in a data-driven form using the policy optimization algorithm in subsequent steps.

[0122] S2: Select the initial control gain of the strategy optimization algorithm

[0123] In this step, for the control gains of the attitude control law in step S1 (i.e., the linear control gain and the non - linear control gain), roughly select the control gains that do not depend on the accurate spacecraft moment of inertia as the warm start of the strategy optimization algorithm.

[0124] S3: Quantitatively analyze the moment of inertia boundary to ensure the performance of the strategy optimization algorithm

[0125] In this step, explicitly define the allowable spacecraft moment of inertia boundary to ensure the optimization performance and convergence performance of the strategy optimization algorithm.

[0126] S4: Optimize the control gain based on the strategy optimization algorithm

[0127] In this step, use the strategy optimization algorithm to automatically adjust the initial control gain to the optimal control gain. Among them, based on the above first formula to the fourth formula, the gradient estimation in the strategy optimization algorithm can be determined by Algorithm 1 described below.

[0128] Algorithm 1 Policy Gradient Estimation

[0129] Input: Initial control gain Number of trajectories N, smoothing parameter r, iteration length T;

[0130] 1: For i = 1, 2,..., N, execute:[[]]

[0131] 2: Set the control gain where U (i) is uniformly randomly sampled from the set of matrices with F - norm of r (i.e., the set of specified matrices);

[0132] 3: Sample the initial state x0 ∼ D;

[0133] 4: For t = 1, 2,..., T, execute:[[]]

[0134] 5: Set the attitude control law u t =-K x x t -K J φ(ω t );

[0135] 6: Collect the closed - loop trajectory data x t+1 and the instantaneous cost function value c t ;

[0136] 7: End;

[0137] 8: Determine the cost function estimate

[0138] 9: End;

[0139] 10: Return the gradient estimate

[0140] In the above Algorithm 1, N independent trajectories are sampled offline to estimate the policy gradient. Each trajectory provides data on the system's response under the perturbed policy. Among them, the value of N plays an important role in balancing the accuracy and efficiency of the policy gradient estimation. The iteration length T represents the total number of time steps in each trajectory, which can be used to calculate the data volume of the trajectory cost and can significantly affect the accuracy of the cost estimation; the smoothing parameter r is used to specify the perturbation size applied to the sampling policy, and the choice of r represents a trade-off between exploration and stability in the policy search; the initial state distribution D represents the set of possible initial states of the system; the closed-loop trajectory data x t+1 is obtained by calculating through the closed-loop system; the immediate cost function value c t is calculated from the collected state data and input data (i.e., the state weight matrix and the control input weight matrix). In practical applications, η = 10 -5 , N = 10, T = 2000, r = 2.45×10 -4 .

[0141] The spacecraft attitude error quaternion, angular velocity, and control torque change trajectory diagrams obtained by using the gain adjustment method provided in this application are as follows Figures 3 to 5 shown.

[0142] According to Figure 3 the spacecraft error quaternion change trajectory diagram shown, it can be seen that the spacecraft error quaternion successfully converges to a relatively close equilibrium point, and anti-unwinding is successfully achieved. Among them, this rotation behavior minimizes the control consumption by ensuring that the system follows the shortest path in the quaternion space; according to Figure 4 the spacecraft actual angular velocity change trajectory diagram shown, it can be seen that the spacecraft angular velocity successfully converges to the origin, and spacecraft attitude control is successfully achieved; Figure 5 the spacecraft actual control torque change trajectory diagram shown intuitively displays the control torque change curve.

[0143] Based on the above examples and embodiments, considering that the traditional method for accurately identifying the moment of inertia of a spacecraft is the ground test swing method, which is extremely time-consuming and laborious. Therefore, this application focuses on the spacecraft attitude reorientation control task when the moment of inertia is unknown, and proposes a control gain adjustment method, which can reduce (or avoid) the dependence of traditional model-based control methods on accurately identifying the moment of inertia of the spacecraft. On this basis, a cost function suitable for automatic adjustment of the control gain, a specific iterative form of the control gain, and a policy gradient estimation algorithm are designed to realize the automatic selection of the control gain based on data by means of a policy optimization algorithm, so as to directly use the data to achieve spacecraft attitude control. In this way, the need for manual trial-and-error adjustment of the control gain in traditional model-based control methods can be avoided, enabling the control gain adjustment method provided in this application to better solve the attitude control problem of the spacecraft with unknown moment of inertia and reduce the dependence on expert experience in the control gain adjustment process, thus simplifying the control design process.

[0144] It should be noted that for the method embodiments, for simplicity of description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the embodiments of this application are not limited by the described action sequence, because according to the embodiments of this application, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions involved are not necessarily essential for the embodiments of this application.

[0145] The embodiments of this application also provide a control gain adjustment device, as Figure 6 shown. The device includes:

[0146] A model establishment module, configured to establish a target dynamics model, which is used to characterize the correlation relationship between closed-loop trajectory data, multiple system matrices, a nonlinear basis determined based on spacecraft physical information, an attitude control law, and state variables. The state variables are determined according to the error quaternion vector part in the spacecraft kinematics and dynamics models and the spacecraft angular velocity. The attitude control law is determined according to a sign function matrix gain, the state variables, a nonlinear part determined based on the nonlinear basis, a linear control gain, and a nonlinear control gain;

[0147] A first processing module, configured to determine a representative value or an average value within the design range of the spacecraft's principal axis moment of inertia as the nominal value of the principal axis, and based on the nominal value of the principal axis, convert the system matrices corresponding to the nonlinear basis and the attitude control law in the target dynamics model into nominal system matrices to obtain a nominal dynamics model of the spacecraft attitude;

[0148] A second processing module, configured to determine a linear quadratic regulation gain of the linear part of the nominal spacecraft attitude dynamics model under a cost function as an initial linear control gain, and determine a zero matrix as an initial non-linear control gain, where the cost function is used to characterize the correlation relationship between the cost function value, the state weight matrix, the control input weight matrix, the state variable, and the attitude control law;

[0149] A gain adjustment module, configured to use the initial linear control gain and the initial non-linear control gain as the initial control gains of the spacecraft attitude controller, and adjust the spacecraft attitude controller from the initial control gains to the target control gains.

[0150] Optionally, the model establishment module is further configured to perform the following steps:

[0151] Use predefined state variables to decompose the spacecraft kinematics and dynamics model into a linear part and a non-linear part, where q ev represents the error quaternion vector part in the spacecraft kinematics and dynamics model, ω represents the spacecraft angular velocity, and the superscript represents vector transpose;

[0152] Discretize each part obtained by the decomposition through a sampling time h to establish a target dynamics model, where the target dynamics model is expressed as follows:

[0153] x t+1 = A 1t x t + C t φ1(x t ) + B t u t , q e0 (0) ≥ 0

[0154] x t+1 = A 2t x t + C t φ2(x t ) + B t u t , q e0 (0) < 0

[0155] where x t+1 represents the closed-loop trajectory data, φ1(x t ), φ2(x t ) both represent non-linear bases, u t represents the attitude control law, q e0 represents the error quaternion scalar part in the spacecraft kinematics and dynamics model, A 1t , A 2t , Bt , C t represent the system matrices respectively, and:

[0156]

[0157] where the constant matrix C J1 and C J2 are respectively related to the moment of inertia of the spacecraft and are unknown. The matrices I3 and 03 represent the 3×3 identity matrix and the zero matrix respectively.

[0158] Optionally, the nominal attitude dynamics model of the spacecraft is expressed as follows:

[0159]

[0160] where x t+1 represents the closed-loop trajectory data, x t represents the state variable, φ1(x t ) represents the non-linear basis, u t represents the attitude control law, A it represents the system matrix, represents the nominal system matrix, and h represents the sampling time. The matrices I3 and 03 represent the 3×3 identity matrix and the zero matrix respectively, and J0 represents the nominal value of the principal axis;

[0161] The cost function is expressed as follows:

[0162]

[0163] where C(K s ) represents the cost function value, x t represents the state variable at time step t, Q and R represent the state weight matrix and the control input weight matrix respectively, u t represents the attitude control law at time step t, and the superscript represents vector transpose;

[0164] The second processing module is further configured to determine the linear quadratic regulation gain of the linear part of the nominal attitude dynamics model of the spacecraft under the cost function as the initial linear control gain

[0165] Optionally, the kinematics and dynamics model of the spacecraft is expressed as follows:

[0166]

[0167] Based on the kinematics and dynamics model of the spacecraft, the non-linear basis is expressed as follows:

[0168]

[0169] Based on the non - linear basis, the non - linear part φ(ω t ) is expressed as follows:

[0170]

[0171] where, represents the derivative of the scalar part q e0 of the error quaternion, represents the derivative of the vector part of the error quaternion, represents the derivative of the spacecraft angular velocity , and ω × respectively represent the skew - symmetric matrices of q ev and ω, and J s represents the spacecraft moment of inertia, u represents the control torque, and the superscript represents vector transpose.

[0172] Optionally, the attitude control law is determined by the following formula:

[0173] u t =-diag[sgn(q e0 (0))I3, I3]K x x t -K J φ(ω t )

[0174] where, u t represents the attitude control law, diag[sgn(q e0 (0))I3, I3] represents the sign - function matrix gain, K x and K J respectively represent the linear control gain and the non - linear control gain, x t represents the state variable, and φ(ω t ) represents the non - linear part.

[0175] Optionally, the gain adjustment module is further configured to automatically adjust the spacecraft attitude controller from the initial control gain to the target control gain based on a policy optimization algorithm;

[0176] where, the iterative form of the control gain in the policy optimization algorithm is expressed as follows:

[0177]

[0178] where, respectively represent the (n + 1)-th and n-th iterative values of the control gain, η represents the iteration step size, represents the gradient estimation.

[0179] Optionally, the gain adjustment module is further configured to, when a set condition is satisfied, automatically adjust the spacecraft attitude controller from the initial control gain to the target control gain based on a policy optimization algorithm, and the set condition is expressed as follows:

[0180]

[0181] where, ΔJ represents the maximum deviation between the true value and the nominal value of the principal axis inertia, h represents the sampling time, J0 represents the nominal value of the principal axis, C B ,C C represents a constant, and:

[0182]

[0183] where, the constant ρ1 satisfies ρ1 ∈ [(ρ ini +1)*2, 1), c ini ,ρ ini respectively represent the linear convergence coefficient and the exponential convergence coefficient of the linear part of the nominal dynamic model of the spacecraft attitude under the initial linear control gain and the constant c2 satisfies constant

[0184] Optionally, the gain adjustment module is further configured to perform the following steps:

[0185] For each trajectory, determine the control gain through the first formula and sample the initial state variable x0;

[0186] For each trajectory, when the number of time steps t does not reach the set iteration length T, iteratively execute the following steps:

[0187] Determine the attitude control law through the second formula;

[0188] Collect the closed-loop trajectory data x t+1 and the instantaneous cost function value c t ;

[0189] For each trajectory, when the number of time steps t reaches the set iteration length, determine the cost function estimate through the third formula

[0190] Determine the gradient estimation through the fourth formula;

[0191] The first formula is expressed as follows:

[0192]

[0193] The second formula is expressed as follows:

[0194] u t =-K x x t -K J φ(ω t )

[0195] The third formula is expressed as follows:

[0196]

[0197] The fourth formula is expressed as follows:

[0198]

[0199] Wherein, represents the (i - 1)-th iteration value of the control gain K, U s represents a matrix randomly selected from a set of preset matrices, K (i) and K x respectively represent a linear control gain and a non-linear control gain, x J represents a state variable, φ(ω t ) represents the non-linear part, t represents a gradient estimate, N represents the number of trajectories, p and n respectively represent the row dimension and the column dimension of the control gain K , and r represents a smoothing parameter. s The row dimension and the column dimension, r represents a smoothing parameter.

[0200] Optionally, the closed-loop trajectory data x t+1 is determined by the following formula:

[0201] x t+1 =A 1t x t +C t φ(x t )+B t u t

[0202] Wherein, x t represents a state variable, φ(x t ) represents the non-linear basis in the target dynamic model, u t represents an attitude control law, A 1t , B t , C t represent different system matrices in the target dynamic model;

[0203] The instantaneous cost function value c tDetermined by the following formula:

[0204]

[0205] where x t represents the state variable, u t represents the attitude control law, Q and R respectively represent the state weight matrix and the control input weight matrix, and the superscript represents the vector transpose.

[0206] It can be seen from the above technical solutions that in this application, the representative value or average value within the design range of the spacecraft's principal axis moment of inertia is determined as the nominal value of the principal axis, and based on this, the nominal attitude dynamics model of the spacecraft is determined. Furthermore, the linear quadratic regulation gain of the linear part of the spacecraft's nominal attitude dynamics model under the cost function is determined as the initial linear control gain, and the zero matrix is determined as the initial non-linear control gain. Thus, without relying on the precise moment of inertia, a relatively close initial control gain to the target control gain is reasonably determined. Thereby, while ensuring the control gain adjustment efficiency, the dependence on the precise moment of inertia for control gain adjustment is reduced, thereby improving the control gain adjustment effect of the spacecraft attitude controller, and further improving the attitude control effect of the spacecraft.

[0207] The embodiment of this application also provides a spacecraft attitude controller. Referring to Figure 7 , Figure 7 is a schematic diagram of the spacecraft attitude controller proposed in the embodiment of this application. As shown in Figure 7 , the spacecraft attitude controller 100 includes: a memory 110 and a processor 120. The memory 110 is communicatively connected to the processor 120 via a bus. A computer program is stored in the memory 110, and this computer program can run on the processor 120, thereby implementing the steps in the control gain adjustment method disclosed in the embodiment of this application.

[0208] The embodiment of this application also provides a computer-readable storage medium, on which a computer program / instructions are stored. When the computer program / instructions are executed by a processor, the control gain adjustment method disclosed in the embodiment of this application is implemented.

[0209] The embodiment of this application also provides a computer program product, including a computer program / instructions. When the computer program / instructions are executed by a processor, the control gain adjustment method disclosed in the embodiment of this application is implemented.

[0210] Each embodiment in this specification is described in a progressive manner. The key point of each embodiment is to illustrate the differences from other embodiments. The same or similar parts among the embodiments can be referred to each other.

[0211] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, apparatus, or computer program product. Therefore, the embodiments of the present application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Moreover, the embodiments of the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0212] The embodiments of the present application are described with reference to the flowcharts and / or block diagrams of methods, systems, devices, storage media, and program products according to the embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram, as well as the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing terminal devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing terminal devices generate a device for implementing the specified functions in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0213] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing terminal device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device, and the instruction device implements the specified functions in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0214] These computer program instructions can also be loaded onto a computer or other programmable data processing terminal device, so that a series of operation steps are executed on the computer or other programmable terminal device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable terminal device provide steps for implementing the specified functions in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0215] Finally, it should also be noted that in this text, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or terminal device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or terminal device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or terminal device comprising the said element.

[0216] The above has introduced in detail a control gain adjustment method and a spacecraft attitude controller provided by the present application. Specific examples are used in this text to elaborate on the principle and implementation manner of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application; at the same time, for those of ordinary skill in the art, according to the idea of the present application, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present application.

Claims

1. A control gain adjustment method, characterized in that, The method includes: Establishing a target dynamic model, which is used to characterize the correlation relationship between closed-loop trajectory data, multiple system matrices, a non-linear basis determined based on spacecraft physical information, an attitude control law, and state variables. The state variables are determined according to the error quaternion vector part in the spacecraft kinematic and dynamic models and the spacecraft angular velocity. The attitude control law is determined according to a sign function matrix gain, the state variables, a non-linear part determined based on the non-linear basis, a linear control gain, and a non-linear control gain; Determining a representative value or an average value within the design range of the spacecraft's principal axis moment of inertia as the nominal value of the principal axis, and converting the system matrices corresponding to the non-linear basis and the attitude control law in the target dynamic model into nominal system matrices based on the nominal value of the principal axis to obtain a nominal dynamic model of the spacecraft attitude; Determining the linear quadratic regulation gain of the linear part of the nominal dynamic model of the spacecraft attitude under a cost function as the initial linear control gain, and determining a zero matrix as the initial non-linear control gain. The cost function is used to characterize the correlation relationship between the cost function value, a state weight matrix, a control input weight matrix, state variables, and an attitude control law; Taking the initial linear control gain and the initial non-linear control gain as the initial control gains of the spacecraft attitude controller, and adjusting the spacecraft attitude controller from the initial control gains to the target control gains.

2. The method according to claim 1, wherein The target dynamic model is established through the following steps: Using predefined state variables Decompose the kinematic and dynamic models of the spacecraft into linear and non - linear parts, where q ev represents the error quaternion vector part in the kinematic and dynamic models of the spacecraft, ω represents the angular velocity of the spacecraft, and the superscript represents vector transpose; Discretizing each decomposed part through a sampling time h to establish a target dynamic model, and the target dynamic model is expressed as follows: x t+1 = A 1t x t + C t φ1(x t ) + B t u t , q e0 (0) ≥ 0 x t+1 = A 2t x t + C t φ2(x t ) + B t u t , q e0 (0) < 0 where x t+1 represents the closed-loop trajectory data, and φ1(x t ), φ2(x t ) both represent the non-linear basis, u t represents the attitude control law, q e0 represents the scalar part of the error quaternion in the spacecraft kinematics and dynamics model, A 1t , A 2t , B t , C t represent the system matrices respectively, and: Among them, the constant matrix C J1 and C J2 are respectively related to the moment of inertia of the spacecraft and unknown. The matrices I3 and 03 represent the 3×3 identity matrix and the zero matrix respectively.

3. The method according to claim 2, wherein The nominal dynamic model of the spacecraft attitude is expressed as follows: Among them, x t+1 represents the closed-loop trajectory data, x t represents the state variable, φ1(x t ) represents the non-linear basis, u t represents the attitude control law, A it represents the system matrix, represents the nominal system matrix, and h represents the sampling time, the matrices I3 and 03 represent the 3-by-3 identity matrix and the zero matrix respectively, and J0 represents the spindle nominal value; The cost function is expressed as follows: Among them, C(K s ) represents the cost function value, x t represents the state variable at time step t, Q and R respectively represent the state weight matrix and the control input weight matrix, u t represents the attitude control law at time step t, and the superscript represents the vector transpose; The step of determining the linear quadratic regulation gain of the linear part of the nominal dynamic model of the spacecraft attitude under a cost function as the initial linear control gain includes: The linear part of the nominal spacecraft attitude dynamics model The linear quadratic regulation gain under the cost function is determined as the initial linear control gain 4. The method according to claim 2, wherein The spacecraft kinematic and dynamic models are expressed as follows: Based on the spacecraft kinematic and dynamic models, the non-linear basis is expressed as follows: Based on the non-linear substrate, the non-linear part φ(ω t ) is expressed as follows: Among them, represents the derivative of the scalar part q of the error quaternion, e0 and represents the derivative of the vector part of the error quaternion, represents the derivative of the angular velocity of the spacecraft and and ω × respectively represent the skew-symmetric matrices of q ev and ω, and J s represents the moment of inertia of the spacecraft, u represents the control torque, and the superscript represents the vector transpose.

5. The method according to claim 1, characterized in that, The attitude control law is determined through the following formula: u t = - diag[sgn(q e0 (0))I3, I3]K x x t -K J φ(ω t ) where, u t represents the attitude control law, diag[sgn(q e0 (0))I3, I3] represents the sign function matrix gain, K x and K J represent the linear control gain and the nonlinear control gain respectively, x t represents the state variable, φ(ω t ) represents the nonlinear part.

6. The method according to any one of claims 1-5, characterized in that The step of adjusting the spacecraft attitude controller from the initial control gains to the target control gains includes: Automatically adjusting the spacecraft attitude controller from the initial control gains to the target control gains based on a policy optimization algorithm; Wherein, the iterative form of the control gain in the policy optimization algorithm is expressed as follows: Among them, respectively represent the (n + 1)-th and n-th iteration values of the control gain, and η represents the iteration step size. represents the gradient estimation.

7. The method according to claim 6, wherein The step of automatically adjusting the spacecraft attitude controller from the initial control gains to the target control gains based on a policy optimization algorithm includes: Under the condition of meeting the set conditions, automatically adjusting the spacecraft attitude controller from the initial control gains to the target control gains based on a policy optimization algorithm, and the set conditions are expressed as follows: where ΔJ represents the maximum deviation between the true value and the nominal value of the spindle inertia, h represents the sampling time, J0 represents the spindle nominal value, C B , C C represents a constant, and: Among them, the constant ρ1 satisfies ρ1 ∈ [(ρ ini +1) / 2, 1), and c ini , ρ ini respectively represent the linear convergence coefficient and the exponential convergence coefficient of the linear part of the nominal dynamic model of the spacecraft attitude under the initial linear control gain . The constant c2 satisfies The constant 8. The method according to claim 6, wherein The gradient estimation in the policy optimization algorithm is determined through the following steps: For each trajectory, determine the control gain through the first formula and sample the initial state variable x0; For each trajectory, when the number of time steps t has not reached the set iteration length T, iteratively execute the following steps: Determine the attitude control law through a second formula; Collect closed-loop trajectory data x t+1 and the immediate cost function value c t ; For each trajectory, when the number of time steps t reaches the set iteration length, the cost function estimate is determined by the third formula Determine the gradient estimation through a fourth formula; The first formula is expressed as follows: The second formula is expressed as follows: u t = -K x x t -K J φ(ω t ) The third formula is expressed as follows: The fourth formula is expressed as follows: Among them, represents the (i - 1)-th iteration value of the control gain K s , U (i) represents a matrix randomly selected from the set of predefined matrices, K x and K J represent the linear control gain and the non - linear control gain respectively, x t represents the state variable, φ(ω t ) represents the non - linear part, represents the gradient estimation, N represents the number of trajectories, p and n respectively represent the row dimension and the column dimension of the control gain K s and r represents the smoothing parameter.

9. The method according to claim 8, wherein The closed-loop trajectory data x t+1 is determined by the following formula: x t+1 = A 1t x t + C t φ(x t ) + B t u t Among them, x t represents the state variable, φ(x t ) represents the non-linear basis in the target dynamics model, u t represents the attitude control law, A 1t , B t , C t represent different system matrices in the target dynamics model; The immediate cost function value c t is determined by the following formula: where x t represents the state variable, u t represents the attitude control law, Q and R respectively represent the state weight matrix and the control input weight matrix, and the superscript represents the vector transpose.

10. A spacecraft attitude controller, characterized in that, It includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the control gain adjustment method according to any one of claims 1 to 9.