A method for adjusting the attitude of a failed spacecraft based on reinforcement learning

By using reinforcement learning-based methods, a mathematical model and constraint model of spacecraft attitude were established. Combining long-term performance index functions and Critic networks, an adaptive control strategy was designed to solve the problem of rapid attitude adjustment of failed spacecraft under uncertainties in rotational inertia and external disturbances, thus achieving highly reliable and rapid attitude control of the spacecraft.

CN115973454BActive Publication Date: 2026-05-08SHANGHAI AEROSPACE CONTROL TECH INST
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
SHANGHAI AEROSPACE CONTROL TECH INST
Filing Date
2022-12-23
Publication Date
2026-05-08

AI Technical Summary

Technical Problem

Existing technologies cannot effectively solve the problem of rapid attitude adjustment for failed spacecraft, especially in the presence of uncertain rotational inertia and external disturbances, making it difficult to achieve rapid and reliable orbital deorbiting and attitude control.

Method used

A reinforcement learning-based approach is used to establish a mathematical model and constraint model of the spacecraft's attitude. By combining a long-term performance index function and a Critic network, an adaptive control strategy is designed. Using a backstepping control framework and an Action network, the spacecraft's attitude can be rapidly adjusted.

Benefits of technology

In the presence of uncertainties in rotational inertia and external disturbances, this system ensures that the spacecraft can quickly and reliably enter the terminal attitude control region, enhancing the robustness and stability of the control.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115973454B_ABST
    Figure CN115973454B_ABST
Patent Text Reader

Abstract

The application discloses a kind of based on reinforcement learning's failed spacecraft attitude quick adjustment method, comprising: step S1, based on spacecraft attitude end constraint, establish failed spacecraft attitude mathematical model and constraint model;Step S2, based on Long-term performance index function in reinforcement learning algorithm, establish evaluation standard and Critic network;Step S3, based on Backstepping control framework combines Action network and the Critic network, establish adaptive control method, to control failed spacecraft into end constraint domain.The application realizes the quick attitude adjustment of failed spacecraft before attitude motion evolution, enters predetermined ignition maneuver pointing.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of reinforcement learning control technology for control systems, and in particular to a method for rapid attitude adjustment of a failed spacecraft based on reinforcement learning. Background Technology

[0002] Since the launch of the first man-made spacecraft in 1957, spacecraft applications have become increasingly intertwined with the development of human society. However, with the continuous increase in the number of objects entering outer space, the problem of space debris environment has become increasingly prominent. Failed spacecraft are a significant source of space debris in low Earth orbit. After a spacecraft fails, it remains in space for an extended period, occupying orbital resources and potentially triggering a large amount of debris generation, causing serious accidents, and even chain reactions, severely impacting high-value spacecraft and normal spacecraft activities. Therefore, there is an urgent need to develop technologies that enable fully autonomous, highly reliable, and rapid maneuvering control of failed spacecraft.

[0003] Currently under development, deorbiting sails, electro-hydroelectric ropes, solar sails, and electric propulsion systems all have thrust levels in the millinewton range, resulting in poor maneuverability and long orbital deorbit times. When the spacecraft is large or in a high orbit, the extended deorbit time fails to meet the requirements for rapid handling of failed spacecraft. Solid propulsion systems, on the other hand, can generate extremely high total impulse in a short time, enabling rapid ignition maneuvers. Furthermore, they are easily expandable with autonomous functional modules. Under attitude instability conditions, highly reliable autonomous maneuvering decisions through a fully autonomous system eliminate dependence on the spacecraft platform's attitude control capabilities, achieving full autonomy during maneuvers. This makes them an ideal choice for a fully autonomous, highly reliable, and rapid maneuvering system for spacecraft. Summary of the Invention

[0004] This invention addresses the rapid attitude maneuver control of failed spacecraft before their attitude motion evolves. It provides a reinforcement learning-based method for rapid attitude adjustment of failed spacecraft, overcoming uncertainties such as rotational inertia and the influence of external disturbances in the system, and ensuring that the system can reliably and quickly enter the terminal attitude control region.

[0005] To achieve the above objectives, the present invention is implemented through the following technical solution:

[0006] A reinforcement learning-based method for rapid attitude adjustment of a failed spacecraft includes: Step S1, establishing a mathematical model and constraint model of the failed spacecraft's attitude based on end-of-course constraints. Step S2, establishing an evaluation criterion and a Critic network based on the long-term performance index function in the reinforcement learning algorithm. Step S3, establishing an adaptive control method based on the Backstepping control framework combined with the Action network and the Critic network to control the failed spacecraft to enter the end-of-course constraint domain.

[0007] Optionally, step S1 includes: the attitude mathematical model of the failed spacecraft is a dynamic and kinematic model of the attitude of the failed spacecraft, and its calculation formula is as follows:

[0008]

[0009]

[0010] Where q = col(q) v ,q4) is a quaternion-based spacecraft state description, q v =[q1,q2,q3] T The subscript v indicates the quaternion vector part, and q1 to q4 represent the four components of the spacecraft attitude quaternion, respectively; ω = [ω x ,ω y ,ω z ] T ω represents the angular velocity of the spacecraft's own frame B relative to inertial frame I. x ,ω y ,ω z τ, T represent the angular velocities of the spacecraft along the x, y, and z axes, respectively; J represents the positive definite symmetric moment of inertia matrix of the spacecraft; d These are the control torque, external disturbances experienced by the spacecraft, and system modeling errors, respectively; I n Let n represent the n-dimensional identity matrix, where n = 3.

[0011] Optionally, the constraint model of the failed spacecraft includes:

[0012] The terminal constraints of the failed spacecraft are selected based on the thruster installation layout and thrust vector of the failed spacecraft as follows:

[0013] -q m ≤q2≤q m

[0014] -ω m ≤ω y ≤ω m

[0015]

[0016] Where, q m ,ω m ,g min ,g max These are the upper limits for the second attitude quaternion parameter, the upper limit for pitch angular velocity, and the upper limit for the ratio of the third attitude quaternion to the yaw angular velocity.

[0017] The above constraints are simultaneously satisfied through an ellipsoidal constraint domain, wherein the ellipsoidal constraint domain s 2 as follows:

[0018]

[0019] Optionally, step S2 includes: based on the Long-term performance metric function as follows:

[0020]

[0021] Where T > 0 is the small reinforcement learning integral step size; γ ∈ (0,1) is the discount factor; if the control system state enters the attraction domain, the control objective is achieved, and the long-term performance index function J(t) will not increase; if the control system state deviates from the attraction domain, the controller should adjust the control output so that the control system state moves toward the terminal constraint domain or remains in the constraint domain.

[0022] Therefore, the expected performance index J d (t) = 0, define p(s) as a long-term performance index; p(s(ξ)) is as follows:

[0023]

[0024] Among them, s 2 s(t) denotes the ellipsoidal constraint domain at time t, s(ξ) denotes the square root of the ellipsoidal constraint domain at time ξ, ξ is the time variable of the integration, and c p >0 represents the relaxation factor that needs to be designed; that is, p(s(ξ)) = 0 represents a good control output, while p(s(ξ)) = 1 indicates that the current control output is poor; 1 means that the performance index function J(t) continues to increase, which makes the control result worse and the spacecraft attitude deviates from the end-point constraint domain; while 0 means that the performance index function J(t) continues to decrease, which makes the control result better and the spacecraft attitude enters the end-point constraint domain.

[0025] Optionally, step S2 further includes: constructing the Bellman error equation and establishing the relationship between J(tT) and J(t):

[0026] J(tT)=γ -1 (J(t)+p c )

[0027] in, Let the performance index function be the reward / penalty integral over the interval [tT,t].

[0028] The Critic network is solved using the time-difference method:

[0029]

[0030] An RBF neural network is used for estimation to solve for nonlinear performance indices.

[0031]

[0032] Among them, H c (x c (t) is the RBF nonlinear activation function, defined as x c (t)=[s,q v T ,ω T ] T s represents the square root of the ellipsoidal constraint domain. This represents an estimate of the ideal network weights.

[0033] According to the Backstepping control framework, z2 = q v z3=ω-ω c ω c For the design of virtual control quantities

[0034]

[0035] in, k1 is a positive definite diagonal matrix.

[0036] The adaptive law of the RBF neural network is:

[0037]

[0038]

[0039] in, For p c The estimated value, ΔH c (t)=H c (x c (t))-γH c (x c (tT)), Λ c The learning matrix is ​​a positive definite diagonal. Let K be the positive constant to be designed; K = [1,1,1] is the matching matrix of the matrix dimension.

[0040] Optionally, step S3 includes: the adaptive law of the Action neural network is:

[0041]

[0042] Among them, Λ a The learning rate matrix is ​​a positive definite diagonal matrix, k a H is a positive constant to be designed. a (x a ) represents the RBF nonlinear activation function. This represents an estimate of the ideal network weights;

[0043] The highly reliable attitude control law based on reinforcement learning under pre-defined end-state constraints is as follows:

[0044]

[0045]

[0046] in, definition To reduce the computational cost of online estimation, a norm is used for estimation, resulting in... k θ Let represent the positive constant to be designed, and η be the positive definite diagonal learning matrix. For the positive constants to be designed;

[0047] This invention has at least one of the following advantages:

[0048] This invention considers the end-of-course constraints of spacecraft attitude, establishes a mathematical model and constraint model of the attitude of a failed spacecraft, and proposes an evaluation criterion based on the long-term performance index function in reinforcement learning algorithms. It also designs a Critic network, which can solve the attitude maneuver control problem of spacecraft with nonlinear end-of-course constraints.

[0049] Based on the Backstepping control framework, this invention simplifies the controller design process and overcomes uncertainties such as rotational inertia and external disturbances in the system. It ensures high reliability and rapid entry into the end attitude control region, ensures system stability, enhances control robustness, and has potential application prospects. Attached Figure Description

[0050] Figure 1 The flowchart illustrates a reinforcement learning-based attitude adjustment method for a failed spacecraft, as provided in this invention. Detailed Implementation

[0051] The following detailed description, in conjunction with the accompanying drawings and specific embodiments, provides a further detailed explanation of a reinforcement learning-based attitude adjustment method for failed spacecraft proposed in this invention. The advantages and features of this invention will become clearer from the following description. It should be noted that the accompanying drawings are in a very simplified form and use non-precise proportions, used only to facilitate and clearly illustrate the embodiments of this invention. Please refer to the accompanying drawings to make the objectives, features, and advantages of this invention more apparent and understandable. It should be understood that the structures, proportions, sizes, etc., depicted in the accompanying drawings are only for illustrative purposes to aid those skilled in the art and are not intended to limit the implementation conditions of this invention. Therefore, they have no substantial technical significance. Any modifications to the structure, changes in proportions, or adjustments to the size, without affecting the effects and objectives achieved by this invention, should still fall within the scope of the technical content disclosed in this invention.

[0052] This embodiment considers a finite number of large control torques for rapid attitude adjustment when the three-axis angular velocities are large during the self-evolution phase of a failed spacecraft. By presetting the end-stage attitude state and considering uncertainties such as the spacecraft's rotational inertia and external disturbances, a reinforcement learning-based maneuvering control strategy is designed, aiming to propose a learning-based iterative intelligent algorithm. Furthermore, control allocation optimization algorithms such as integer linear programming can be used to achieve highly reliable and rapid entry of the failed spacecraft into the end-stage attitude control region, after which it can be deorbited via thruster ignition. In other words, this embodiment provides a reinforcement learning-based method for rapid attitude adjustment of a failed spacecraft, overcoming uncertainties such as rotational inertia and the influence of external disturbances, ensuring highly reliable and rapid entry into the end-stage attitude control region.

[0053] like Figure 1 As shown, this embodiment provides a method for rapid attitude adjustment of a failed spacecraft based on reinforcement learning, comprising the following steps:

[0054] Step S1: Considering the end-of-course attitude constraints of the spacecraft, establish a mathematical model and constraint model of the failed spacecraft attitude;

[0055] Based on the assumptions and dynamic principles, the attitude dynamics and kinematics model of the failed spacecraft is established as Equation (1):

[0056]

[0057]

[0058] Where q = col(q) v ,q4) is a quaternion-based spacecraft state description, q v =[q1,q2,q3] TThe subscript v denotes the quaternion vector part, and q1 to q4 represent the four components of the spacecraft attitude quaternion, respectively. ω = [ω x ,ω y ,ω z ] T ω represents the angular velocity of the spacecraft's own frame B relative to inertial frame I. x ,ω y ,ω z τ, T represent the angular velocities of the spacecraft along the x, y, and z axes, respectively; J represents the positive definite symmetric moment of inertia matrix of the spacecraft; d These are the control torque, external disturbances experienced by the spacecraft, and system modeling errors, respectively. n Let n represent the n-dimensional identity matrix. In this embodiment, n = 3.

[0059] For any vector χ = [χ1χ2χ3] T , symbol χ × Represent the following oblique symmetric matrix:

[0060]

[0061] The preset end-effector constraints are selected based on the thruster's installation layout and thrust vector as follows:

[0062] -q m ≤q2≤q m

[0063] -ω m ≤ω y ≤ω m (3)

[0064]

[0065] Where, q m ,ω m ,g min ,g max These are the upper limits for the second attitude quaternion parameter, the pitch angular velocity, and the ratio of the third attitude quaternion to the yaw angular velocity, respectively. To ensure that the above constraints are satisfied simultaneously, the following ellipsoidal constraint domain is designed.

[0066]

[0067] For ease of description later, the definition is...

[0068] Step S2: Based on the long-term performance index function in reinforcement learning algorithms, propose evaluation criteria and design the Critic network;

[0069] Based on the long-term performance metric function, the following objective function is designed:

[0070]

[0071] Where T > 0 represents the small reinforcement learning integration step size. γ ∈ (0,1) is the discount factor. If the system state enters the attraction domain, the control objective is achieved, and the long-term performance index function J(t) will not increase. If the system state deviates from the attraction domain, the controller should adjust the control output so that the system state moves toward the terminal constraint domain or remains within the constraint domain.

[0072] Therefore, the expected performance index J d (t) = 0. Define p(s) as a long-term performance metric.

[0073] p(s(ξ)) is designed as follows:

[0074]

[0075] Among them, s 2 s(t) denotes the ellipsoidal constraint domain at time t, s(ξ) denotes the square root of the ellipsoidal constraint domain at time ξ, ξ is the time variable of the integration, and c p A value greater than 0 indicates a relaxation factor that needs to be designed. That is, p(s(ξ)) = 0 represents a good control output, while p(s(ξ)) = 1 indicates a poor current control output. A value of 1 means that the index function J(t) continuously increases, leading to a worse control result and the spacecraft attitude deviating from the terminal constraint domain. Conversely, a value of 0 means that the index function J(t) continuously decreases, leading to a better control result and the spacecraft attitude entering the terminal constraint domain.

[0076] Construct the Bellman error equation and establish the relationship between J(tT) and J(t):

[0077] J(tT)=γ -1 (J(t)+p c (7)

[0078] in, Let the index function be the reward / penalty integral over the interval [tT,t]. This invention proposes to use the time difference method to solve the Critic network.

[0079]

[0080] To facilitate the calculation of the nonlinear performance index J(t), an RBF neural network is used for estimation. H c (x c (t) is the RBF nonlinear activation function, defined as x c (t)=[s,q vT ,ω T ] T s represents the square root of the ellipsoidal constraint domain. This represents an estimate of the ideal network weights.

[0081] Through derivation, the adaptive law of the neural network is designed as follows:

[0082]

[0083]

[0084] in, For p c The estimated value, ΔH c (t)=H c (x c (t))-γH c (x c (tT)), Λ c The learning matrix is ​​a positive definite diagonal. Let K be the positive constant to be designed. K = [1,1,1] is the matching matrix with the dimension of the matrix.

[0085] S3. Based on the Backstepping control framework combined with the Action network, an adaptive control method is designed to ultimately enable the failed spacecraft to enter the terminal constraint domain.

[0086] First, based on the Backstepping control framework, the following coordinate transformations are introduced: z1 = s, z2 = q v z3=ω-ω c ω c For the design of virtual control variables:

[0087]

[0088] in, k1 is a positive definite diagonal matrix.

[0089] Further consideration:

[0090]

[0091] The goal is to make z3→0. Considering that the moment of inertia of the failed spacecraft will change, we assume that the moment of inertia matrix J of the spacecraft is an unknown, positive definite, symmetric constant matrix.

[0092] Define linear multipliers L(a):R 3 →R 3×6

[0093]

[0094] The rotational inertia matrix J of the spacecraft is:

[0095]

[0096] Let α = [J] 11 J 12 J 13 J 22 J 23 J 33 ] T Then there is

[0097] Ja=L(a)α (14)

[0098] definition To reduce the computational cost of online estimation, this invention utilizes a norm for estimation, yielding...

[0099] Further design of Action neural network adaptive law to estimate uncertainties such as disturbances and modeling errors T d

[0100]

[0101] Among them, H a (x a (t) is the RBF nonlinear activation function, x a (t)=[z2 T z3 T ] T .

[0102] Through derivation, the adaptive law of the Action neural network is designed as follows:

[0103]

[0104] Among them, Λ a The learning rate matrix is ​​a positive definite diagonal matrix, k a For the positive constants to be designed, This represents an estimate of the ideal network weights.

[0105] Finally, a highly reliable attitude control law based on reinforcement learning under pre-defined end-state constraints can be designed as follows:

[0106]

[0107]

[0108] in, k θ Let represent the positive constant to be designed, and η be the positive definite diagonal learning matrix. For the positive constants to be designed. The ultimate goal is to achieve the failure of the spacecraft entering the terminal constraint domain.

[0109] This embodiment mainly addresses the problem of rapid attitude maneuver control for a failed spacecraft under conditions of attitude end constraints, rotational inertia uncertainty, and external disturbances. It can be used in spacecraft attitude maneuver control systems.

[0110] This embodiment addresses the rapid attitude maneuver control of a failed spacecraft before its attitude motion evolution. It designs a highly reliable attitude adjustment strategy based on reinforcement learning under pre-set end-state constraints. By pre-setting the attitude end state and considering modeling uncertainties such as the rotational inertia of the failed spacecraft and external disturbances, a spacecraft attitude maneuver control method based on reinforcement learning is designed to achieve rapid attitude adjustment of the failed spacecraft before its attitude motion evolution and to enter the predetermined ignition maneuver direction.

[0111] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0112] It should be noted that the apparatus and methods disclosed in the embodiments herein can also be implemented in other ways. The apparatus embodiments described above are merely illustrative; for example, the flowcharts and block diagrams in the accompanying drawings show the architecture, functionality, and operation of possible implementations of apparatus, methods, and computer program products according to various embodiments herein. In this regard, each block in a flowchart or block diagram may represent a module, program, or part of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than those marked in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in a block diagram and / or flowchart, and combinations of blocks in block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system to perform the specified function or action, or can be implemented using a combination of dedicated hardware and computer instructions.

[0113] In addition, the functional modules in the various embodiments of this article can be integrated together to form an independent part, or each module can exist independently, or two or more modules can be integrated to form an independent part.

[0114] Although the present invention has been described in detail through the preferred embodiments above, it should be understood that the above description should not be considered as a limitation of the present invention. Various modifications and substitutions to the present invention will be apparent to those skilled in the art after reading the above description. Therefore, the scope of protection of the present invention should be defined by the appended claims.

Claims

1. A method for rapid attitude adjustment of a failed spacecraft based on reinforcement learning, characterized in that, Includes the following steps: Step S1: Based on the spacecraft's attitude end constraints, establish a mathematical model and constraint model for the attitude of the failed spacecraft; The constraint models for failed spacecraft include: The terminal constraints of the failed spacecraft are selected based on the thruster installation layout and thrust vector of the failed spacecraft as follows: -q m ≤q2≤q m -oh m ≤ω y ≤ω m Where, q m ,ω m ,g min ,g max These are the upper limits of the second attitude quaternion parameters, the upper limit of the pitch rate, the lower limit of the ratio of the third attitude quaternion to the yaw rate, and the upper limit of the ratio of the third attitude quaternion to the yaw rate, respectively. The above constraints are simultaneously satisfied through an ellipsoidal constraint domain, wherein the ellipsoidal constraint domain s 2 as follows: The mathematical model of the attitude of the failed spacecraft is a dynamic and kinematic model of the attitude of the failed spacecraft, and its calculation formula is as follows: Where q = col(q) v ,q4) is a quaternion-based spacecraft state description, q v =[q1,q2,q3] T The subscript v indicates the quaternion vector part, and q1 to q4 represent the four components of the spacecraft attitude quaternion, respectively; ω = [ω x ,ω y ,ω z ] T ω represents the three-axis rotational angular velocity of the spacecraft's own frame B relative to inertial frame I. x ,ω y ,ω z τ, T represent the angular velocities of the spacecraft along the x, y, and z axes, respectively; J represents the positive definite symmetric moment of inertia matrix of the spacecraft; d These are the control torque, external disturbances experienced by the spacecraft, and system modeling errors, respectively; I n Represents an n-dimensional identity matrix, where n = 3; Step S2: Based on the long-term performance index function in the reinforcement learning algorithm, establish the evaluation criteria and the Critic network; Step S3: Based on the Backstepping control framework combined with the Action network and the Critic network, an adaptive control method is established to control the failed spacecraft to enter the terminal constraint domain.

2. The reinforcement learning-based rapid attitude adjustment method for failed spacecraft as described in claim 1, characterized in that, Step S2 includes: The long-term performance metric function is as follows: Where T > 0 is the small reinforcement learning integral step size; γ ∈ (0,1) is the discount factor; if the control system state enters the attraction domain, the control objective is achieved, and the long-term performance index function J(t) will not increase; if the control system state deviates from the attraction domain, the controller should adjust the control output so that the control system state moves toward the terminal constraint domain or remains in the constraint domain. Therefore, the expected performance index J d (t) = 0, define p(s) as a long-term performance index; p(s(ξ)) is as follows: Among them, s 2 s(t) denotes the ellipsoidal constraint domain at time t, s(ξ) denotes the square root of the ellipsoidal constraint domain at time ξ, ξ is the time variable of the integration, and c p >0 represents the relaxation factor that needs to be designed; that is, p(s(ξ)) = 0 represents a good control output, while p(s(ξ)) = 1 indicates that the current control output is poor; 1 means that the performance index function J(t) continues to increase, which makes the control result worse and the spacecraft attitude deviates from the end-point constraint domain; while 0 means that the performance index function J(t) continues to decrease, which makes the control result better and the spacecraft attitude enters the end-point constraint domain.

3. The reinforcement learning-based rapid attitude adjustment method for failed spacecraft as described in claim 2, characterized in that, Step S2 further includes: Construct the Bellman error equation and establish the relationship between J(tT) and J(t): J(t-T)=γ -1 (J(t)+p c ) in, Let the performance index function be the reward / penalty integral over the interval [tT,t]. The Critic network is solved using the time-difference method: An RBF neural network is used for estimation to solve for nonlinear performance indices. Among them, H c (x c (t) is the RBF nonlinear activation function, defined as x c (t)=[s,q v T ,ω T ] T s represents the square root of the ellipsoidal constraint domain. This represents an estimate of the ideal network weights; According to the Backstepping control framework, z2 = q v z3=ω-ω c ω c For the design of virtual control quantities in, k1 is a positive definite diagonal matrix; The adaptive law of the RBF neural network is: in, For p c The estimated value, ΔH c (t)=H c (x c (t))-γH c (x c (tT)), Λ c The learning matrix is ​​a positive definite diagonal matrix, l c η p , η p ,l Γ Let K be the positive constant to be designed; K = [1,1,1] is the matching matrix of the matrix dimension.

4. The reinforcement learning-based rapid attitude adjustment method for failed spacecraft as described in claim 3, characterized in that, Step S3 includes: The adaptive law of the Action neural network is: Among them, Λ a The learning rate matrix is ​​a positive definite diagonal matrix, k a H is a positive constant to be designed. a (x a ) represents the RBF nonlinear activation function. This represents an estimate of the ideal network weights; The highly reliable attitude control law based on reinforcement learning under pre-defined end-state constraints is as follows: in, definition To reduce the computational cost of online estimation, a norm is used for estimation, resulting in... k θ Let η represent the positive definite diagonal learning matrix to be designed, and l represent the positive definite diagonal learning matrix. h For the positive constants to be designed;

Citation Information

Patent Citations

  • Fault observer and control distribution strategy integrated spacecraft attitude fault tolerance control method

    CN109765920A

  • Reinforced learning-based spatial non-cooperative target parameter self-tuning tracking method

    CN110850719A