A position-attitude integrated control method for space optical imaging formation system based on reinforcement learning

Through a reinforcement learning-based method combined with algebraic graph theory and dual algebra, a position-attitude integrated control method for spacecraft formations is designed, which solves the control performance optimization and timeliness problems of spacecraft formations under nonlinear conditions, and realizes efficient control of the spacecraft formation configuration reconstruction task.

CN117492466BActive Publication Date: 2025-09-12BEIHANG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311402346.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-10-26
Publication Date
2025-09-12
Estimated Expiration
2043-10-26

AI Technical Summary

Technical Problem

Existing spacecraft formation control methods are difficult to simultaneously ensure the optimization of motion nonlinear indicator performance and computational timeliness under highly nonlinear conditions. Existing methods such as sliding mode control and deep reinforcement learning have shortcomings in resource consumption and real-time performance.

Method used

A reinforcement learning-based method is adopted, combined with algebraic graph theory and dual algebra, to establish a position-attitude integrated dynamic model of the spacecraft formation, design a comprehensive reward function, and optimize the controller parameters through online reinforcement learning to achieve collaborative control of the spacecraft formation.

Benefits of technology

The accuracy of the spacecraft formation configuration reconstruction mission and the real-time optimization of the control performance have been achieved, meeting the timeliness and economy requirements of the space mission.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117492466B_ABST
    Figure CN117492466B_ABST
Patent Text Reader

Abstract

The present invention discloses a position-attitude integrated control method for a space optical imaging formation system based on reinforcement learning. First, the communication topology relationship of the spacecraft formation is established based on algebraic graph theory, the position-attitude integrated dynamic model of the spacecraft is expressed based on dual algebra, and a compact dynamic model of the spacecraft formation is established according to the formation configuration error and the position-attitude tracking error. Then, a reward function is designed according to the imaging quality requirements of the space telescope and the tracking error of the formation system. Finally, based on the imaging quality and configuration error, a real-time parameter learning law of the controller is designed using online data, so that the controller is gradually upgraded from a simple control strategy to an optimal controller, thereby improving the execution efficiency of the spacecraft formation mission.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of spacecraft formation control, and in particular relates to a position and attitude integrated control method for a space optical imaging formation system based on reinforcement learning. Background Art

[0002] With the rapid development of aerospace technology, driven by the increasingly complex demands of national defense and scientific research missions such as deep space exploration and planetary observation, the organizational model for space missions has evolved from single spacecraft to spacecraft formations. On the one hand, spacecraft formations are characterized by clustering and spatial distribution. To complete a designated mission, the spacecraft formation system establishes a communication topology and, within this communication relationship, must form a desired configuration. In practical applications such as the construction of spatially distributed optical imaging systems and gravitational wave interferometry, physical performance indicators (such as imaging quality and measurement accuracy) are directly related to the formation configuration and exhibit complex nonlinearities. On the other hand, due to the advantages of small size and light weight of formation spacecraft, their onboard fuel is limited. Considering the limited resources of spacecraft formations and the timeliness of missions, motion control during the spacecraft formation configuration reconstruction process, which combines optimal timeliness and economy, is a critical factor in the design of future spacecraft formation control systems. Therefore, it is particularly important to study a coordinated control problem for spacecraft formations that considers the optimization of highly nonlinear physical indicators and the energy performance of mission completion time.

[0003] In relevant research at home and abroad, current approaches to nonlinear control of spacecraft formations include the following: one is to linearize the system dynamics by selecting nonlinear feedback, but this places high demands on the control system; the other is to construct a sliding surface and make the system slide on the sliding surface to achieve the control target. However, these methods may lead to shock phenomena, and both of the above methods have difficulty optimizing nonlinear performance indicators. Deep reinforcement learning has great potential for processing complex nonlinear systems, but existing methods require large amounts of data and are time-consuming, and are not timely enough to meet the needs of space missions. On the other hand, for the problem of optimizing spacecraft formation performance, existing methods, including optimal control algorithms and model predictive control, consume large amounts of computing resources and lack real-time performance. Therefore, existing spacecraft formation control methods, when the system is highly nonlinear, cannot simultaneously ensure the optimization of nonlinear motion indicators and computational efficiency. Summary of the Invention

[0004] To solve the above technical problems, the present invention provides a position and attitude integrated control method for a space optical imaging formation system based on reinforcement learning, which takes into account the optimization of highly nonlinear physical indicators and the optimization of mission completion time and energy consumption performance to achieve collaborative control of spacecraft formations.

[0005] In order to achieve the above object, the present invention adopts the following technical solutions:

[0006] A method for integrated position and posture control of a space optical imaging formation system based on reinforcement learning comprises the following steps:

[0007] S1: Based on algebraic graph theory, the communication topology of the spacecraft formation is established. According to the dynamic characteristics of the spacecraft formation in the position-attitude integrated control task, a position-attitude integrated dynamic model of the spacecraft formation six-degree-of-freedom maneuver task is established based on formation configuration error and position tracking error based on dual algebra.

[0008] S2: Based on the requirements of formation telescope imaging, the spherical aberration parameters based on the current formation spacecraft posture are designed based on dual algebra. The spherical aberration parameters are incorporated into the design of the reward function, and the comprehensive reward function related to the mission is obtained by combining the configuration error state term and the control cost design.

[0009] S3: Combining the position-attitude integrated dynamics model and comprehensive reward function of the spacecraft formation system, design a spacecraft formation position-attitude integrated control method based on online reinforcement learning.

[0010] Beneficial effects:

[0011] The present invention provides a method for controlling the position and attitude of a space optical imaging formation system based on reinforcement learning. First, based on a dual algebraic framework, a position and attitude integrated dynamic model of a spacecraft formation six-degree-of-freedom maneuvering task based on a formation configuration error is established, and a communication topology relationship of a spacecraft formation is established based on algebraic graph theory. Then, according to the requirements of formation-type space telescope imaging, based on an algebraic framework, spherical aberration parameters based on the current position and attitude of the formation spacecraft are designed, the spherical aberration is integrated into the design of the reward function, and the reward function is obtained by combining the configuration error state term and the control cost design. Finally, the dynamic model of the formation system is combined with the control cost. and reward functions, and designs a spacecraft formation posture integrated control method based on online reinforcement learning; in this way, by designing a reward function related to the task function, constructing a residual cost function, and using online data to design the real-time parameter learning law of the controller, it is possible to solve the spacecraft formation collaborative control problem that takes into account the optimization of highly nonlinear physical indicators and the optimization of task completion time and energy consumption performance. By autonomously improving the performance of the optimized controller through real-time learning, the controller can be gradually upgraded from a simple control strategy to a suboptimal controller by utilizing online data, thereby improving the execution efficiency of the spacecraft formation control system configuration reconstruction task. Compared with existing methods such as sliding mode control, feedback linearization, and deep reinforcement learning, the present invention uses a method based on online learning control, which can not only optimize nonlinear indicators, but also effectively improve control performance, meet the needs of real-time solution, and improve the economy of the control system and task execution. In summary, the present invention can realize the design of a real-time online learning controller for a highly nonlinear spacecraft formation system, ensure the accuracy of posture control in the spacecraft formation configuration reconstruction task, and improve the control performance of the spacecraft formation in real time based on online data. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] Figure 1 A schematic flow chart of a method for integrated position and posture control of a space optical imaging formation system based on reinforcement learning provided by the present invention;

[0013] Figure 2 This is a principle block diagram of the position and posture integrated control method of the space optical imaging formation system based on reinforcement learning provided by the present invention. DETAILED DESCRIPTION

[0014] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only for illustration and are not intended to limit the present invention.

[0015] The present invention provides a method for controlling the position and posture of a space optical imaging formation system based on reinforcement learning. Figure 1 As shown, the following steps are included:

[0016] Step S1: Based on algebraic graph theory, the communication topology of the spacecraft formation is established. According to the dynamic characteristics of the spacecraft formation in the position-attitude integrated control task, a position-attitude integrated dynamic model of the spacecraft formation six-degree-of-freedom maneuver task based on the formation configuration error and the position-attitude tracking error is established based on dual algebra.

[0017] Step S2: Based on the requirements of formation telescope imaging, the spherical aberration parameters based on the current formation spacecraft posture are designed based on dual algebra. The spherical aberration parameters are incorporated into the design of the reward function, and the comprehensive reward function related to the mission is obtained by combining the configuration error state term and the control cost design.

[0018] Step S3: Combining the position-attitude integrated dynamics model and comprehensive reward function of the spacecraft formation system, a spacecraft formation position-attitude integrated control method based on online reinforcement learning is designed.

[0019] The following describes in detail the specific implementation of the above-mentioned reinforcement learning-based position and posture integrated control method for the space optical imaging formation system provided by the present invention through a specific embodiment.

[0020] Example 1:

[0021] In a possible implementation, step S1 of the embodiment of the present invention includes:

[0022] Design a leader-follower spacecraft formation configuration, and the communication diagram is composed of and a set of nodes where v i is the i-th spacecraft in the communication graph, n is the number of spacecraft in the formation, and a set of edges and the adjacency matrix Composition, where a ij Represents the adjacency relationship between spacecraft i and j; if (v j ,v i )∈ε, then there is a communication path from the jth spacecraft to the ith spacecraft, and the in-degree matrix D is defined as D=diag(d1,d2,...,d i ,...,d n ), where diag(·) is used to construct a diagonal matrix and return the diagonal elements of a matrix. represents the in-degree of the i-th spacecraft, N i represents the set of nodes adjacent to i; the communication authority of the follower to the leader is defined as B = [b1, b2, ..., b n ], where b i >0 means that the i-th follower can obtain the leader's information, otherwise b i = 0. The Laplace matrices L and B can be chosen as:

[0023]

[0024] Based on the dual algebraic properties, the compact form error dynamics equation of the spacecraft formation posture is established as:

[0025]

[0026] in, represents the coordinate transformation matrix between the body coordinate systems of follower spacecraft i and j, Represents the state difference of spacecraft i and j, and the superscript represents the coordinate system of the spacecraft. denote the control inputs of spacecraft i and j respectively, Represents a 12-dimensional real vector, and the superscript T represents the transpose of the matrix. represents the relative pose error, represents the relative position deviation between follower spacecraft i and j, represents the deviation of the follower spacecraft from the desired pose, represents the multiplication of dual quaternions, represents the conjugate of the dual quaternion, represents the commutative operation of the dual quaternion, that is, the real part and the dual part are interchanged. The operation represents taking the vector part of the dual quaternion, is the unit quaternion, represents the tracking error relative to the dual angular velocity, represents the dual angular velocity of the follower spacecraft i, is the dual inertia matrix of the spacecraft, defined as The value can be m i =20kg, J i =diag([20.8,21.1,32.6])kg·m 2 , represents the coordinate transformation matrix between the body coordinate systems of follower spacecraft i and j, It represents the control force and control torque of follower spacecraft i, and the control matrix is ​​G i =G j =[0 6×6 ,I 6×6 ] T , nonlinear function in:

[0027]

[0028]

[0029] in, represents the deviation of the follower spacecraft from the desired angular velocity, represents the deviation of the angular velocity of follower spacecraft i and j, is the cross-product matrix operator of the dual vector,

[0030] In a possible implementation, step S2 of the embodiment of the present invention specifically includes:

[0031] The follower equipped with a reflector follows the leader to adjust its direction and position so that the light reflected by the follower is focused on the image sensor of the leader. The imaging accuracy of the telescope is determined by the spherical aberration parameter t i It is expressed as follows: This parameter is obtained by algebraizing the dual quaternion of the distance between the reflected light and the focus:

[0032]

[0033] Among them, ||·|| represents the 2-norm of the vector, · represents the inner product of the vector, represents the dual quaternion of the i-th follower's coordinate system relative to the leader, represents the real part, represents the dual part, z i F Bi ={O Bi -X Bi Y Bi Z Bi}Coordinate system lower edge Z Bi The unit vector in the axis direction, c is F L ={O L -X L Y L Z L}Coordinate system lower edge Z L Unit vector in the direction of the axis;

[0034] The comprehensive reward function of the formation in the configuration reconstruction task consists of two parts: the expected motion state term and the control cost term. Combined with the definition of the spherical aberration parameter of imaging quality, the comprehensive reward function of the expected state is designed as follows:

[0035]

[0036] Where, α ei >0,α di >0, Represents the dual vector multiplication algorithm, represents the weight factor of the state-related reward function, considers the control cost in the reward function, and gives the instantaneous performance index function including the desired motion state and control cost As a reward function:

[0037]

[0038] Where, α ii >0,α ij ≥0 is the weight factor of the control cost-related reporting function. Considering that the number of formation follower spacecraft is 4, the weight corresponding to the state variable in the above parameters can be set to α e1 =0.012,α e2 =0.015,α e3 =0.003,α e4 =0.007,α ti =0.001, i=1,2,3,4, control cost weight α ii ,α ij The weight matrix α u for:

[0039]

[0040] In a possible implementation, step S3 of the embodiment of the present invention specifically includes:

[0041] Based on the design of the above comprehensive reward function, the residual cost function corresponding to the control strategy in the pose approximation task during the spacecraft formation configuration reconstruction process can be defined as follows:

[0042]

[0043] It is expected to find an optimal control to minimize the cost function:

[0044]

[0045] When the above formula satisfies When the inequality is true, the N-tuple of the optimal control strategy is The Nash equilibrium solution of the spacecraft formation system and the optimal control strategy are expressed as:

[0046]

[0047] in, Indicates V i * about The partial derivative of V is calculated using the following network form. i An approximate estimate of:

[0048]

[0049] in, represents the network basis function; Represents the estimated weight vector corresponding to the network basis, Combining the approximate estimate of and the optimal controller, the approximate optimal controller for the spacecraft formation pose approach task is obtained as follows:

[0050]

[0051] The Bellman error is defined as:

[0052]

[0053]

[0054] in, is the network weight estimation error, represents the induced reconstruction error.

[0055] The estimated weight vector corresponding to the network basis The learning update law is designed as follows:

[0056]

[0057] Among them, β i1 >0,β i2 >0 is the learning update rate parameter, η i =θ i / (1+θ i T θ i ), ψ i =η i / (1+θ i T θ i ), augmented vector Ξ i is defined as:

[0058]

[0059] in, t k2 ,t k1 To satisfy the moment of this formula, κ is a positive constant. When i is 4, the parameter in the learning update rate of the network weight vector is designed to be β 11 =1,β 12 =0.005,β 21 =1.5,β 22 =0.001,β 31 =0.4,β 32 =0.001,β 41 =1,β 12 =0.001,κ=10 -4 ,In addition, the network weight vectors corresponding to the initial control strategies of the four follower spacecraft are assigned as follows:

[0060]

[0061]

[0062]

[0063]

[0064] Through the learning update law of formula (11) and the selection of the initial network weights mentioned above, stable integrated control of the spacecraft formation position and attitude can be achieved.

[0065] like Figure 2 The figure below shows the principle block diagram of the reinforcement learning-based position-integrated control method for the space optical imaging formation system provided by the present invention. It primarily consists of a judgment network, a reward network and design add-ons, a learner, a controller, and a six-degree-of-freedom dynamic model for the spacecraft formation. First, the initial controller performs control tasks for the spacecraft. The judgment network and reward network collect data to evaluate the overall control performance of each spacecraft in the formation. Simultaneously, each spacecraft's learner uses the evaluation results to learn its own network weights in real time and update the control parameters in the controller to achieve online performance improvement.

[0066] Obviously, those skilled in the art may make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if such changes and modifications fall within the scope of the claims and their equivalents, the present invention is intended to include such changes and modifications.

Claims

1. A method for integrated position and posture control of a space optical imaging formation system based on reinforcement learning, characterized in that: The steps include: S1: Based on algebraic graph theory, the communication topology of the spacecraft formation is established. According to the dynamic characteristics of the spacecraft formation in the position-attitude integrated control task, based on dual algebra, a position-attitude integrated dynamic model of the spacecraft formation six-degree-of-freedom maneuvering task based on the formation configuration error and position-attitude tracking error is established. The method includes: First, the communication topology graph is designed based on algebraic graph theory: Design a leader-follower spacecraft formation configuration, and the communication diagram is composed of and a set of nodes ,in In the communication diagram A spacecraft, represents the number of spacecraft in the formation, a set of edges and the adjacency matrix Composition, of which Represents spacecraft adjacency relationship; if , then there exists Spacecraft to The communication path of the spacecraft, the in-degree matrix Defined as ,in Used to construct a diagonal matrix and return the diagonal elements of a matrix. Indicates the The in-degree of the spacecraft, Represents A set of adjacent nodes; the follower's communication authority to the leader is defined as ,in Indicates the Followers get information from the leader, and vice versa ; Based on the dual algebraic properties, the compact form error dynamics equation of the spacecraft formation posture is established as: (1) in, , , Indicates follower spacecraft and The coordinate transformation matrix between the body coordinate system, Represents spacecraft The state difference, the superscript represents the body coordinate system, Represents spacecraft The control input, ,in, represents the relative pose error, represents the multiplication of dual quaternions, represents the conjugate of the dual quaternion, represents the commutative operation of the dual quaternion, that is, the real part and the dual part are interchanged. The operation represents taking the vector part of the dual quaternion, Indicates follower spacecraft and The relative posture deviation between represents the deviation of the follower spacecraft from the desired pose, , is the unit quaternion; represents the tracking error relative to the dual angular velocity, Indicates follower spacecraft The dual angular velocity, is the dual inertia matrix of the spacecraft, defined as , is the unit dual operator, for The complementary operator of Indicates follower spacecraft The control force and control torque, the control matrix is , nonlinear function ,in: in, represents the deviation of the follower spacecraft from the desired angular velocity, Indicates follower spacecraft and The deviation of the angular velocity, is the cross-product matrix operator of the dual vector, ; S2: Based on the requirements of formation telescope imaging, the spherical aberration parameters based on the current formation spacecraft posture are designed based on dual algebra. The spherical aberration parameters are incorporated into the design of the reward function, and the comprehensive reward function related to the mission is obtained by combining the configuration error state term and the control cost design. S3: Combining the position-attitude integrated dynamics model and comprehensive reward function of the spacecraft formation system, design a spacecraft formation position-attitude integrated control method based on online reinforcement learning.

2. The method for integrated position and posture control of a space optical imaging formation system based on reinforcement learning as claimed in claim 1, characterized in that: The S2 includes: The follower equipped with a reflector follows the leader to adjust its direction and position so that the light reflected by the follower is focused on the image sensor of the leader; the imaging accuracy of the telescope is determined by the spherical aberration parameter This parameter is obtained by the dual algebraization of the distance between the reflected light and the focus: (2) in, represents the 2-norm of the vector, represents the inner product of vectors, Indicates the The dual quaternion of the follower's coordinate system relative to the leader, represents the real part, represents the dual part, , for Bottom edge of the coordinate system The unit vector in the direction of the axis, for Bottom edge of the coordinate system Unit vector in the direction of the axis; The comprehensive reward function of the formation in the configuration reconstruction task consists of two parts: the expected motion state term and the control cost term. Combined with the definition of spherical aberration parameters of imaging quality, the reward function of the expected state is designed as follows: (3) In the formula represents the weight factor of the state-dependent reward function, Represents the dual vector multiplication algorithm, considers the control cost in the reward function, and gives the instantaneous performance index function including the desired motion state and control cost As a reward function: (4) Where, is the weight factor for controlling the cost-related reporting function.

3. The position and posture integrated control method of the space optical imaging formation system based on reinforcement learning according to claim 2, characterized in that: The S3 includes: Based on the design of the above comprehensive reward function, the residual cost function corresponding to the control strategy in the pose approximation task during the spacecraft formation configuration reconstruction is defined as follows: It is expected to find an optimal control to minimize the cost function: (5) When the above formula satisfies When the inequality is true, the N-tuple of the optimal control strategy is The Nash equilibrium solution of the spacecraft formation system and the optimal control strategy are expressed as: (6) in, express about The partial derivative of , using the following network form as an approximate estimate of : (7) in, represents the network basis function; Represents the estimated weight vector corresponding to the network basis, which will be approximately estimated Combined with the optimal controller, the approximate optimal controller for the spacecraft formation attitude approach task is obtained as follows: (8) in, , define the Bellman error as: (9) (10) in, is the network weight estimation error, represents the induced reconstruction error; The estimated weight vector corresponding to the network basis The learning update law is designed as follows: (11) in, To learn the update rate parameter, , , augmented vector is defined as: (12) in, , To satisfy the moment of this formula, is a positive constant. By using the learning update law of Equation (11) and the reasonable selection of the initial network weights, stable integrated control of the spacecraft formation attitude is achieved.

Citation Information

Patent Citations

  • Finite-time posture fault-tolerant control method for spacecraft formation

    CN109459931A

  • Spacecraft formation attitude directional cooperative control method and related equipment

    CN116466735A