A Spacecraft Attitude and Orientation Integrated Control Method Based on SE(3) Orbit and Attitude Coupling Modeling and A3C Training Framework
By using Lie group SE(3)-based orbital attitude coupling modeling and A3C training framework to dynamically adjust the weights of the MPC cost function, the accuracy and robustness issues of orbit and attitude control in spacecraft rendezvous and docking were solved, achieving efficient and stable integrated orbital attitude control.
Patent Information
- Application Number
- CN202511958913.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-24
- Publication Date
- 2026-03-06
- Estimated Expiration
- 2045-12-24
AI Technical Summary
In the rendezvous and docking process of spacecraft, the traditional orbit and attitude control methods have failed to effectively handle the errors caused by orbit and attitude coupling, resulting in low control accuracy. Furthermore, the traditional MPC algorithm has poor robustness when facing complex models, making it difficult to meet the requirements of high precision and rapid maneuvering.
A spacecraft orbit-attitude integrated control method is constructed by adopting the orbit-attitude coupling modeling method based on Lie group SE(3) and dynamically adjusting the weight coefficients of the MPC cost function using the A3C training framework. The relative dynamic model of the spacecraft orbit-attitude integrated representation is constructed in the Lie group SE(3) space, and the pre-trained cost weight network is used to dynamically adjust the weight coefficients of the MPC cost function to optimize the MPC controller.
It improves control accuracy and robustness, maintains optimized control performance in complex environments, reduces energy consumption during control, enhances adaptability to different state variables, and ensures high-precision tracking of both trajectory and attitude.
Smart Images

Figure CN121386429B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of spacecraft control and relates to a final tracking control method for spacecraft during the rendezvous and docking phase. In particular, it relates to an integrated spacecraft attitude control method based on SE(3) attitude coupling modeling and A3C training framework. Background Technology
[0002] For close-range space operations such as rendezvous and docking, high-precision position and attitude control is crucial. During the final approach phase at close range (less than 100 meters), nonlinear terms such as the Coriolis force and centrifugal force become significant, leading to non-negligible errors in some commonly used simplified models. Therefore, a high-precision nonlinear relative dynamics model is necessary. Simultaneously, the timescales of position and attitude motions in such missions are highly similar, meaning the coupling between them cannot be completely ignored; otherwise, the resulting large errors will not meet the requirements for high-precision control. Furthermore, these missions demand high maneuverability from the spacecraft, especially when dealing with non-cooperative targets, requiring rapid tracking of targets undergoing large-angle tumbling. Therefore, the controller design must ensure global stability and rapid adaptability. These challenges highlight the need for both high-precision modeling and the design of highly robust and adaptive controllers.
[0003] In terms of spacecraft attitude and orbit control, traditional methods model attitude and orbital dynamics separately and design independent controllers, ignoring the errors caused by attitude coupling during the close approach phase, resulting in low control accuracy. To address this issue, many studies have been conducted in recent years on modeling and control methods for attitude coupling. The mainstream methods can be divided into three categories: vector algebra-based separation-integration, dual quaternion-based, and Lie group SE(3)-based attitude coupling modeling methods. Among them, the Lie group SE(3)-based attitude coupling method uses a singularity-free direction cosine matrix and position vector to form a matrix belonging to the Lie group SE(3) in three-dimensional Euclidean space, which uniformly expresses attitude and position while avoiding the possibility of singularities in the calculation.
[0004] For control scenarios of spacecraft rendezvous and docking, most research based on orbital attitude coupling spacecraft modeling focuses on designing orbital attitude coupling control laws, or designing sliding mode control, terminal sliding mode control, etc. for orbital attitude coupling control. Such controllers do not give full play to the advantages of high modeling accuracy, but instead increase the complexity of controller design. Summary of the Invention
[0005] This invention leverages the advantages of orbital attitude coupling modeling based on Lie groups (SE(3)) to establish relative dynamic equations for spacecraft orbital attitude coupling based on SE(3), thereby meeting the control requirements during the rendezvous and docking phase. However, existing controller designs fail to fully utilize the advantage of high modeling accuracy. Therefore, this invention deploys the controller design based on Model Predictive Control (MPC). The MPC control algorithm can predict future dynamics based on the model and achieve online replanning, thereby optimizing the control input sequence over a future period. Its characteristics fully utilize the advantages of orbital attitude coupling modeling on SE(3) and improve control accuracy.
[0006] However, the control effectiveness of traditional MPC algorithms largely depends on the design of the cost function weight coefficients. When faced with complex models or complex controlled objects and objectives, the complexity of this design cannot guarantee optimal control performance, resulting in poor robustness to different external disturbances and situations where the order of magnitude difference between state variables is too large. Furthermore, current research on adjusting MPC weight coefficients based on neural networks mostly treats the MPC computation process as a "black box." The generalization ability and control safety of the network training results are highly dependent on the quality of the dataset, which also poses a challenge to control effectiveness.
[0007] Therefore, in order to overcome the defects of the prior art, this invention provides a spacecraft orbit-attitude integrated control method based on SE(3) orbit-attitude coupling modeling and A3C training framework with better control effect. A dynamic model of the spacecraft orbit-attitude integrated representation is constructed in the Lie group SE(3) space. The policy network trained based on A3C is used to dynamically adjust the weight coefficient of each term of the MPC cost function, thereby using the improved MPC controller to realize the orbit-attitude integrated control of the spacecraft in the final tracking process during the rendezvous and docking phase.
[0008] The objective of this invention can be achieved through the following technical solutions:
[0009] A spacecraft attitude and orbit integrated control method based on SE(3) orbital attitude coupling modeling and A3C training framework includes the following steps:
[0010] Construct a relative dynamic model of spacecraft orbital attitude integration in the Lie group SE(3) space;
[0011] An improved MPC controller is obtained by dynamically adjusting the weight coefficients of the MPC cost function using a cost weight network pre-trained based on the A3C training framework.
[0012] Based on the aforementioned relative dynamics model and the improved MPC controller, integrated orbital attitude control is achieved during the spacecraft tracking control process.
[0013] The A3C training framework is an asynchronous A3C framework that embeds MPC gradient calculation information.
[0014] Furthermore, the process of constructing the relative dynamic model includes:
[0015] Based on the mathematical theory of Lie group SE(3), the dynamic equations of single spacecraft orbital attitude coupling are constructed;
[0016] The spacecraft configuration expression is transformed to exponential coordinates on SE(3) to obtain the kinematic equations of single spacecraft orbital attitude coupling in exponential coordinates;
[0017] Based on the dynamic equations and kinematic equations of single spacecraft orbital attitude coupling, a relative dynamic model based on SE(3) is established to represent the integrated orbital attitude of the spacecraft.
[0018] Furthermore, the dynamic equations for the single spacecraft's orbital attitude coupling are expressed as follows:
[0019]
[0020] in, The configuration of the spacecraft is represented by a homogeneous matrix. The symbols represent the rotational angular velocity and translational velocity of an integrated spacecraft. Indicates coordinate system mapping; Represents the action of the Lie algebra se(3) The inverse adjoint operator; This represents the vector composed of the gravitational gradient moment and gravity. The control vector consists of the control torque and the control force. The disturbance vector is composed of the disturbance moment and the disturbance force, and the matrix is... , Represents the spacecraft's moment of inertia matrix. Indicates the mass of the spacecraft. Represents the identity matrix.
[0021] Furthermore, the kinematic equations of the single spacecraft's orbital attitude coupling are expressed as follows:
[0022]
[0023] in, This represents the spacecraft's attitude vector in exponential coordinates. and position vector ; Represents the action on the Lie algebra se(3). The adjoint operator; These represent the expressions used in the derivation process, respectively, concerning the principal rotation angle. The functions of the attitude vector and position vector in exponential coordinates are as follows:
[0024]
[0025]
[0026] Furthermore, the expression for the relative dynamics model representing the spacecraft's integrated orbital attitude is as follows:
[0027]
[0028] in, This represents the attitude and position errors of the two spacecraft. This represents the angular velocity and velocity error between the two spacecraft. For se(3) to act on The finite displacement spinor matrix, To track the inverse of the configuration error matrix between the spacecraft and the target spacecraft, For se(3) to act on The adjoint operator, For se(3) to act on The inverse adjoint operator.
[0029] Furthermore, the discrete mathematical expression of the optimization problem corresponding to the MPC controller is determined based on the state variables and control variables of the relative dynamic model and the scenario requirement constraints.
[0030] Furthermore, the discrete mathematical expression is as follows:
[0031]
[0032] in, It predicts the time domain. It refers to the control trajectory in the prediction time domain. These are model state variables, defined as follows: , These are the weights of the operating cost and terminal cost components of the cost function, respectively. For relative dynamics models, This is the cost function for MPC.
[0033] Furthermore, the pre-training of the cost weight network specifically includes the following steps:
[0034] Build a global policy network and global value network , These represent network parameters, and multiple local networks are initialized based on the global network to form a local Actor-Critic training framework.
[0035] Design a reward function, and perform parallel training on each local network based on the reward function;
[0036] Once all local networks have completed training, the global network that converges based on the reward value is used as the cost weight network that can be used to dynamically generate the weight coefficients of the MPC cost function.
[0037] Furthermore, the expression for the reward function is:
[0038]
[0039] in, Representing the state space and action space respectively. This is the weight value for that item. and These are the baseline values for the returns of the two return functions, respectively.
[0040] Furthermore, the parallel training specifically includes the following steps:
[0041] Step 300: Assign the latest global network parameters to the local network. For each local network, randomly initialize the state variables of the control scenario during training. The training process for the k-th training step includes steps 301-306, as follows:
[0042] Step 301: Based on the current relative state of the spacecraft and the local network parameters, the policy network outputs the weight coefficients of the MPC cost function and updates the MPC cost function accordingly.
[0043]
[0044] Among them, with This represents the weighting coefficients updated to the MPC cost function; For the action space of the training framework; The action space represents the policy network that updates the action space based on the state input in the previous training step. The output obtained;
[0045] Step 302: Apply the MPC controller to the relative dynamic model of the orbital attitude coupling, and construct the cost function using the weight coefficients updated in Step 301 to form the NCMPC solver, obtaining the optimal control quantity and the spacecraft state after control:
[0046]
[0047] in, The control quantity is solved by NCMPC. For the controlled spacecraft attitude error With position error ; Let be the state space after NCMPC control, and also the initial state space for the (k+1)th training step, given by . , and will As the current initial control quantity The three components constitute a symbol. This indicates that the solved control variables are applied to the relative dynamics model to obtain the state variables after the control action. ,symbol This represents the optimal control sequence obtained by NCMPC, denoted by [symbol]. This represents the process of summing the cost function in the prediction time domain and using it for optimization of the NCMPC solver when solving for NCMPC.
[0048] Step 303: Based on the relative states of the spacecraft before and after control and the action space obtained in step 301, the value network outputs a corresponding evaluation, and the advantage of the current strategy is expressed using temporal difference error, defined as follows:
[0049]
[0050] in, This is a discount factor for future value.
[0051] The training objective of the value network is to minimize the temporal difference error, and the gradient policy update formula is:
[0052]
[0053] in, For prediction in the time domain;
[0054] Step 304: Combining the MPC solution information from Step 302 and the definition of policy advantage from Step 303, update the local network parameters according to the policy network update formula with correction terms, and define the process of obtaining the control quantity through MPC solution as a function. The policy network update formula with correction terms is expressed as:
[0055]
[0056] in, To the number of training iterations, This is a correction item;
[0057] Step 305: If the MPC solution is successful, return to step 301 for iterative calculation; if the MPC solution fails, accumulate the number of failures.
[0058] Step 306: If the control requirements are met, update the trained network parameters to the global network and return to step 300 to train the next training step; if the cumulative number of MPC solution failures reaches the set number, exit the loop and obtain the latest network weights from the global network again, and return to step 301 to retrain the scenario.
[0059] Compared with the prior art, the present invention has the following beneficial effects:
[0060] (1) A relative dynamic equation for spacecraft orbital attitude coupling based on SE(3) was established, and MPC control was introduced, which can give full play to the advantages of orbital attitude coupling modeling on SE(3) and improve control accuracy. At the same time, in view of the design problem of the cost function weight of traditional MPC, the weight is adaptively adjusted in each control step of MPC, thereby improving the robustness of MPC control algorithm. Faced with different state variables, the cost value can be kept reasonable and relatively stable during the MPC optimization process, thus achieving better control effect to a certain extent.
[0061] (2) This invention solves the problem of low generalization of the network model when constructing the network mapping of state-cost function weights. This invention uses the A3C asynchronous framework with embedded MPC gradient information to train the cost weight network, which makes the network more generalizable, thereby further enhancing the robustness of the control algorithm and the stability of the control effect. Attached Figure Description
[0062] Figure 1 This is a schematic diagram of the training framework and process of the policy network in this invention;
[0063] Figure 2 These are the positional relationships between the target spacecraft and the tracking spacecraft in the simulations of Examples 1 and 2;
[0064] Figure 3 These are the convergence graphs of the training reward of the policy network used in the simulations of Examples 1 and 2;
[0065] Figure 4 This is the test control response diagram of the NCMPC algorithm constructed from the policy network used in the simulations of Examples 1 and 2;
[0066] Figure 5 This is a comparison chart of the position, attitude, velocity, and angular velocity control responses of the NCMPC, MPC, and MLP algorithms in the simulation of Example 1;
[0067] Figure 6 This is a graph showing the applied torque and force changes during the NCMPC and MPC algorithm control process in the simulation of Example 1;
[0068] Figure 7This is a comparison chart of the cost values of NCMPC and MPC algorithms at each control time step in the simulation of Example 1;
[0069] Figure 8 This is a graph showing the variation of the terminal control quantities of the NCMPC and MPC algorithms with the initial position and attitude errors in the simulation of Example 2;
[0070] Figure 9 This is a graph showing the changes in the terminal control quantities of the NCMPC and MPC algorithms with the initial position and velocity errors during the simulation of Example 2. Detailed Implementation
[0071] The present invention will now be described in detail with reference to the accompanying drawings and specific embodiments. These embodiments are based on the technical solution of the present invention and provide detailed implementation methods and specific operating procedures. However, the scope of protection of the present invention is not limited to the following embodiments.
[0072] This invention provides a spacecraft attitude and orbit integrated control method based on SE(3) attitude and orbit coupling modeling and A3C training framework, comprising the following steps:
[0073] Construct a relative dynamic model of spacecraft orbital attitude integration in the Lie group SE(3) space;
[0074] An improved MPC controller is obtained by dynamically adjusting the weight coefficients of the MPC cost function using a cost weight network pre-trained based on the A3C training framework. The A3C training framework is an asynchronous A3C framework that embeds gradient information for MPC computation.
[0075] Based on the aforementioned relative dynamics model and the improved MPC controller, integrated orbital attitude control is achieved during the spacecraft tracking control process.
[0076] The specific technical means of the above method are described below.
[0077] 1. Relative spacecraft dynamics modeling based on Lie group SE(3) orbital attitude coupling
[0078] As the controlled object of the control loop, the relative dynamic model between the tracking spacecraft and the target spacecraft must first be completed. Combining the Lie group SE(3) orbital attitude coupling modeling method, the specific steps are explained as follows:
[0079] Step 1-1: Based on the mathematical theory of Lie group SE(3), using homogeneous matrices... Indicate the configuration of the spacecraft, and in The rotational angular velocity and translational velocity of the spacecraft are integrated into a single spacecraft orbital attitude coupled dynamic equation, as shown in the following expression:
[0080]
[0081]
[0082]
[0083] in, Represents the action of the Lie algebra se(3) The anti-adjoint operator; symbol Indicates from The mapping to the Lie group exponential coordinate system, applied to the spacecraft velocity vector, is expressed as follows. as follows:
[0084]
[0085] Let be the direction cosine matrix representing the rotation matrix from the body coordinate system fixed to the spacecraft to the inertial reference coordinate system fixed to the Earth's center; This represents the spacecraft's position vector in the inertial reference frame; and Describe the angular velocity and linear velocity in the body coordinate system, respectively; These are the three components of the angular velocity in the body coordinate system. Represents the gravitational gradient moment and gravity The vector formed Control torque and control Forming control vectors, Disturbance torque and disturbance forces Form the perturbation vector. Represents the spacecraft's moment of inertia matrix. Indicates the mass of the spacecraft. Represents the identity matrix. The calculation can follow the following expression:
[0086]
[0087] in, This represents the trace of the spacecraft's moment of inertia matrix. This represents the spacecraft's position vector in the inertial coordinate system. This represents the position vector in the body coordinate system. This represents the perturbation force caused by the J2 term (second-order nonspherical gravitational perturbation coefficient). The perturbation acceleration caused by this force, These are the acceleration vectors of the perturbation acceleration in the inertial coordinate system, respectively. and These are the Earth's gravitational constant and the Earth's equatorial radius, respectively. To perform a matrix antisymmetric transformation on the position vector in the body coordinate system.
[0088] Steps 1-2: To avoid singularities when using parameters such as Euler angles, the spacecraft configuration expression is transformed to exponential coordinates on SE(3), and the dynamic equations are constructed in a globally smooth parameterized manner. This represents the spacecraft's attitude vector in exponential coordinates. and position vector The kinematic equations of a single spacecraft in exponential coordinates are obtained as follows:
[0089]
[0090] in, For the Lie algebra se(3) acting on The adjoint operator; defining the magnitude of the principal rotation angle. , The expanded expression is as follows:
[0091]
[0092]
[0093] Step 1-3: Based on the dynamic and kinematic models of single spacecraft orbital attitude coupling in Step 1-1 and Step 1-2, establish a relative dynamic model based on SE(3) to represent the integrated orbital attitude of the spacecraft.
[0094] In the SE(3) exponential coordinate system, with The attitude and position errors of the two spacecraft are represented by... Let ω represent the angular velocity and velocity error of the two spacecraft. Then the expression for the relative dynamics model is (in the formula, the subscript "t" represents the target spacecraft, and the subscript "e" represents the error):
[0095]
[0096] in, For se(3) to act on The finite displacement spinor matrix, To track the inverse of the configuration error matrix between the spacecraft and the target spacecraft, For se(3) to act on The adjoint operator, For se(3) to act on The inverse adjoint operator. The meanings and expressions of the remaining symbols are the same as in steps 1-2.
[0097] 2. Mathematical expression of the MPC controller and formal design of the cost function
[0098] Based on the constructed relative dynamics model, corresponding state variables and control variables are selected, and scenario requirement constraints are designed to form the discrete mathematical expression of the optimization problem corresponding to the MPC controller of orbital attitude coupling control, as follows:
[0099]
[0100] in, It predicts the time domain; The control trajectory in the prediction time domain is the control vector in step 1-1. These are the lower and upper limits of the control torque, i.e., the minimum control torque and the maximum control torque; These are model state variables, defined as follows: ; These are the weights of the operating cost and terminal cost components of the cost function, respectively. These are the relative dynamics and kinematic equations for the spacecraft's orbital attitude coupling in the aforementioned exponential coordinate system. The cost function for MPC is expressed as follows:
[0101]
[0102] 3. Pre-training of cost weight networks based on embedded MPC computation information
[0103] To address the issue of the MPC algorithm's reliance on cost function weight design, a cost weight network needs to be pre-trained based on the A3C training framework to control the weight coefficients of the dynamically generated cost function (i.e., ...). ),like Figure 1 As shown, the specific training steps are as follows:
[0104] Step 3-1: Construct a global policy network based on the specific state space and action space. and global value network , These represent network parameters, and multiple local networks are initialized based on the global network to form a local Actor-Critic training framework.
[0105] Step 3-2: Design a suitable reward function for the policy network according to the scenario requirements. Based on the application scenario of this invention, the reward function is designed as follows:
[0106]
[0107] in, Represent the state space and action space respectively; This is the weight value for that item. and These are the baseline values for the returns of the two return functions, respectively.
[0108] Step 3-3: Conduct parallel training of each local network. The specific steps are as follows:
[0109] Step 300: Assign the latest global network parameters to the local network. For each local network, randomly initialize the state variables of the control scenario during training. The training process for the k-th training step includes steps 301-306, as follows:
[0110] Step 301: Based on the current relative state of the spacecraft and the local network parameters, the policy network outputs the weight coefficients of the MPC cost function and updates the MPC solver.
[0111]
[0112] Among them, with This represents the weighting coefficients updated to the MPC cost function; The action space of the training framework is composed of two sets of weight coefficients. The action space represents the policy network that updates the action space based on the state input in the previous training step. The output is obtained.
[0113] Step 302: Apply the MPC controller to the orbital attitude coupled spacecraft relative dynamics model, and construct the cost function using the weight coefficients updated in step 301 to form the NCMPC solver, obtaining the optimal control quantity and the controlled spacecraft state:
[0114]
[0115] in, The control quantity is solved by NCMPC. For the controlled spacecraft attitude error With position error ; Let be the state space after NCMPC control, and also the initial state space for the (k+1)th training step, given by . , and will As the current initial control quantity It consists of three elements. Symbol This means applying the solved control quantity to the relative dynamic model constructed in steps 1-3 to obtain the state quantity after the control action. ,symbol This represents the optimal control sequence obtained by NCMPC, denoted by [symbol]. This represents the process of summing the cost function in the prediction time domain and using it for optimization of the NCMPC solver when solving for NCMPC.
[0116] Step 303: Based on the relative states of the spacecraft before and after control and the action space obtained in step 301, the value network outputs a corresponding evaluation, and the advantage of the current strategy is expressed using temporal difference error, defined as follows:
[0117]
[0118] in, This is a discount factor for future value.
[0119] The training objective of the value network is to minimize the temporal difference error, and the gradient policy update formula is:
[0120]
[0121] in, For prediction in the time domain.
[0122] Step 304: Combining the MPC solution information from Step 302 and the definition of policy advantage from Step 303, update the local network parameters according to the policy network update formula with correction terms, and define the process of obtaining the control quantity through MPC solution as a function. The policy network update formula with correction terms can be expressed as:
[0123]
[0124] in, To the number of training iterations, For the correction term, the expanded expression and derivation process are as follows:
[0125]
[0126] in, It is obtained by differentiating the designed reward function and using This represents the gradient of the optimal solution in MPC with respect to the action space. The solution to this gradient can be derived by combining the KKT conditions and implicit function theorem for the corresponding optimization problem in MPC. and The two items are respectively corresponding middle and The location index value is in The corresponding calculated values in the solution can be specifically expressed as follows:
[0127]
[0128] in, The KKT matrix corresponding to the MPC optimization problem. H is the weight derivative matrix; H is the Hessian matrix of the Lagrangian function; and C is the Jacobian matrix of the constraint equations. The term is the mixed second derivative of the cost function;
[0129] Step 305: If the MPC solution is successful, return to step 301 for iterative calculation; if the MPC solution fails, accumulate the number of failures.
[0130] Step 306: If the control requirements are met, update the trained network parameters to the global network and return to step 300 to train the next training step; if the cumulative number of MPC solution failures reaches a certain number, exit the loop and obtain the latest network weights from the global network again, and return to step 301 to retrain the scenario.
[0131] Steps 3-4: After all local networks have completed training, the reward value of the global network converges, resulting in a cost-weighted network that can be used to dynamically generate the weight coefficients of the MPC cost function.
[0132] 4. Construction of the orbital attitude coupling control loop
[0133] The pre-trained cost-weight network is combined with the constructed MPC controller to obtain NCMPC, which is then used in the constructed orbital-attitude coupled relative dynamics model to complete the construction of the closed-loop control loop for integrated orbital-attitude tracking control of the spacecraft. The specific execution process of this control loop includes the following steps:
[0134] Step 4-1: Initialize the tracking scene.
[0135] The aforementioned spacecraft relative dynamics model based on SE(3) orbital attitude coupling uses the relative position, relative attitude Euler angle, relative velocity and relative angular velocity between the target spacecraft and the tracking spacecraft as state variables, and randomly initializes these four state variables for simulating the tracking scenario.
[0136] Step 4-2: Dynamically update the weight coefficients of the cost function of the MPC controller.
[0137] Input the current state variable into the pre-trained cost weight network, and it outputs the weight coefficients of the corresponding cost function. Update the constructed cost function expression.
[0138] Step 4-3: Execute control using the MPC controller with the updated cost function.
[0139] Based on the latest state variables and cost function, perform the MPC optimization solution to obtain the current optimal control variable. It also acts on the constructed relative dynamics model to update the spacecraft's state.
[0140] Step 4-4: Iterative loop execution control.
[0141] Return to step 4-2 until the control termination time.
[0142] The above control method establishes the spacecraft orbital attitude coupling relative dynamic equation based on SE(3), and adaptively adjusts the weights in each control step of MPC to improve the robustness of the MPC control algorithm and achieve better control performance.
[0143] If the above methods are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0144] To illustrate the cost-weighted network training process in this invention (e.g.) Figure 1 To illustrate the advantages of the proposed control framework (as shown), the following two examples demonstrate the benefits of this algorithm in terms of control effectiveness and robustness during the final tracking phase of spacecraft rendezvous and docking. To highlight the optimized control effect of this invention, three control algorithms are used for comparison: NCMPC (the control framework proposed in this invention); MPC (a traditional model predictive control algorithm); and MLP (a control algorithm that directly maps state variables to control variables using a cost-weighted network trained under the AC framework).
[0145] To control the impact of other variables on the control effect, the following two embodiments use the same tracking spacecraft and target spacecraft, with parameters shown in Table 1. The orbital parameters of the target spacecraft are shown in Table 2. Its initial attitude is such that the target spacecraft's body coordinate system coincides with its orbital coordinate system. The control objective in both embodiments is for the tracking spacecraft to complete the tracking of the target spacecraft's attitude and position, while simultaneously achieving zero velocity and angular velocity errors between the two spacecraft. The positional relationship and variable signs between the target and tracking spacecraft are shown in Table 2. Figure 2 As shown.
[0146] Table 1 Spacecraft Parameters
[0147]
[0148] Table 2 Target orbit parameters
[0149]
[0150] The embodiment simulates the external disturbance forces experienced by the spacecraft. and torque It can be expressed as follows, where The initial angular velocity of the target spacecraft, To control simulation time:
[0151]
[0152]
[0153] The NCMPC algorithm in this embodiment uses the same pre-trained cost-weight network, and the training process is as follows: Figure 1 The prediction time domain of MPC is set as follows: The relative state variables between the target spacecraft and the tracking spacecraft are randomly initialized, and the training reward curve converges, as shown in the example. Figure 3 As shown. To test the usability of this pre-trained cost-weight network, Figure 4 The control curves of each state variable under the control of the NCMPC algorithm using this cost-weighted network are shown, and the results indicate that the control of each state variable converges. During the tracking process, the average accuracy of the steady-state errors of the attitude angle, position, angular velocity, and velocity of the two spacecraft remained at [percentage missing]. , , and .
[0154] Example 1
[0155] This embodiment compares the control responses of three control algorithms—NCMPC, MPC, and MLP—under the same initial scenario. The initial relative state parameters between the tracking spacecraft and the target spacecraft are set as shown in Table 3. The prediction time domain of the MPC controller in both the NCMPC and MPC algorithms is set to be the same. Taking the x-axis as an example, the comparison chart of control responses of NCMPC, MPC, and MLP is as follows: Figure 5 .
[0156] Table 3 Initial state variables of the simulation
[0157]
[0158] Figure 5The comparison clearly shows that NCMPC control accelerates the speed of error convergence and the amplitude of oscillation in the control process. Table 4 shows the control indicators of the three algorithms. Rows 1-4 are the norm values of the three-axis position error, attitude error, velocity error and angular velocity error, respectively. Rows 5-8 are the settling time of each state variable control.
[0159] Table 4 Comparison of control indicators of three control algorithms (x-axis)
[0160]
[0161] The comparative results show that the lack of a dynamic model in the training process causes MLP control to fail to accurately map the policy network to the appropriate control force and torque every time when facing control processes with large randomness in the state variables. This results in a much longer settling time than MPC-based control, and the lowest final control accuracy. Compared to traditional MPC control, the NCMPC described in this invention can adjust the weights of the cost function in real time according to the current state during the control process, further shortening the settling time, and achieving relatively high overall control accuracy. Traditional MPC, when faced with large data orders between state variables (e.g., distances in meters versus angles in radians), cannot guarantee that the manually adjusted cost function weights will be applicable to all initial state variables, leading to an inability to guarantee control accuracy for all state variables. This is unacceptable for spacecraft rendezvous and docking missions, where trajectory and attitude tracking accuracy must be ensured simultaneously. In addition to significant advantages in control response speed and average control accuracy, NCMPC also exhibits relatively smaller amplitude and magnitude variations in the applied control torque and force during the control process, which is closely related to energy consumption in spacecraft control. Figure 6 A comparison of the torques and forces applied during the control process by MPC and NCMPC shows that, to achieve the same control objective, the NCMPC control process consumes significantly less energy, and the smaller amplitude variation can, to some extent, avoid instability in the spacecraft control process. Figure 7 The cost function value for each control time step during simulation is represented by a cost weight network. Adjusting the weights results in a significantly smaller cost value at each step, smoother changes, and faster convergence to zero. In contrast, traditional MPC, with its fixed weights, exhibits significantly larger fluctuations in cost value. This is the fundamental reason why NCMPC outperforms traditional MPC in terms of adjustment time and overall control accuracy.
[0162] Example 2
[0163] To test the network generalization and control robustness of the NCMPC algorithm of this invention, this embodiment randomly sets multiple sets of initial state error values for Monte Carlo experiments and compares them with traditional MPC control algorithms.
[0164] The first set of simulations simultaneously changed the initial position and velocity errors between the two spacecraft, using MPC and NCMPC control algorithms respectively. The error values of various state variables at the terminal are displayed as follows: Figure 8 As shown, the left column displays the MPC control results, and the right column displays the NCMPC control results. Similarly, the second set of simulations simultaneously changes the initial position error and attitude error values between the two spacecraft, and the error values of various state variables at the terminal are displayed as follows. Figure 9 As shown in the figure, the blank cells in the results graph indicate that the control results did not converge under the initial conditions, which to some extent suggests that the weight values in the cost function were inappropriate. To ensure the accuracy of the terminal error, the control duration selected for this experiment was 800s.
[0165] pass Figure 8 and Figure 9 The comparison of the two sets of figures shows that, under the NCMPC control framework proposed in this invention, the accuracy of the steady-state error remains basically consistent. However, when using traditional MPC control with fixed weights, the tracking accuracy fluctuates significantly due to the influence of the initial value, and the probability of non-convergence of the control result is relatively high. This indicates that the pre-trained cost-weight network in this invention effectively solves the problem of the MPC control algorithm's sensitivity to cost function weights by obtaining adaptive weight coefficients, enhances the solution effect of MPC and the generalization ability of the controller, and can better meet the control requirements of spacecraft rendezvous and docking terminal tracking missions. Furthermore, in addition to these two sets of experiments, the above conclusions can still be obtained in experiments with different initial variable sets.
[0166] Matters not covered in this invention are common knowledge.
[0167] The preferred embodiments of the present invention have been described in detail above. It should be understood that those skilled in the art can make numerous modifications and variations based on the concept of the present invention without creative effort. Therefore, all technical solutions that can be obtained by those skilled in the art based on the concept of the present invention through logical analysis, reasoning, or limited experimentation on the basis of existing technology should be within the scope of protection defined by the claims.
Claims
1. A spacecraft orbit and attitude integrated control method based on SE(3) orbit-attitude coupling modeling and A3C training framework, characterized in that, The method comprises the following steps: a relative dynamics model of a spacecraft orbit attitude integrated representation is constructed in a Lie group SE(3) space; a cost weight network pre-trained based on an A3C training framework is used to dynamically adjust the weight coefficients of an MPC cost function, so as to obtain an improved MPC controller; orbit attitude integrated control in a spacecraft tracking control process is realized based on the relative dynamics model and the improved MPC controller; the A3C training framework is an A3C asynchronous framework embedded with MPC calculation gradient information; the construction process of the relative dynamics model comprises: a single-spacecraft orbit attitude coupling dynamics equation is constructed based on the mathematical theory of the Lie group SE(3); a single-spacecraft orbit attitude coupling kinematics equation is obtained by converting a spacecraft configuration representation to exponential coordinates on SE(3); a relative dynamics model of a spacecraft orbit attitude integrated representation based on SE(3) is established based on the single-spacecraft orbit attitude coupling dynamics equation and the kinematics equation; the pre-training of the cost weight network specifically comprises the following steps: building a global policy network and a global value network , represent network parameters, and initialize multiple local networks according to the global network to form a local Actor-Critic training framework; a reward function is designed, and each local network is trained in parallel based on the reward function; after all the local networks are trained, a global network with converged reward values is used as the cost weight network that can be used to dynamically generate the weight coefficients of the MPC cost function.
2. The spacecraft orbit and attitude integrated control method based on SE(3) orbit- attitude coupling modeling and A3C training framework according to claim 1, wherein, The expression of the single-spacecraft orbit attitude coupling dynamics equation is as follows: wherein denotes the configuration of the spacecraft, which is a homogeneous matrix; denotes the rotational and translational velocities of the integrated spacecraft, the symbol denotes the coordinate system mapping; denotes the adjoint operator in the Lie algebra se(3) acting on ; denotes the vector composed of the gravitational gradient and the gravitational force, the control vector is composed of the control moment and the control force, the disturbance vector is composed of the disturbance moment and the disturbance force, the matrix , denotes the spacecraft's moment of inertia matrix, denotes the spacecraft's mass, denotes the identity matrix.
3. The spacecraft orbit and attitude integrated control method based on SE(3) orbit- attitude coupling modeling and A3C training framework of claim 1, wherein, The expression of the single-spacecraft orbit attitude coupling kinematics equation is as follows: where denotes the attitude vector of the spacecraft in exponential coordinates and the position vector ; denotes the adjoint operator in the Lie algebra se(3) acting on ; are functions of the principal rotation angles , the attitude vector in exponential coordinates and the position vector in exponential coordinates, respectively, as follows: 。 4. The spacecraft orbit and attitude integrated control method based on SE(3) orbit- attitude coupling modeling and A3C training framework of claim 2, wherein, The expression of the relative dynamics model of the spacecraft orbit attitude integrated representation is as follows: where denote the two spacecraft attitude and position errors, denote the two spacecraft angular velocity and velocity errors; is the finite displacement screw matrix acting on in se(3), is the inverse of the configuration error matrix between the chaser spacecraft and the target spacecraft, is the adjoint operator acting on in se(3), is the co-adjoint operator acting on in se(3).
5. The spacecraft orbit and attitude integrated control method based on SE(3) orbit- attitude coupling modeling and A3C training framework of claim 1, wherein, The discrete mathematical expression of the optimization problem corresponding to the MPC controller is determined based on the state quantity and the control quantity of the relative dynamics model and the scene demand constraint condition.
6. The spacecraft orbit and attitude integrated control method based on SE(3) orbit- attitude coupling modeling and A3C training framework according to claim 5, wherein, The discrete mathematical expression is as follows: wherein is the prediction horizon, is the control trajectory in the prediction horizon, is the model state quantity defined as , are weight values for the running cost and terminal cost parts of the cost function, respectively, is the relative dynamics model, is the cost function of the MPC, are lower and upper limit constraints for the control torque, i.e. the minimum and maximum control torque, respectively.
7. The spacecraft orbit and attitude integrated control method based on SE(3) orbit- attitude coupling modeling and A3C training framework of claim 1, wherein, The expression of the reward function is as follows: wherein, respectively denote the state space and the action space, is a weight value for the term, and are respectively the return baseline values for the two return functions.
8. The spacecraft orbit and attitude integrated control method based on SE(3) orbit- attitude coupling modeling and A3C training framework of claim 1, wherein, The parallel training specifically comprises the following steps: Step 300: assign the current latest global network parameters to the local networks, and randomly initialize the state quantity of the control scene for each local network training; Step 301: output the weight coefficients of the MPC cost function from the policy network according to the current spacecraft relative state and the local network parameters, and update the MPC cost function: wherein, with represents the weight coefficient updated to the MPC cost function; is the action space of the training framework; represents the output obtained by the policy network of the action space updated by the previous training step according to the state input . Step 302: use the MPC controller for the orbit attitude coupling relative dynamics model, and form an NCMPC solver with the weight coefficients obtained in step 301 to form the cost function, so as to obtain the optimal control quantity and the spacecraft state after control: wherein, is the control quantity solved by NCMPC, is the spacecraft attitude error after control and position error ; is the state space after NCMPC control, also the initial state space of the k+1 training step, which is composed of , and the current initial control quantity , and the symbols , the symbol represents the state quantity after applying the to-be-solved control quantity to the relative dynamics model to obtain the control action , the symbol represents the optimal control sequence solved by NCMPC, and the symbol represents the summation of the cost function in the prediction time domain and is used in the optimization process of the NCMPC solver. Step 303: output the corresponding evaluation from the value network according to the spacecraft relative state before and after control and the action space obtained in step 301, and use the time difference error to express the advantage of the current policy, which is defined as follows: wherein, is a discount factor for future value; The training target of the value network is to minimize the time difference error, and the gradient policy update formula is as follows: wherein is the prediction time domain; Step 304: According to the policy network update formula with a correction term, update the local network parameters based on the MPC solution information in step 302 and the definition of the strategy advantage in step 303, and define the process of obtaining the control quantity by MPC solution as a function , and the policy network update formula with a correction term is expressed as: wherein, is the number of training iterations, is the correction term; Step 305: if the MPC solving is successful, return to step 301 for iterative calculation; if the MPC solving fails and the cumulative failure number is greater than a preset threshold, return to step 300 for iterative calculation. Step 306: if the control requirement is completed, the trained network parameters are updated to the global network, and the next training step is returned to step 300 for training; if the accumulated failure number of MPC solving reaches the set number, the loop is exited, the current latest network weight is obtained from the global network, and step 301 is returned to retrain the scene.
Citation Information
Patent Citations
Relative orbit and attitude tracking control method for final approaching section of rendezvous and docking of spacecrafts
CN113485396A
Attitude and orbit integrated tracking control method and device under multi-constraint condition
CN114153222A