Spacecraft autonomous obstacle avoidance rendezvous hierarchical motion planning method fused with reinforcement learning

By combining the dual-delay deep deterministic policy gradient method and the hierarchical planning framework of optimal control theory, the problems of high computational complexity and low solution efficiency in spacecraft rendezvous trajectory planning are solved, and efficient and interpretable autonomous obstacle avoidance rendezvous trajectory planning for spacecraft is realized.

CN121632164APending Publication Date: 2026-03-10CENT SOUTH UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511888189.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-15
Publication Date
2026-03-10

AI Technical Summary

Technical Problem

Existing spacecraft rendezvous trajectory planning methods suffer from high computational complexity and low solution efficiency when dealing with complex scenarios. They also have problems such as strong data dependence, insufficient robustness, and poor strategy interpretability.

Method used

By combining the dual-delay deep deterministic policy gradient method with optimal control theory, a model-data hybrid driven hierarchical motion planning framework is designed. The upper-level planner generates virtual waypoints to handle obstacle avoidance constraints, while the lower-level planner generates locally optimal trajectories, simplifying high-dimensional optimization problems and satisfying dynamic constraints.

Benefits of technology

It improves the computational efficiency and generalization ability of autonomous obstacle avoidance and rendezvous trajectory planning for spacecraft, generates high-quality globally optimized trajectories, and enhances the quality and interpretability of autonomous obstacle avoidance and rendezvous.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121632164A_ABST
    Figure CN121632164A_ABST
Patent Text Reader

Abstract

The invention relates to a spacecraft space rendezvous trajectory planning technology, in particular to a spacecraft autonomous obstacle avoidance rendezvous hierarchical motion planning method fused with reinforcement learning, which comprises the following steps: modeling a spacecraft autonomous obstacle avoidance rendezvous trajectory problem, and constructing a hierarchical planning framework comprising an upper planner and a lower planner; designing an upper planner based on a double-delay depth deterministic strategy gradient method, and obtaining a virtual waypoint through the upper planner; designing a lower-layer planner based on a linear quadratic regulator, and obtaining a local optimal track and a control sequence through the lower-layer planner; and combining the optimal trajectory fragments output by the lower planner into a global optimized trajectory, and controlling the spacecraft through the control sequence so as to enable the spacecraft to move according to the planned trajectory. According to the method, while the overall optimality of the rendezvous trajectory is ensured, the generalization ability to a complex task environment is remarkably enhanced, and finally the quality of the autonomous obstacle avoidance rendezvous trajectory of the spacecraft is improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to spacecraft space rendezvous trajectory planning technology, in particular to a spacecraft autonomous obstacle avoidance rendezvous hierarchical motion planning method fusing reinforcement learning. BACKGROUND

[0002] Spacecraft safe rendezvous is a key link in various space activity scenarios such as manned spaceflight, deep space exploration, on-orbit service, etc., and its trajectory planning problem aims to find a collision-free motion trajectory from the initial state to the target state while meeting relevant constraint conditions. At present, the representative methods of spacecraft rendezvous trajectory planning method mainly include indirect method based on Pontryagin maximum principle, direct method represented by pseudo-spectral method, and heuristic algorithm. With the complication of rendezvous scenarios, higher requirements for autonomy and intelligence are put forward for future space technology development, and it is urgent to promote the paradigm shift of rendezvous trajectory planning method from "model-based" to "model and data-driven fusion". For example, the Actor-Critic reinforcement learning framework fusing expert data around the multi-impulse rendezvous trajectory planning problem realizes the rapid calculation of trajectory under random initial state.

[0003] Through the analysis of the current spacecraft space rendezvous trajectory planning method, the existing technology characteristics and the problems are summarized as follows: the indirect method based on the minimum principle can convert the optimization problem into Hamilton boundary value problem, but its initial value of the adjoint variable is sensitive and the derivation process is complex, which makes it difficult to obtain an analytical solution when dealing with complex problems; the pseudo-spectral method has inherent difficulties in dealing with complex non-convex space obstacle avoidance constraints, and the solving effect is highly dependent on the adopted collocation scheme and initial guess; the heuristic algorithm has the problems of large search space, difficulty in designing fitness function, easy to fall into local optimum, and low solving efficiency in spacecraft space rendezvous trajectory planning; single reinforcement learning method in spacecraft rendezvous trajectory planning faces the challenges of strong data dependence, insufficient convergence robustness, poor strategy interpretability and on-board computing power limitations.

[0004] In view of the above problems, the present application is aimed at spacecraft space rendezvous trajectory planning task, and a model-data hybrid driven hierarchical motion planning framework is designed by combining the Twin Delayed Deep Deterministic policy gradient (TD3) method in reinforcement learning and the linear quadratic regulator in optimal control theory. The upper layer of the framework generates virtual waypoints through the policy network, realizes global exploration and obstacle avoidance constraint processing; the lower layer generates local optimal trajectory based on the model, ensures the dynamics feasibility and trajectory optimality, effectively reduces the computational complexity through problem decoupling, and has good generalization ability and interpretability advantage. SUMMARY

[0005] The present application aims at the deficiencies in the prior art, and provides a spacecraft rendezvous trajectory planning task-oriented, model-data hybrid driven spacecraft rendezvous hierarchical motion planning scheme combining a Twin Delayed Deep Deterministic policy gradient (TD3) method in reinforcement learning and a linear quadratic regulator in optimal control theory.

[0006] To achieve the above-mentioned purpose, the present application provides a spacecraft autonomous obstacle avoidance rendezvous hierarchical motion planning method combining reinforcement learning, comprising the following steps:

[0007] S1, modeling the spacecraft autonomous obstacle avoidance rendezvous trajectory problem, and constructing a hierarchical planning framework including an upper planner and a lower planner;

[0008] S2, designing the upper planner based on a Twin Delayed Deep Deterministic policy gradient method, and obtaining virtual waypoints through the upper planner;

[0009] S3, designing the lower planner based on a linear quadratic regulator, and obtaining a locally optimal trajectory and a control sequence through the lower planner;

[0010] S4, combining the optimal trajectory segments output by the lower planner into a globally optimized trajectory, and controlling the spacecraft through the control sequence to make the spacecraft move according to the planned trajectory.

[0011] Further, in S1, the relative motion of the tracking spacecraft is described in the Earth synchronous orbit spacecraft The relative motion of the tracking spacecraft is described in the Earth synchronous orbit spacecraft

[0012] The relative motion of the tracking spacecraft is described in the Earth synchronous orbit spacecraft The relative motion of the tracking spacecraft

[0013] is described in the Earth synchronous orbit spacecraft wherein, represents the relative position vector of in the orbital coordinate system, is the orbital angular velocity of , and is the control input.

[0014] Further, in S1, the spacecraft state vector is defined, and the spacecraft autonomous obstacle avoidance rendezvous trajectory planning model is represented as:

[0015]

[0016] wherein, is the planning initial time, to plan the termination time, denotes the spacecraft transfers from the initial state to the terminal state with the convergence bias, , , are self-defined weight matrices, is the obstacle avoidance constraint, denotes the set of collision-free rendezvous trajectories, is the orbit dynamics constraint, and are path constraints, which mainly limit the mapping of the spacecraft motion path in time, and are boundary constraints, which ensure that the spacecraft can move from the initial state to the target state.

[0017] Further, the upper planner in S2 includes state space design, action space design and reward function design, the state space is a complete description of the environment by the upper planner at each time step in the iteration process; the action space is the specific decision behavior made by the upper planner at each time step; the reward function is used to provide feedback to the upper planner.

[0018] Further, the reward function in S2 includes distance guidance reward, performance index reward, collision constraint reward and path constraint reward; the distance guidance reward is a position-oriented dense reward; the performance index reward is used to feedback an reward to the upper planner at each time step, which is negatively related to the minimization of the performance index; the collision constraint reward corresponds to a hard constraint, when the constraint is violated, the upper planner receives a negative reward and ends the search; the path constraint reward is used to make the upper planner meet the path constraint condition in the exploration process.

[0019] Further, the upper planner in S2 includes policy training, and the policy training objects include policy network and value network, the policy network outputs the action to be executed according to the state of each step, and updates itself by maximizing the value generated by policy execution; the value network is used to evaluate the value generated by each step of policy execution, and updates itself by time difference algorithm.

[0020] Further, the lower planner in S3 obtains the optimal time-varying state feedback gain by solving the time-varying Riccati differential equation, minimizes the performance index in the form of quadratic form in the limited time domain, and performs closed-loop simulation based on the initial state and the stored feedback gain matrix in the preset time period to obtain the optimal trajectory segment and the control sequence.

[0021] Further, S3 specifically includes the following sub-steps:

[0022] The state deviation system equation is obtained based on the orbit dynamics equation in S1 as follows:

[0023]

[0024] The state deviation system is a linear system, is a system matrix, is an input matrix, and

[0025]

[0026] An optimal control sequence is searched in a finite time domain to minimize a quadratic form performance index of the spacecraft autonomous rendezvous and avoidance trajectory planning model in S1, and according to the Pontryagin maximum principle, has the following state feedback form:

[0027]

[0028] wherein, is a symmetric semi-positive definite matrix, and is given by the solution of a differential Riccati equation:

[0029]

[0030] Meanwhile, the following conditions are satisfied:

[0031]

[0032] The integral time interval is set as:

[0033] ,

[0034] The terminal condition is configured as:

[0035]

[0036] The solution is obtained by integrating from to in the reverse time using the fourth-order Runge-Kutta method, and the time-varying matrix function is obtained by mapping the solution to the actual time ;

[0037] The feedback gain matrix at each time is calculated by the following formula:

[0038]

[0039] Based on the initial state and the stored feedback gain matrix, the closed-loop simulation is performed in to obtain the optimal trajectory segment​​ and control sequence .

[0040] The above-mentioned scheme of the present application has the following beneficial effects:

[0041] The spacecraft autonomous obstacle avoidance rendezvous hierarchical motion planning method provided by the present application simplifies the high-dimensional optimization problem into multiple low-dimensional sub-problems through the design of two-layer planner, takes into account global exploration and local optimization, does not need to perform collision detection, and improves the calculation efficiency of the motion planning process; and combines reinforcement learning with optimal control, trains the upper-layer decision network to generate virtual waypoints to process obstacle avoidance constraints through the double-delay deep deterministic policy gradient method, and adopts a linear quadratic regulator to generate optimal trajectory segments to process orbital dynamics, path and boundary constraints, etc., thereby significantly enhancing the generalization ability to complex task environments while ensuring the overall optimality of the rendezvous trajectory, and ultimately improving the quality of the spacecraft autonomous obstacle avoidance rendezvous trajectory.

[0042] Other beneficial effects of the present application will be described in detail in the subsequent specific embodiments. BRIEF DESCRIPTION OF DRAWINGS

[0043] Figure 1 is a step flowchart of the present application;

[0044] Figure 2 is a hierarchical planning framework schematic diagram of the present application;

[0045] Figure 3 is a spacecraft autonomous obstacle avoidance rendezvous scene schematic diagram in the embodiment of the present application;

[0046] Figure 4 is a reinforcement learning process space rendezvous trajectory planning success rate and cumulative reward curve diagram in the embodiment of the present application;

[0047] Figure 5 is a spacecraft relative distance change curve of a space target in the embodiment of the present application;

[0048] Figure 6 is a spacecraft relative speed change curve of a space target in the embodiment of the present application;

[0049] Figure 7 is a spacecraft relative distance change curve of a space obstacle in the embodiment of the present application. DETAILED DESCRIPTION

[0050] The following specific examples illustrate the implementation of this disclosure. Those skilled in the art can easily understand other advantages and effects of this disclosure from the content disclosed in this specification. Obviously, the described embodiments are only a part of the embodiments of this disclosure, and not all of them. This disclosure can also be implemented or applied through other different specific embodiments, and the details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of this disclosure. It should be noted that, in the absence of conflict, the following embodiments and features in the embodiments can be combined with each other. Based on the embodiments in this disclosure, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this disclosure.

[0051] It should be noted that various aspects of embodiments within the scope of the appended claims are described below. It will be apparent that the aspects described herein can be embodied in a wide variety of forms, and any particular structure and / or function described herein is merely illustrative. Based on this disclosure, those skilled in the art will understand that one aspect described herein can be implemented independently of any other aspect, and two or more of these aspects can be combined in various ways. For example, any number of aspects set forth herein can be used to implement the device and / or practice the method. Additionally, this device and / or method can be implemented using structures and / or functionalities other than one or more of the aspects set forth herein.

[0052] It should also be noted that the illustrations provided in the following embodiments are merely schematic representations of the basic concept of this disclosure. The illustrations only show components relevant to this disclosure and are not drawn according to the actual number, shape, and size of components in implementation. In actual implementation, the type, quantity, and proportion of each component can be arbitrarily changed, and the component layout may be more complex. Furthermore, specific details are provided in the following description to facilitate a thorough understanding of the examples. However, those skilled in the art will understand that the described aspects can be practiced without these specific details.

[0053] like Figure 1 As shown, embodiments of the present invention provide a hierarchical motion planning method for autonomous obstacle avoidance and rendezvous of spacecraft that integrates reinforcement learning. It employs a model-data hybrid-driven hierarchical planning framework. The upper layer of this framework primarily generates virtual waypoints to achieve global exploration and obstacle avoidance constraint handling; the lower layer generates locally optimal trajectories to ensure dynamic feasibility and trajectory optimality. Problem decoupling effectively reduces computational complexity, while also possessing good generalization ability and interpretability. Specifically, this step includes the following sub-steps:

[0054] S1 models the autonomous obstacle avoidance and rendezvous trajectory problem of spacecraft and constructs a hierarchical planning framework that includes an upper-level planner and a lower-level planner.

[0055] In this step, the tracking spacecraft is based on the Clohessy-Wiltshire equations. Spacecraft in geosynchronous orbit Describing the relative motion in the vicinity, we obtain the orbital dynamics equations:

[0056]

[0057] In the formula, In orbital coordinate system and The relative position vector, for orbital angular velocity, To control the input.

[0058] Simultaneously define the spacecraft state vector Using planning time, state deviation weighting, and control input weighting as the performance minimization index, the spacecraft autonomous obstacle avoidance and rendezvous trajectory planning model can be expressed as:

[0059]

[0060] in, To plan the initial moment, To plan the termination time, Indicates the spacecraft from its initial state To terminal status Convergence bias of the transfer , , All are custom weight matrices. To avoid obstacles and constraints, This represents the set of intersecting trajectories that will not collide with obstacles. For orbital dynamic constraints, and Path constraints primarily restrict the time mapping of the spacecraft's motion path. and These are boundary constraints to ensure that the spacecraft can move from its initial state to its target state.

[0061] S2 is designed based on the double-delay deep deterministic policy gradient method to obtain virtual waypoints.

[0062] As mentioned earlier, the main goal of the upper-level planner is to provide virtual waypoints for the lower-level planner, thereby guiding the lower-level planner to generate segmented trajectories that can avoid obstacles. Therefore, in this step, the upper-level planner is designed based on the dual-delay deep deterministic policy gradient method in reinforcement learning, specifically including state space design, action space design, reward function design, and policy training. The configuration process is as follows:

[0063] For the state space, during the iteration process, the upper-level planner provides a complete description of the environment at each time step; therefore, the state space... It contains the following two parts:

[0064]

[0065] in, This indicates the relative position information of the obstacle in the orbital coordinate system. Indicates the number of obstacles. That is, the spacecraft state vector, which includes the spacecraft's state vector. Real-time relative position and velocity information at the start of each time step. By designing a state space, the upper-level planner can make corresponding intelligent decisions under different environments (i.e., different combinations of target states and obstacle positions), thus possessing strong generalization ability.

[0066] The action space represents the specific decision-making actions made by the upper-level planner at each time step. Therefore, in this step, the action space of the upper-level planner is used to specify the spacecraft. The virtual waypoint to be reached.

[0067] The reward function is used to provide feedback to the upper-level planner. Therefore, in this step, the formula for the reward function is as follows, consisting of four parts:

[0068]

[0069] in, Indicates distance-guided reward. This indicates a performance-based reward. This indicates a collision constraint reward. Indicates path constraint reward. , , , These are the corresponding weighting coefficients.

[0070] Furthermore, regarding the distance to the introductory reward This indicates a dense reward system for target location orientation, when the spacecraft... The closer The higher the reward value, the better. The specific formula is as follows:

[0071]

[0072] in, and Represent the spacecraft in the orbital coordinate system respectively The position vectors at the previous time step and the current time step.

[0073] Performance Indicator Rewards This means that at each time step, a reward negatively correlated with minimizing the performance metric is fed back to the upper-level planner to generate the most optimal trajectory possible. The specific formula is as follows:

[0074]

[0075] The constant term 1 represents the number of time steps taken, corresponding to the shortest time in the S1 performance index. The other terms quantify the convergence of the spacecraft's state deviation and control input at the current moment. The terminal moment for planning the trajectory segment for the lower-level planner.

[0076] For collision constraint rewards Considering that the goal of the upper-level planner is to find a collision-free global path, the collision constraint is a hard constraint. When this constraint is violated, the upper-level planner will receive a negative reward and terminate the search. Therefore, the specific formula for the collision constraint reward is:

[0077]

[0078] in, This is the collision constraint penalty coefficient. Indicates obstacles The relative position vector in the orbital coordinate system This indicates the minimum safe distance for avoiding collisions with spatial obstacles.

[0079] For path constraint rewards In this embodiment, the geometric path formed by virtual waypoints will be used as the initial solution in subsequent trajectory optimization. Therefore, it is desirable to satisfy the path constraints in S1 as much as possible during the exploration process. The specific formula is as follows:

[0080]

[0081] in, This is the path constraint penalty coefficient, which means that at each time step, if the path constraint condition is not met, a negative reward will be given.

[0082] Therefore, at each time step, when the upper-level planner interacts with the environment, it receives a reward from the environment for each action it performs. When the total number of time steps in the task exceeds a preset value, the reward received by the upper-level planner becomes 0, and the task terminates.

[0083] In this embodiment, the policy training object includes the policy network during policy training. and value network The training method is based on the Dual-Delay Deep Deterministic Policy Gradient (TD3) method. On one hand, the TD3 method is suitable for continuous action spaces; on the other hand, it is based on a deterministic policy, enabling training using offline data, making it suitable for spacecraft rendezvous scenarios. Specifically, the TD3 method employs an Actor-Critic framework. The Actor is the policy network, which outputs the action to be executed based on the state at each step, updating the policy network by maximizing the value generated by policy execution. The Critic is the value network, used to evaluate the value generated by policy execution at each step, and updated using the Temporal Difference (TD) algorithm.

[0084] The specific process of policy training is as follows: The upper-level planner observes the initial state. The policy network generates virtual waypoints based on the initial state. The lower-level planner generates a trajectory; after interacting with the training environment and executing the trajectory, the upper layer observes the reward of the value network. And the next step in the status Save the experience array The policy network and value network are updated. When a termination signal is detected, the environment is initialized and a new round of interaction and training begins. The above steps are repeated until policy training is completed and convergence is achieved.

[0085] S3 is based on a linear quadratic regulator to design a lower-level planner, which obtains the local optimal trajectory and control sequence.

[0086] Understandably, in S2, the collision constraint of the reward function is set as a hard constraint, so the virtual waypoint path planned by the upper-level planner can meet the obstacle avoidance requirements. However, boundary value and orbital dynamics constraints are not included in the training process. Therefore, in the lower-level trajectory optimization stage, the lower-level planner uses the virtual waypoint path as the initial condition, and under the premise of satisfying boundary value and orbital dynamics constraints, further generates locally optimal trajectory segments, and finally combines them into a globally optimized trajectory. Based on this, in this step, the lower-level planner obtains the optimal time-varying state feedback gain by solving the time-varying Riccati differential equation, thereby minimizing the quadratic form performance index in the finite time domain.

[0087] Specifically, the state deviation system equations are obtained based on the orbital dynamics equations in S1 as follows:

[0088]

[0089] Among them, the state deviation system is a linear system. For system matrix, Given an input matrix, and

[0090]

[0091] Finding the optimal control sequence within a finite time domain To minimize the performance index of the quadratic form of the spacecraft autonomous obstacle avoidance and rendezvous trajectory planning model in S1, according to Pontryagin's maximum principle, It has the following status feedback forms:

[0092]

[0093] in, It is a symmetric positive semi-definite matrix, given by the solution of the differential Riccati equation:

[0094]

[0095] Simultaneously satisfy:

[0096]

[0097] When solving the differential Riccati equation, it is transformed into vector form to adapt to the standard ordinary differential equation solver. Specifically, in this embodiment, the integration time interval is set as follows:

[0098] ,

[0099] Terminal configuration requirements:

[0100]

[0101] Using the fourth-order Runge-Kutta method, from the reverse time... Points to Obtain the solution Mapping the solution to real time Obtaining time-varying matrix functions .

[0102] The feedback gain matrix at each time step is calculated using the following formula. :

[0103]

[0104] Finally, based on the initial state and the stored feedback gain matrix, in Closed-loop simulation was performed to obtain the optimal trajectory segment. and control sequence .

[0105] S4 combines the optimal trajectory segments into a globally optimized trajectory, and controls the spacecraft through a control sequence to make the spacecraft move along the planned trajectory.

[0106] As described above, the hierarchical motion planning method for autonomous obstacle avoidance and rendezvous of spacecraft provided in this embodiment, which integrates reinforcement learning, simplifies the high-dimensional optimization problem into multiple low-dimensional sub-problems through the design of a two-layer planner. It takes into account both global exploration and local optimization, eliminates the need for collision detection, and improves the computational efficiency of the motion planning process. Furthermore, it combines reinforcement learning with optimal control, trains the upper-layer decision network to generate virtual waypoints to handle obstacle avoidance constraints through a dual-delay deep deterministic policy gradient method, and uses a linear quadratic regulator to generate optimal trajectory segments to handle orbital dynamics, path, and boundary constraints. While ensuring the overall optimality of the rendezvous trajectory, it significantly enhances the generalization ability to complex mission environments, ultimately improving the quality of the autonomous obstacle avoidance and rendezvous trajectory of spacecraft.

[0107] The effects of the present invention are further illustrated below through specific examples. Figure 3 This demonstrates a scenario of a spacecraft autonomously avoiding obstacles and safely rendezvous. The upper-level planner generates a collision-free geometric path. ,in , All are virtual waypoints; the lower-level planner outputs spatial intersection trajectories that satisfy boundary value, path, and orbital dynamics constraints.

[0108] The simulation parameters for the spacecraft's autonomous obstacle avoidance and safe rendezvous trajectory planning problem are designed as follows:

[0109] Table 1. Relevant parameters for spacecraft rendezvous scenarios

[0110] Table 2 Hyperparameters of the reinforcement learning algorithm training process

[0111] Table 3. Parameters related to the lower-level planner

[0112] A positive reward of 200 is awarded when the spacecraft reaches the target state. Simultaneously, parameters are set. , , , , , The planning success rate is defined as the percentage of spacecraft that successfully reach the target area in 50 random missions without any collisions. The cumulative reward is defined as the average reward obtained by the spacecraft in 50 random missions.

[0113] Figure 4 The success rate and cumulative reward curve for spacecraft space rendezvous trajectory planning are shown. Both the success rate and cumulative reward curve converge as the number of training steps increases, with the final success rate reaching approximately 95%. Figure 4 The curve showing the change in distance between the spacecraft and the space target is displayed. After 212 seconds of flight, the spacecraft successfully reached the vicinity of the target while meeting the velocity requirements relative to the target. Figure 5 The curve showing the spacecraft's velocity change relative to a space target is displayed. After 212 seconds of flight, the spacecraft simultaneously met the velocity and position requirements relative to the space target. Figure 6 The curves showing the change in distance between the spacecraft and two space obstacles are presented. During the 212-second flight time, the relative distance between the spacecraft and the two space obstacles remained greater than the safety threshold of 50 meters, meeting the collision avoidance requirements.

[0114] Based on the same inventive concept, this embodiment also provides an apparatus, including: a memory for storing a computer program; and a processor for executing the computer program to implement the relevant steps of the spacecraft autonomous obstacle avoidance and rendezvous hierarchical motion planning method fused with reinforcement learning as described above.

[0115] The processor may include one or more processing cores, such as a quad-core processor or an octa-core processor. The processor can be implemented using at least one hardware form of Digital Signal Processing (DSP), Field-Programmable Gate Array (FPGA), or Programmable Logic Array (PLA). The processor may also include a main processor and coprocessors. The main processor, also known as the Central Processing Unit (CPU), is used to process data in the wake-up state; the coprocessors are low-power processors used to process data in the standby state. In some embodiments, the processor may integrate a Graphics Processing Unit (GPU), which is responsible for rendering and drawing the content to be displayed on the screen. In some embodiments, the processor may also include an Artificial Intelligence (AI) processor, which handles computational operations related to machine learning.

[0116] The memory may include one or more computer-readable storage media, which may be non-transitory. The memory may also include high-speed random access memory and non-volatile memory, such as one or more disk storage devices or flash memory devices. In this embodiment, the memory is used to store at least the following computer program, which, after being loaded and executed by the processor, is capable of implementing the aforementioned software method steps. In addition, the resources stored in the memory may also include operating systems and data, and the storage method may be temporary or permanent storage. The operating system may include Windows, Unix, Linux, etc.

[0117] This embodiment also provides a computer-readable storage medium storing a computer program. When the computer program is executed by a processor, it implements the relevant steps of the spacecraft autonomous obstacle avoidance and rendezvous hierarchical motion planning method fused with reinforcement learning as described above.

[0118] The system, apparatus, computer-readable storage medium, etc. provided in this embodiment have the same inventive concept and beneficial effects as the aforementioned method, and will not be repeated here.

[0119] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0120] The above embodiments are merely illustrative of several implementation methods of this application, and their descriptions are relatively specific and detailed, but they should not be construed as limiting the scope of the application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.

Claims

1. A spacecraft autonomous rendezvous and docking layered motion planning method based on fusion of reinforcement learning, characterized in that, Comprising the following steps: S1, modeling the spacecraft autonomous obstacle avoidance rendezvous trajectory problem, and constructing a hierarchical planning framework including an upper planner and a lower planner; S2, designing the upper planner based on the double-delay deep deterministic policy gradient method, and obtaining virtual waypoints through the upper planner; S3, designing the lower planner based on the linear quadratic regulator, and obtaining a locally optimal trajectory and a control sequence through the lower planner; S4, combining the optimal trajectory segments output by the lower planner into a globally optimized trajectory, and controlling the spacecraft through the control sequence to make the spacecraft move according to the planned trajectory.

2. The method according to claim 1, wherein, Tracking spacecraft in S1 Geosynchronous orbit spacecraft The orbital dynamics equations describing the relative motion in the vicinity of a spacecraft are given by: ; in, In orbital coordinate system and The relative position vector, for orbital angular velocity, For controlling input.

3. The spacecraft autonomous rendezvous and proximity maneuver planning method based on fusion of reinforcement learning according to claim 2, wherein, S1 defines the spacecraft state vector The spacecraft autonomous obstacle avoidance rendezvous trajectory planning model is represented as: ; wherein, is the initial time instant, is the final time instant, denotes the convergence bias of the spacecraft from the initial state to the terminal state , , , are self-defined weight matrices, is the obstacle avoidance constraint, denotes the set of collision-free rendezvous trajectories, is the orbital dynamics constraint, and are path constraints, mainly limiting the mapping of the spacecraft motion path in time, and are boundary constraints, ensuring that the spacecraft can move from the initial state to the target state.

4. The spacecraft autonomous rendezvous and proximity maneuver planning method based on fusion of reinforcement learning according to claim 3, characterized in that, The upper planner design in S2 includes state space design, action space design and reward function design, the state space is a complete description of the environment by the upper planner at each time step in the iteration process; The action space is the specific decision behavior made by the upper planner at each time step; the reward function is used to provide feedback to the upper planner.

5. The spacecraft autonomous rendezvous and proximity maneuver planning method based on fusion of reinforcement learning according to claim 4, wherein, The reward function in S2 includes distance guidance reward, performance index reward, collision constraint reward and path constraint reward; the distance guidance reward is a dense reward oriented to the target position; the performance index reward is used to feedback a reward negatively related to the minimum performance index to the upper planner at each time step; the collision constraint reward corresponds to a hard constraint, when the constraint is violated, the upper planner receives a negative reward and ends the search; the path constraint reward is used to make the upper planner meet the path constraint condition in the exploration process.

6. The method of claim 4, wherein, The upper planner design in S2 includes policy training, and the policy training objects include policy network and value network, the policy network outputs the action to be executed according to the state of each step, and updates itself by maximizing the value generated by policy execution; the value network is used to evaluate the value generated by each step of policy execution, and updates itself through the time difference algorithm.

7. The method of claim 3, wherein, The lower planner in S3 obtains the optimal time-varying state feedback gain by solving the time-varying Riccati differential equation, minimizes the performance index in the form of quadratic form in the limited time domain, and performs closed-loop simulation based on the initial state and the stored feedback gain matrix within the preset time period to obtain the optimal trajectory segment and the control sequence.

8. The spacecraft autonomous rendezvous and proximity maneuver planning method of claim 7, wherein, S3 specifically includes the following sub-steps: Based on the orbit dynamics equation in S1, the state deviation system equation is: ; wherein the state deviation system is a linear system, is a system matrix, is an input matrix, and ; Finding optimal control sequences in finite time domain to minimize the performance index in a quadratic form of the spacecraft autonomous collision avoidance rendezvous trajectory planning model in S1, it is known from Pontryagin's maximum principle that, has the following state feedback form: ; where is a symmetric positive semi-definite matrix, given by the solution of the differential Riccati equation ; Meanwhile, the following conditions are met: ; Set the integral time interval: , ; Configure the terminal condition: ; Using the fourth order Runge-Kutta method, the solution is obtained by integrating from to the backward time , mapping the solution to the real time , and obtaining the time-varying matrix function ;​ The feedback gain matrix at each time instant is calculated by the following equation : ; Based on the initial state and the stored feedback gain matrix, a closed-loop simulation is performed in to obtain an optimal trajectory segment and control sequence .