An Autonomous Decision-Making Method for Spacecraft Multi-Space Debris Collision Avoidance Based on Near-End Strategy Optimization
By constructing a spacecraft-space debris orbital dynamics model and neural network decision-making, the problem of low computational efficiency in space debris collision avoidance was solved, enabling autonomous avoidance and real-time decision-making, thereby improving the avoidance success rate and energy utilization efficiency.
Patent Information
- Application Number
- CN202310103998.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-02-13
- Publication Date
- 2025-12-02
- Estimated Expiration
- 2043-02-13
AI Technical Summary
Existing space debris collision avoidance methods for spacecraft have low computational efficiency, cannot achieve autonomous avoidance, and traditional methods cannot meet the requirements for real-time obstacle avoidance decision-making.
An autonomous decision-making method for spacecraft collision avoidance based on near-end strategy optimization is adopted. This method involves constructing orbital dynamics models of spacecraft and space debris, designing a collision probability calculation module, generating space debris simulation parameters, and using neural networks to achieve online decision-making. The system is then trained using a Markov decision process model and a near-end strategy optimization algorithm to establish an autonomous decision-making system for spacecraft collision avoidance.
It effectively reduces the decision-making time for avoidance, improves the success rate of spacecraft in avoiding multiple space debris and energy utilization efficiency, and has real-time decision-making capabilities.
Smart Images

Figure CN116125811B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of spacecraft collision avoidance, and more specifically to an autonomous decision-making method for spacecraft multi-space debris collision avoidance based on near-end strategy optimization. Background Technology
[0002] With the rapid development of the global space industry, the number of satellite launches worldwide has been increasing year by year, with more than 30 countries and regions conducting launch missions. Entering the 21st century, driven by national military strategic security needs, satellite launches have become increasingly urgent and frequent. However, due to the limited resources of outer space, particularly near-Earth space and geostationary orbit, the amount of space debris near Earth has increased rapidly. These ineffective payloads severely pollute the space environment around Earth, having a wide-ranging and serious impact on the safe operation of spacecraft in orbit, satellite mission execution, and rocket launch windows. Existing research on space debris collision avoidance is mostly based on simplified relative kinematic models and uses offline mathematical optimization methods to derive optimal maneuvers. However, the solution speed of traditional Gaussian pseudospectral methods and genetic algorithms cannot meet the real-time obstacle avoidance decision-making requirements of spacecraft in orbit, and it is also difficult to provide instantaneous high thrust for spacecraft in engineering. Therefore, it is necessary to study the real-time autonomous obstacle avoidance maneuvering decision-making of spacecraft with limited thrust in orbit.
[0003] Therefore, designing an autonomous decision-making method for spacecraft multi-space debris collision avoidance based on near-end strategy optimization can achieve autonomous avoidance and effectively reduce maneuver decision time and optimize maneuver energy consumption. Summary of the Invention
[0004] The purpose of this invention is to provide an autonomous decision-making method for spacecraft collision avoidance of multiple space debris based on near-end strategy optimization. This invention solves the problems of low computational efficiency and inability to autonomously avoid space debris in existing spacecraft space debris avoidance methods. This invention mainly achieves offline training by constructing orbital dynamics models of spacecraft and space debris, designing collision probability calculation modules, and generating space debris simulation parameters, and uses neural networks to achieve online decision-making. Through simulation case verification, it is proved that this method uses fewer spacecraft computing resources, effectively reduces avoidance decision time, has real-time decision-making capability, and improves the success rate of spacecraft avoidance of multiple space debris and energy utilization efficiency.
[0005] This reduces the time spent generating optimal evasive maneuvers and improves the energy efficiency of aircraft.
[0006] The present invention adopts the following technical solution:
[0007] An autonomous decision-making method for spacecraft multi-space debris collision avoidance based on near-end strategy optimization, the method comprising the following steps:
[0008] Step 1: Construct the spacecraft's space dynamics model in the geocentric inertial coordinate system as follows:
[0009]
[0010] Where r is the spacecraft's spatial position vector; μ is the Earth's gravitational constant, with a value of 3.986 × 10⁻⁶. 5 km 3 / s 2 ;f t The engine thrust acceleration vector is used in this invention, which employs pulse maneuvering, and the total maneuvering amount is set to F. max ;f p It is the J2 perturbation acceleration vector acting on the spacecraft;
[0011] Step 2: Construct a mathematical model of the collision probability based on the orbital dynamics of the spacecraft and space debris;
[0012] Step 3: Generating space debris simulation parameters based on collision time;
[0013] Step 4: Construct a mathematical model of the reward function for the collision probability and energy loss;
[0014] Step 5: Establish a spacecraft collision avoidance autonomous decision-making training system based on the near-end strategy optimization algorithm;
[0015] The spacecraft collision avoidance autonomous decision-making training system selects the optimal action in the current state and enables the spacecraft to successfully avoid space debris in the best state through continuous decision-making.
[0016] Step Six: Apply the models established in Steps One, Two, Three, and Four to the system in Step Five to train the spacecraft collision avoidance autonomous decision-making system offline;
[0017] Step 7: Apply the spacecraft collision avoidance autonomous decision-making system trained in Step 6 to multiple space debris collision avoidance scenarios for online spacecraft to obtain optimized maneuver trajectories for successful autonomous avoidance.
[0018] Furthermore, step two involves constructing a mathematical model for the collision probability, the specific process of which is as follows:
[0019] At each time step, obtain the position and velocity of the spacecraft and space debris in the geocentric coordinate system at the current moment;
[0020] The closest moment between the spacecraft and space debris, as well as its position and velocity at the closest moment, are obtained by propagating forward from the orbital dynamics equations.
[0021] The position and velocity of the spacecraft at the closest moment to the space debris are converted into relative position and relative velocity in a relative coordinate system, and the joint position error covariance of the two is calculated.
[0022] Using the first term of the infinite series of the two-dimensional Gaussian probability density function as an approximation of the probability integral, the mathematical model of the collision probability P at the closest moment is calculated according to the following formula. c ;
[0023]
[0024] Where, μ x and μ y These represent the x-axis and y-axis coordinates of the spacecraft and space debris in the encounter coordinate system, respectively, σ x and σ y These represent the standard deviations of the joint positional errors of the spacecraft and space debris along the x and y axes in the encounter coordinate system, respectively, r. A It is the sum of the radii of the spacecraft and space debris.
[0025] Furthermore, the specific process of generating space debris simulation parameters based on the collision time in step three is as follows:
[0026] The space debris collision time t is obtained by orbital propagation over a certain period of time based on the spacecraft's initial state. c ;
[0027] Based on the space debris collision time t c The position of the spacecraft at time R s and speed V s Add a certain random perturbation R ε and V ε ;
[0028] Based on this, a random orbital inclination angle is selected. To obtain the final location R′ of the space debris d and speed V d ′;
[0029] Based on the forward propagation of space debris t c The initial position R of the space debris is obtained in seconds. d and speed V d .
[0030] Furthermore, the mathematical model for the reward function based on collision probability and energy loss in step four is as follows:
[0031]
[0032] Where, r p P is the reward value for the collision probability. sumThe total collision probability of multiple space debris is calculated using the following formula: P i r represents the collision probability of a single space debris. c As a reward for energy loss, F max F represents the total energy value. ac F represents the cumulative energy consumption value. sc F represents the energy consumption per maneuver. smax r represents the maximum energy consumption for a single maneuver. s For step size reward; r t As a conditional reward for the terminal, t step For environmental steps, coll flag This is a collision occurrence flag.
[0033] Furthermore, step five, establishing a spacecraft collision avoidance autonomous decision-making training system, involves the following specific process:
[0034] 501. The spacecraft collision avoidance decision-making process is modeled as a Markov decision process model, which includes: a state set, an action set, a state transition equation, a reward function mathematical model, and a discount factor;
[0035] The state set consists of twenty-six variables, including the relative three-dimensional position coordinates and three-dimensional velocity values of a spacecraft and three space debris generated in the geocentric coordinate system through the spacecraft's space dynamics model, and the spacecraft's remaining fuel value.
[0036] The relative distance, collision probability, and total collision probability of the spacecraft and the three space debris at the closest moment are obtained through the collision probability mathematical model.
[0037] The action set consists of three variables, including the spacecraft's pulse maneuver values in the x-direction, y-direction, and z-direction coordinate systems.
[0038] use Calculate the total maneuver loss in all three directions during a single maneuver;
[0039] The state transition equation is based on the spacecraft's space dynamics model, meaning that when an action is input, the state will transition to the next state with 100% probability according to the orbital dynamics equation.
[0040] The mathematical model of the reward function, wherein the total reward value is composed of the collision probability reward value r p Energy loss reward value r c Step size reward value r s Terminal conditional reward value r t composition;
[0041] The discount factor is set to 0.95;
[0042] 502. A spacecraft collision avoidance autonomous decision-making training system is established by training the spacecraft collision avoidance model using a near-end strategy optimization algorithm; wherein:
[0043] The spacecraft collision avoidance autonomous decision-making training system consists of a Critic network and an Actor network. The Actor network is used to output the spacecraft's maneuver values, and the Critic network is used to evaluate the quality of the current state. The Actor network and the Critic network continuously interact with the simulation environment composed of the first four steps, collect experience samples, and further train and update the parameters of the Actor network and the Critic network through experience samples.
[0044] The spacecraft collision avoidance autonomous decision-making training system first initializes the parameters of the Actor network and Critic network, and initializes the experience pool space in the initial stage of training, where each set of data D in the experience pool... t ={s t ,s t+1 ,a t ,r t} represents the current state s t New state s t+1 Current mobility value a t and the current reward value r t ;
[0045] The spacecraft and space debris are initialized to a state s0, and this state is input into the Actor network and the Critic network. The Actor network outputs a maneuver value a0 based on the input state, and the Critic network outputs an evaluation value based on the input state.
[0046] The spacecraft collision avoidance autonomous decision training system inputs the maneuver value output by the Actor network into the collision probability mathematical model to obtain a new state, and obtains the reward value r0 of the maneuver value through the reward function mathematical model in step four.
[0047] The experience pool stores the above data;
[0048] The spacecraft collision avoidance autonomous decision-making training system further determines whether the new state has reached the terminal state, namely, the three states of collision, energy depletion, and simulation round end. If the terminal state has not been reached, the Actor network and Critic network continue to interact with the environment. If the terminal state has been reached, the state of the spacecraft and space debris needs to be reinitialized.
[0049] The spacecraft collision avoidance autonomous decision-making training system determines the number of experience pools. If the number of experience pools is reached, the Actor network and Critic network are updated through an algorithm; otherwise, the system continues to collect data.
[0050] During training, the Actor and Critic networks are updated according to the proximal policy optimization algorithm;
[0051] Clear the experience pool after updating the Actor and Critic networks;
[0052] The system determines whether the maximum number of training rounds has been reached. If it has, training stops; otherwise, training continues.
[0053] Furthermore, both the Actor and Critic networks employ fully connected neural network models:
[0054] The Critic network is designed as a fully connected neural network. The number of nodes in the input layer equals the number of variables in the state set, i.e., the input variables are twenty-six state variables. The number of nodes in the output layer is an evaluation value used to judge the quality of the current state. The number of hidden layers and nodes can be defined by the user. Here, three hidden layers are designed, with 256, 128, and 128 nodes in each layer, respectively. The ReLU function is used as the activation function of the network, and the Adam optimizer is used to train the neural network.
[0055] The Actor network is designed as a fully connected neural network with twenty-six state variables as input variables. It has two hidden layers with 256 and 128 nodes respectively. The output is the mean and standard deviation of the impulse maneuvers in three directions. The actual impulse maneuver values can be obtained by probability sampling. The ReLU function is used as the activation function of the network, and the Adam optimizer is used to train the neural network.
[0056] The beneficial effects of this invention are:
[0057] This invention addresses the problems of low computational efficiency and inability to autonomously avoid space debris in existing spacecraft collision avoidance methods. The aim of this invention is to provide a collision avoidance method capable of offline training and online decision-making. Offline training is achieved through steps such as constructing orbital dynamics models of the spacecraft and space debris, designing a collision probability calculation module, and generating space debris simulation parameters. Online decision-making is then implemented using a neural network. This method uses fewer spacecraft computational resources, effectively reducing avoidance decision-making time, providing real-time decision-making capabilities, and improving the success rate and energy efficiency of spacecraft avoiding multiple pieces of space debris. Attached Figure Description
[0058] Figure 1 This is a flowchart for calculating collision probability.
[0059] Figure 2 This is a flowchart of the spacecraft collision avoidance autonomous decision-making training process.
[0060] Figure 3This is a training result diagram of the spacecraft collision avoidance autonomous decision-making system.
[0061] Figure 4 This is a simulation case's motor result diagram.
[0062] Figure 5 This is a graph showing the change in the miss distance in a simulation case.
[0063] Figure 6 This is a graph showing the changes in collision probability in a simulation case.
[0064] Figure 7 This is a graph showing the relationship between the incremental maneuver value and the reward value in 100 simulations. Detailed Implementation
[0065] The embodiments of the present invention will be further described below with reference to the accompanying drawings.
[0066] Specific Implementation Method 1: This invention provides an autonomous decision-making method for spacecraft multi-space debris collision avoidance based on near-end strategy optimization, including the following steps:
[0067] Step 1: Construct orbital dynamics models of spacecraft and space debris;
[0068] Step 2: Design the collision probability calculation module;
[0069] Step 3: Generating space debris simulation parameters based on collision time;
[0070] Step 4: Design a reward function based on collision probability and energy loss;
[0071] Step 5: Establish a spacecraft collision avoidance autonomous decision-making training system based on the near-end strategy optimization algorithm;
[0072] Step Six: Apply the models established in Steps One, Two, Three, and Four to the system in Step Five to train the spacecraft collision avoidance autonomous decision-making model offline.
[0073] Step 7: Apply the spacecraft collision avoidance autonomous decision-making model trained in Step 6 to multiple space debris collision avoidance scenarios for online spacecraft to obtain optimized maneuver trajectories for successful autonomous avoidance.
[0074] This invention mainly consists of two stages: offline training and online decision-making. A large amount of diverse simulation data is generated using a dynamic model to train the near-end policy optimization algorithm offline, ultimately resulting in a trained neural network model. This trained neural network model can then be used to implement autonomous collision avoidance decisions online, effectively improving the collision avoidance success rate.
[0075] Specific Implementation Method Two: This implementation method differs from Specific Implementation Method One in that: the specific process of constructing the orbital dynamics model of the spacecraft and space debris in step one is as follows:
[0076] First, this method establishes the orbital dynamics equations for spacecraft and space debris based on the J2000 geocentric inertial coordinate system. This method can directly use spatial position and velocity to describe the on-orbit state of space objects, and can be performed more intuitively and conveniently in the self-learning system, thereby improving the system's solution speed. In the geocentric inertial coordinate system, the spacecraft's space dynamics model is as follows:
[0077]
[0078] Where r is the spacecraft's spatial position vector; μ is the Earth's gravitational constant, with a value of 3.986 × 10⁻⁶. 5 km 3 / s 2 ;f t The engine thrust acceleration vector is used in this invention, which employs pulse maneuvering, and the total maneuvering amount is set to F. max ;f p It is the J2 perturbation acceleration vector acting on the spacecraft, and its specific expression is:
[0079]
[0080] Where x, y, and z are the components of the spacecraft's position vector along the coordinate axes of the J2000 coordinate system, respectively, and f px f pz f pz The components of the perturbation acceleration along the three-dimensional coordinate axes, R e The radius of the Earth is 6378.137 km, and J2 = 1.08262668 × 10 -3 .
[0081] Since the orbital altitude of space debris near a spacecraft is approximately the same as that of the spacecraft, the orbital kinematic equations of the space debris are consistent with the orbital dynamic equations of the spacecraft.
[0082] Specific Implementation Method Three: This implementation method differs from Specific Implementation Method Two in that: Step Two involves designing a collision probability calculation module, the specific process of which is as follows:
[0083] Regarding satellite collision risk assessment, the widely accepted approach is to evaluate and analyze the collision risk between satellites and space targets using collision probability estimation. The expression for the three-dimensional Gaussian probability density function is:
[0084]
[0085] Among them, t TCA C represents the moment when two spatial objects are closest. rr (tTCA S(t) represents the joint positional error covariance of the two. TCA Let be the relative positions of the two. Integrating this equation over the spatial region traversed by the joint envelope sphere, we can obtain the collision probability at the moment of encounter as:
[0086]
[0087] To address the time-consuming integration of the above formula, considering the relatively high velocity of objects in space, we can treat them as linear relative motion. Based on this assumption, the problem of calculating the collision probability can be transformed into calculating the integral of the two-dimensional probability density function within a circular domain. The collision probability calculation formula can then be simplified as follows:
[0088]
[0089] Where, μ x and μ y These represent the x-axis and y-axis coordinates of the spacecraft and space debris in the encounter coordinate system, respectively, σ x and σ y These represent the standard deviations of the joint positional errors of the spacecraft and space debris along the x and y axes in the encounter coordinate system, respectively, r. A It is the sum of the radii of the spacecraft and space debris.
[0090] Based on previous work, the above formula can be approximated by taking the first term of the infinite series as the probability integral, and the specific expression is as follows:
[0091]
[0092] The above is the formula for calculating collision probability. The flowchart for using collision probability as a collision warning is as follows: Figure 1 As shown.
[0093] In each time step, the system first obtains the position and velocity of the spacecraft and space debris in the geocentric coordinate system at the current moment. Then, it propagates forward according to the orbital dynamics equation to obtain the closest time (TCA) between the spacecraft and the space debris, as well as the position and velocity at the closest time. Then, it converts these into relative position and relative velocity in the relative coordinate system (NTW) and calculates the joint position error covariance of the two. Finally, it calculates the collision probability at the closest time according to equation (6).
[0094] Based on existing research, the collision avoidance threshold for spacecraft is generally divided into three cases: when the collision probability reaches 10... -4 When the probability of collision reaches the danger threshold, the spacecraft must perform corresponding evasive maneuvers; when the probability of collision reaches 10... -5When the proximity threshold for danger is reached, a collision warning is required, along with further tracking and analysis of the hazardous target. After implementing evasive maneuvering strategies, the probability of a collision between the spacecraft and space debris at the closest possible moment must be less than 10. -7 This method will use 10 -4 As a threshold for the probability of a dangerous collision, when the calculated collision probability exceeds this threshold, it indicates that evasive maneuvers are required.
[0095] Specific Implementation Method Four: This implementation method differs from Specific Implementation Method Three in that step three is based on the generation of space debris simulation parameters at the collision time. The specific process is as follows:
[0096] This method employs reinforcement learning as an optimization tool for evasive maneuvers, thus requiring offline training of the reinforcement learning algorithm. To generate a large number of simulation scenarios, this section designs a corresponding method for generating space debris parameters.
[0097] In this method, the spacecraft's orbital parameters are fixed during training, while the other three space debris fragments are randomly generated in each training round. At the initial moment, the spacecraft's state is obtained, and its orbit is propagated over a certain period. The system then randomly selects a time as the collision time t. c Then according to t c The position of the spacecraft at time R s and speed V s Add a certain random perturbation R ε and V ε Based on this, a random orbital inclination angle is generated. The location R′ of the space debris is obtained. d and speed V d Finally, the space debris will propagate forward. c The initial position and velocity R of the space debris are obtained after seconds. d and speed V d .
[0098] This approach generates different spatial fragments in each training round, ensuring the diversity of training samples and improving the applicability of reinforcement learning. Compared to traditional optimization algorithms, data-based reinforcement learning achieves better generalization performance.
[0099] Specific Implementation Method Five: This implementation method differs from Specific Implementation Method Four in that step four is based on the design of a reward function that considers collision probability and energy loss. The specific process is as follows:
[0100] As a spacecraft collision avoidance problem, the primary goal is to successfully avoid space debris to ensure the spacecraft's normal on-orbit service. Secondly, it's crucial to minimize energy loss (i.e., reduce velocity increments) while successfully avoiding collisions. Therefore, collision probability and energy consumption are used as two optimization metrics, and a reward function is designed for them. The specific design results are as follows:
[0101]
[0102] Where, r p P is the reward value for the collision probability. sum The total collision probability of multiple space debris is calculated using the following formula: P i r represents the collision probability of a single space debris. c As a reward for energy loss, F max F represents the total energy value. ac F represents the cumulative energy consumption value. sc F represents the energy consumption per maneuver. smax r represents the maximum energy consumption for a single maneuver. s For step size reward; r t As a conditional reward for the terminal, t step For environmental steps, coll flag This is a collision occurrence flag.
[0103] The reward function defined above fully considers task avoidance and energy optimization metrics, while also facilitating the learning of reinforcement learning algorithms. Specifically, to enable the agent to learn to avoid spatial debris, r is set... p Reward value. When P sum >10 -4 At that time, r p For negative rewards; when P sum <10 -4 At that time, r p Positive rewards are used to encourage agents to learn how to correctly avoid space debris. While avoiding space debris, energy consumption also needs to be optimized; therefore, r is set... c The reward value includes cumulative energy loss and single maneuver loss. When P sum >10 -4 At that time, r c The overall negative reward value is relatively small, encouraging the agent to boldly take evasive maneuvers, but when P sum <10 -4 At this point, the agent should not perform any additional maneuvers to reduce energy consumption; therefore, in this situation, r c This will result in a significant negative reward. Furthermore, to encourage the agent to continuously run towards the simulation terminal, a time step reward r is specifically set. sThe longer the agent runs, the greater the reward. t The system's terminal condition reward value is determined when the spacecraft collides with space debris or runs out of energy, resulting in a negative reward; when the spacecraft successfully avoids space debris and reaches the simulation terminal moment, a larger positive reward is given.
[0104] Specific Implementation Method Six: This implementation method differs from Specific Implementation Method Five in that step five establishes a spacecraft collision avoidance autonomous decision-making training system based on reinforcement learning theory. The specific process is as follows:
[0105] The most representative algorithm architecture in reinforcement learning is the Actor-Critic algorithm. Its core is that an "actor" generates an action policy, and then a "critic" evaluates the current policy, guiding adjustments to the action policy. This framework has spawned numerous algorithms, including the Trust Region Policy Optimization (TRPO), Proximal Policy Optimization (PPO), and Deep Deterministic Policy Gradient (DDPG). Among them, the PPO algorithm stands out due to its ease of implementation and superior performance; therefore, this method adopts PPO as the spacecraft collision avoidance algorithm.
[0106] The purpose of the spacecraft collision avoidance autonomous decision-making training system is to select the optimal action in the current state and enable the spacecraft to successfully avoid space debris in the best state through continuous decision-making. This process satisfies the characteristics of stochastic sequential decision-making.
[0107] The spacecraft collision avoidance decision-making process is modeled as a Markov decision process model, which includes: a state set, an action set, a state transition equation, a reward function, and a discount factor.
[0108] The state set consists of twenty-six variables, including the relative three-dimensional position coordinates and three-dimensional velocity values of a spacecraft and three space debris generated in the geocentric coordinate system by the spacecraft's space dynamics model, the remaining fuel value of the spacecraft, the relative distance, collision probability, and total collision probability of the spacecraft and the three space debris at the closest moment obtained by the collision probability mathematical model.
[0109] The action set consists of three variables: the spacecraft's pulse maneuver values in the x-direction, y-direction, and z-direction coordinates. (Using...) Calculate the total maneuver loss in all three directions during a single maneuver.
[0110] The state transition equation adopts the orbital dynamics equation (1), that is, when an action is input, the state will transition to the next state with 100% probability according to the orbital dynamics equation.
[0111] The reward function is the same as the reward function designed in step four, where the total reward value is composed of the collision probability reward value r. p Energy loss reward value r c Step size reward value r s Terminal conditional reward value r t composition.
[0112] The discount factor is set to 0.95.
[0113] The above five elements constitute the entire decision update process of the training system.
[0114] The spacecraft collision avoidance autonomous decision-making training system uses the PPO algorithm as its training algorithm, which consists of a Critic network and an Actor network. The Actor network outputs the spacecraft's maneuver values, while the Critic network evaluates the quality of the current state. The Actor and Critic networks continuously interact with the simulation environment comprised of the first four steps, collecting experience samples. These experience samples are then used to further train and update the parameters of the Actor and Critic networks.
[0115] The flowchart of the spacecraft collision avoidance autonomous decision-making training system is as follows: Figure 2 As shown, in the early stages of training, the parameters of the Actor network and Critic network are first initialized, and the experience pool space is initialized, where each set of data D in the experience pool is... t ={s t ,s t+1 ,a t ,r t} represents the current state s t New state s t+1 Current mobility value a t and the current reward value r t ;
[0116] In each simulation round, the state s0 of the spacecraft and space debris is first initialized and then input into the Actor network and the Critic network. The Actor network outputs a maneuver value a0 based on the input state, and the Critic network outputs an evaluation value based on the input state.
[0117] The system inputs the maneuver value output by the Actor network into the orbital dynamics equation to obtain a new state, and obtains the reward value r0 of the maneuver value through the reward function in step four;
[0118] Store this group of data in the experience pool;
[0119] The system further determines whether the new state has reached the terminal state, namely, one of three states: collision, energy depletion, or end of simulation round. If the terminal state has not been reached, the Actor network and Critic network continue to interact with the environment; if the terminal state has been reached, the states of the spacecraft and space debris need to be reinitialized.
[0120] The system determines the size of the experience pool. If the size of the experience pool is reached, the Actor network and Critic network are updated using the PPO algorithm; otherwise, the system continues to collect data.
[0121] The training objective of the autonomous decision-making training system is to obtain the Actor network parameters corresponding to the maximum total expected return J(θ) and the minimum expected error L(θ). Here, parameter θ is the policy function approximated by the Actor network, and the expected error is used to update the Critic network parameters. Therefore, the PPO algorithm updates the Actor network and Critic network as follows:
[0122] This method defines the loss function of the Actor network as L. actor (θ), its specific expression is:
[0123]
[0124] in, π represents the ratio of the probabilities of the new and old strategies. θ (a t |s t ) is represented by parameter θ in state s t Choose action a under the condition t The probability of; It is the parameter θ old Indicated in state s t Choose action a under the condition t The probability, θ old The historical values of θ are passed to θ after a certain number of training steps. old . Let be the dominance function, which represents the current action 'a'. t Compared to strategy π θ The advantages.
[0125] L critic (θ) represents the state s t The difference between the corresponding true value function and the estimated value is used to update the Critic network parameters. Because s t The corresponding value function V π (s t If the value is unknown, it is generally estimated using a neural network. It can be expressed as a function of the neural network weight parameters θ, i.e. For a given trajectory, state s t The true value function at a given location can be estimated using the following formula:
[0126]
[0127] The loss function of the value function can then be expressed as:
[0128]
[0129] The maximum total expected return J(θ) and the minimum expected error L(θ) are expressed as follows:
[0130]
[0131] The above is the derivation of the loss functions for updating the Actor and Critic networks. During training, gradient descent is used to update the Actor and Critic networks.
[0132] Clear the experience pool after updating the Actor and Critic networks;
[0133] The system determines whether the maximum number of training rounds has been reached. If it has, training stops; otherwise, training continues.
[0134] Specific Implementation Method Seven: This implementation method differs from Specific Implementation Method Six in that both the Actor network and the Critic network employ fully connected neural network models.
[0135] For the Critic network, a fully connected neural network is designed. The number of nodes in the input layer equals the number of variables in the state set, i.e., the input variables are twenty-six state variables. The number of nodes in the output layer is an evaluation value used to judge the quality of the current state. The number of hidden layers and nodes can be defined by the user; here, three hidden layers are designed, with 256, 128, and 128 nodes in each layer, respectively. The ReLU function is used as the activation function of the network, and the Adam optimizer is used to train the neural network.
[0136] For the Actor network, a fully connected neural network is designed with 26 state variables as input variables and two hidden layers with 256 and 128 nodes respectively. The output is the mean and standard deviation of the impulse maneuvers in three directions. The actual impulse maneuver values can be obtained by probability sampling. The ReLU function is used as the activation function of the network, and the Adam optimizer is used to train the neural network.
[0137] To facilitate the description of the training process in step six, the following simulation scenario is designed:
[0138] This simulation scenario considers the collision avoidance problem of three space debris. At the initial moment, the spacecraft's orbital parameters are set as [6868.76, -1801.98, -3153.79, 0.20, 7.62, 0.13] (position km, velocity km / s). According to step four, the orbital parameters of the space debris can be generated by setting the collision time. Therefore, the relevant parameters for generating the three space debris are shown in Table 1.
[0139] Table 1 Range of parameters related to space debris generation
[0140] serial number Collision time track inclination Space Fragment 1 6500s~6600s 35°~55° Space Fragment 2 7000s~7100s 50°~70° Space Fragment 3 7500s~7600s 260°~280°
[0141] In addition, set position R s and speed V s The standard deviations were 0.00005 and 0.00001, respectively.
[0142] The spacecraft radius is set to 100m, and the radii of the three space debris fragments are 0.1m. The maximum maneuver increment per round for the spacecraft is set to 1m / s, and the maximum maneuver increment in a single direction during a single maneuver is set to 0.03m / s. The position uncertainty covariance of the spacecraft and space debris on the x, y, and z axes is set to [200, 200, 200, 300, 300, 300] (in meters). The kinematic simulation step size is set to 200s, and the simulation cycle per round is 9000s. The hyperparameters for training the reinforcement learning algorithm are shown in Table 2.
[0143] Table 2 Hyperparameter settings for reinforcement learning algorithms
[0144]
[0145]
[0146] The beneficial effects of the present invention are verified using the following embodiments:
[0147] Example 1:
[0148] 1) Experimental Environment
[0149] The simulation experimental environment described in step five is adopted.
[0150] 2) Analysis of experimental results
[0151] The training obtained through step six of this invention is as follows: Figure 3 The average return curve shown is composed of... Figure 3 As can be seen, the training algorithm proposed in this invention begins to converge around 2000 rounds;
[0152] To visually demonstrate the training results, a simulation case was selected. The orbital parameters for space debris 1 were set to [5456.76, 883.44, -4080.35, 1.36, 6.90, 3.01] (position km, velocity km / s), for space debris 2 to [3458.48, 288.65, -5930.48, 0.36, 7.58, 0.58], and for space debris 3 to [1486.74, -2978.73, -6015.82, -3.09, 5.90, 3.67]. Figure 4 It can be seen that the spacecraft successfully avoided three pieces of space debris by performing four maneuvers, with a total maneuver increment of 0.2245 m / s. Figure 4 and Figure 5 The data shows the changes in the miss distance and collision probability after the spacecraft maneuvers. After four maneuvers, the collision probability dropped to 10%. -7 Meanwhile, the simulated maneuver decision-making time was 1.2s, demonstrating the effectiveness and speed of the invention.
[0153] To verify the robustness of the algorithm, the trained network was subjected to 100 evasion maneuver simulations, and the results were as follows: Figure 7 The scatter plot shows the avoidance maneuver increment and reward value for 100 simulation tests. As can be seen from the scatter plot, the method described in this invention achieves a 100% success rate in avoiding collisions. The average maneuver increment for 100 simulation tests is 0.2186 m / s, indicating that the algorithm has good robustness and can significantly improve the collision avoidance capability of spacecraft.
[0154] The above specific embodiments are used to explain and illustrate the present invention, but not to limit the present invention. Any modifications and changes made to the present invention within the spirit and scope of the claims shall fall within the protection scope of the present invention.
Claims
1. An autonomous decision-making method for spacecraft multi-space debris collision avoidance based on near-end strategy optimization, characterized in that, The method includes the following steps: Step 1: Construct the spacecraft's space dynamics model in the geocentric inertial coordinate system as follows: Where r is the spacecraft's spatial position vector; μ is the Earth's gravitational constant, with a value of 3.986 × 10⁻⁶. 5 km 3 / s 2 ;f t The engine thrust acceleration vector is used in this invention, which employs pulse maneuvering, and the total maneuvering amount is set to F. max ;f p It is the J2 perturbation acceleration vector acting on the spacecraft; Step 2: Construct a mathematical model of the collision probability based on the orbital dynamics of the spacecraft and space debris; Step 3: Generating space debris simulation parameters based on collision time; Step 4: Construct a mathematical model of the reward function for the collision probability and energy loss; Step 5: Establish a spacecraft collision avoidance autonomous decision-making training system based on the near-end strategy optimization algorithm; The spacecraft collision avoidance autonomous decision-making training system selects the optimal action in the current state and enables the spacecraft to successfully avoid space debris in the best state through continuous decision-making. Step Six: Apply the models established in Steps One, Two, Three, and Four to the system in Step Five to train the spacecraft collision avoidance autonomous decision-making system offline; Step 7: Apply the spacecraft collision avoidance autonomous decision-making system trained in Step 6 to multiple space debris collision avoidance scenarios for online spacecraft to obtain optimized maneuver trajectories for successful autonomous avoidance.
2. The autonomous decision-making method for spacecraft multi-space debris collision avoidance based on near-end strategy optimization according to claim 1, characterized in that, The specific process of constructing the collision probability mathematical model in step two is as follows: At each time step, obtain the position and velocity of the spacecraft and space debris in the geocentric coordinate system at the current moment; The closest moment between the spacecraft and space debris, as well as its position and velocity at the closest moment, are obtained by propagating forward from the orbital dynamics equations. The position and velocity of the spacecraft at the closest moment to the space debris are converted into relative position and relative velocity in a relative coordinate system, and the joint position error covariance of the two is calculated. Using the first term of the infinite series of the two-dimensional Gaussian probability density function as an approximation of the probability integral, the mathematical model of the collision probability P at the closest moment is calculated according to the following formula. c ; Where, μ x and μ y These represent the x-axis and y-axis coordinates of the spacecraft and space debris in the encounter coordinate system, respectively, σ x and σ y These represent the standard deviations of the joint positional errors of the spacecraft and space debris along the x and y axes in the encounter coordinate system, respectively, r. A It is the sum of the radii of the spacecraft and space debris.
3. The autonomous decision-making method for spacecraft multi-space debris collision avoidance based on near-end strategy optimization according to claim 1, characterized in that, The specific process of generating space debris simulation parameters based on collision time in step three is as follows: The space debris collision time t is obtained by orbital propagation over a certain period of time based on the spacecraft's initial state. c ; Based on the space debris collision time t c The position of the spacecraft at time R s and speed V s Add a certain random perturbation R ε and V ε ; Based on this, a random orbital inclination angle is selected. To obtain the final location R′ of the space debris d and speed V d ′; The initial position R of the space debris is obtained based on the time it returned to the initial position. d and speed V d .
4. The autonomous decision-making method for spacecraft multi-space debris collision avoidance based on near-end strategy optimization according to claim 1, characterized in that, The mathematical model for the reward function in step four, based on collision probability and energy loss, is as follows: Where, r p P is the reward value for the collision probability. sum Let be the total collision probability of multiple space debris, calculated using the following formula: P i r represents the collision probability of a single space debris. c As a reward for energy loss, F max F represents the total energy value. ac F represents the cumulative energy consumption value. sc F represents the energy consumption per maneuver. smax r represents the maximum energy consumption for a single maneuver. s For step size reward; r t As a conditional reward for the terminal, t step For environmental steps, coll flag This is a collision occurrence flag.
5. The autonomous decision-making method for spacecraft multi-space debris collision avoidance based on near-end strategy optimization according to claim 1, characterized in that, Step five, establishing the spacecraft collision avoidance autonomous decision-making training system, involves the following process:
501. The spacecraft collision avoidance decision-making process is modeled as a Markov decision process model, which includes: a state set, an action set, a state transition equation, a reward function mathematical model, and a discount factor; The state set consists of twenty-six variables, including the relative three-dimensional position coordinates and three-dimensional velocity values of a spacecraft and three space debris generated in the geocentric coordinate system through the spacecraft's space dynamics model, and the spacecraft's remaining fuel value. The relative distance, collision probability, and total collision probability of the spacecraft and the three space debris at the closest moment are obtained through the collision probability mathematical model. The action set consists of three variables, including the spacecraft's pulse maneuver values in the x-direction, y-direction, and z-direction coordinate systems. use Calculate the total maneuver loss in all three directions during a single maneuver; The state transition equation is based on the spacecraft's space dynamics model, meaning that when an action is input, the state will transition to the next state with 100% probability according to the orbital dynamics equation. The mathematical model of the reward function, wherein the total reward value is composed of the collision probability reward value r p Energy loss reward value r c Step size reward value r s Terminal conditional reward value r t composition; The discount factor is set to 0.95; 502. A spacecraft collision avoidance autonomous decision-making training system is established by training the spacecraft collision avoidance model using a near-end strategy optimization algorithm; wherein: The spacecraft collision avoidance autonomous decision-making training system consists of a Critic network and an Actor network. The Actor network is used to output the spacecraft's maneuver values, and the Critic network is used to evaluate the quality of the current state. The Actor network and the Critic network continuously interact with the simulation environment composed of the first four steps, collect experience samples, and further train and update the parameters of the Actor network and the Critic network through experience samples. The spacecraft collision avoidance autonomous decision-making training system first initializes the parameters of the Actor network and Critic network, and initializes the experience pool space in the initial stage of training, where each set of data D in the experience pool... t ={s t ,s t+1 ,a t ,r t } represents the current state s t New state s t+1 Current mobility value a t and the current reward value r t ; The spacecraft and space debris are initialized to a state s0, and this state is input into the Actor network and the Critic network. The Actor network outputs a maneuver value a0 based on the input state, and the Critic network outputs an evaluation value based on the input state. The spacecraft collision avoidance autonomous decision training system inputs the maneuver value output by the Actor network into the collision probability mathematical model to obtain a new state, and obtains the reward value r0 of the maneuver value through the reward function mathematical model in step four. The experience pool stores the above data; The spacecraft collision avoidance autonomous decision-making training system further determines whether the new state has reached the terminal state, namely, the three states of collision, energy depletion, and simulation round end. If the terminal state has not been reached, the Actor network and Critic network continue to interact with the environment. If the terminal state has been reached, the state of the spacecraft and space debris needs to be reinitialized. The spacecraft collision avoidance autonomous decision-making training system determines the number of experience pools. If the number of experience pools is reached, the Actor network and Critic network are updated through an algorithm; otherwise, the system continues to collect data. During training, the Actor and Critic networks are updated according to the proximal policy optimization algorithm; After updating the Actor and Critic networks, the experience pool was cleared. The system determines whether the maximum number of training rounds has been reached. If it has, training stops; otherwise, training continues.
6. The autonomous decision-making method for spacecraft multi-space debris collision avoidance based on near-end strategy optimization according to claim 5, characterized in that, Both the Actor and Critic networks employ fully connected neural network models. The Critic network is designed as a fully connected neural network. The number of nodes in the input layer is equal to the number of variables in the state set, i.e., the input variables are twenty-six state variables. The number of nodes in the output layer is an evaluation value, which is used to judge the quality of the current state. The number of hidden layers and nodes can be defined by the user. Here, three hidden layers are designed, with 256, 128, and 128 nodes in each layer, respectively. The ReLU function is used as the activation function of the network, and the Adam optimizer is used to train the neural network. The Actor network is designed as a fully connected neural network with twenty-six state variables as input variables. It has two hidden layers with 256 and 128 nodes respectively. The output is the mean and standard deviation of the impulse maneuvers in three directions. The actual impulse maneuver values can be obtained by probability sampling. The ReLU function is used as the activation function of the network, and the Adam optimizer is used to train the neural network.
Citation Information
Patent Citations
Autonomous avoidance maneuvering method for multiple interceptors by spacecraft based on reinforcement learning
CN112001120A
Simulation method and system for spacecraft on-orbit game and storage medium
CN113268859A