An Intelligent Planning Method for Emergency Tasks of Space Debris Removal
Through the joint cleaning of contact-grabbing derail removal based on reinforcement learning and non-contact laser ablation-driven derail cleaning, the problems of high cost and poor mobility in space waste cleaning are solved, and efficient and low-cost cleaning effect is achieved.
Patent Information
- Application Number
- CN202311756578.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-20
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2043-12-20
AI Technical Summary
The prior art is difficult to effectively and at low cost to clean up space waste, especially when dealing with large-scale mission allocation and migration problems.
The joint cleaning of contact-grabbing off-rail removal based on reinforcement learning and contactless laser ablation-driven off-rail cleaning is achieved through an intelligent planning platform.
At lower cost, efficient space waste cleaning is achieved, significantly shortening the spacecraft reinforcement learning training time and having good migration.
Smart Images

Figure CN117875615B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of aerospace science and technology, and specifically relates to an intelligent planning method for emergency tasks of space debris cleaning. Background Technique
[0002] Humans have enjoyed various conveniences brought by various space missions in space activities. However, due to the poor protection of the space environment, the debris in outer space has shown an increasing trend, seriously polluting the space environment and posing a serious threat to the use of on-orbit spacecraft (including space stations and manned spacecraft) for future human use. Currently, there are more and more space missions with high costs and great values. Timely, low-cost, and efficient cleaning of space debris is an important part of space work.
[0003] The current research on the problem of space debris cleaning mainly focuses on the following three sub-problems: (1) Target screening and grouping problem: determining the set of space debris to be removed and allocating ADR tasks to specific service spacecraft; (2) Access sequence planning problem: determining the removal order and removal time of service spacecraft for space debris; (3) Trajectory transfer optimization problem: determining the optimal transfer trajectory between two target debris removal operations.
[0004] The research on the task planning problem for space debris cleaning focuses on how to allocate space debris cleaning units to effectively clean the targets to achieve the best cleaning effect. Its essence is a non-linear combinatorial optimization problem. As the number of debris cleaning units and targets increases, the size of the solution space also grows exponentially, which is an NP-complete problem. Common task allocation methods include traditional methods such as enumeration method, Hungarian algorithm, and dynamic programming algorithm, as well as heuristic algorithms such as genetic algorithm, ant colony algorithm, and particle swarm algorithm. The traditional methods have simple principles but cumbersome processes and are difficult to handle large-scale task allocation problems; heuristic algorithms use some heuristic rules to search for optimization in the solution space and can efficiently handle relatively large-scale task allocation problems. Their execution is one-time. When there is a new task, it needs to be executed from scratch and it is difficult to achieve migration. Summary of the Invention
[0005] In view of the above problems, the present invention proposes an intelligent planning method for emergency tasks of space debris cleaning, adopting a combined cleaning scheme of contact capture and off-orbit removal and non-contact laser ablation and drive off-orbit, with the processing ability of "multiple to one", and completing the cleaning task at a lower cost.
[0006] The present invention is realized through the following technical solutions:
[0007] An intelligent planning platform for emergency tasks of space debris cleaning:
[0008] The planning platform includes a contact debris cleaning tool task planning module based on reinforcement learning and a non-contact laser ablation deorbiting task planning module based on reinforcement learning, which are used to clean space debris in the scene;
[0009] The contact debris cleaning tool task planning module based on reinforcement learning includes several contact capture deorbiting and removal aircraft and launch vehicles;
[0010] The contact capture deorbiting and removal aircraft is specifically a rocket carrying a contact debris cleaning tool. The rocket is carried on the launch vehicle. First, it maneuvers from the station to the launch site, and then each contact capture deorbiting and removal aircraft selects a space debris target to execute the launch, calculates the launch window, and cleans the space debris;
[0011] The non-contact laser ablation deorbiting task planning module based on reinforcement learning includes several non-contact ablation-driven deorbiting lasers and non-contact ablation-driven deorbiting laser vehicles
[0012] The non-contact ablation-driven deorbiting laser is specifically a laser carrying a non-contact laser debris cleaning tool. The non-contact laser is carried on the non-contact ablation-driven deorbiting laser vehicle. First, it moves from the station to the launch site, and then each non-contact ablation-driven deorbiting laser selects a space debris target to execute the emission, calculates the emission window, and cleans the space debris.
[0013] Furthermore, when the contact debris cleaning tool task planning module based on reinforcement learning conducts space debris cleaning, the task constraints it is subject to include maneuvering speed, retraction time, deployment time, action altitude, cleaning probability, and action object;
[0014] The maneuvering speed refers to the speed of the launch vehicle during ground maneuvering, and this constraint does not exceed 30 m / s;
[0015] The retraction time refers to the time required for the vehicle to retract back to the station after the rocket carrying the contact debris cleaning tool is launched; in the task planning problem, it will be fixed at 64 seconds;
[0016] The deployment time refers to the time required for the equipment of the rocket carrying the contact debris cleaning tool to conduct necessary pre-launch work after maneuvering to the launch site. This constraint does not exceed 45 minutes;
[0017] The action altitude refers to the altitude that can be affected when the rocket carrying the contact debris cleaning tool executes the task. In the space junk cleaning task, it is determined by the initialization parameters of the rocket carrying the contact debris cleaning tool;
[0018] The cleaning probability refers to the probability of the rocket equipped with the contact debris cleaning tool cleaning the debris when it flies near the debris, generally 70%;
[0019] The object of action is the debris that has not been cleaned by other mission equipment during the mission.
[0020] Furthermore, the contact debris cleaning tool needs to perform reinforcement learning before mission planning: assume that all aircraft in the entire system are not affected by external forces; regard the rocket equipped with the contact debris cleaning tool as an independent agent, and abstract the working process of the rocket equipped with the contact debris cleaning tool into a mission planning problem, thereby transforming the problem into a task sequence for training a contact interceptor;
[0021] The state space of the rocket agent equipped with the contact debris cleaning tool includes its own spatial position information and the position information of the debris in the environment; the position information of the rocket agent equipped with the contact debris cleaning tool and the debris is described using the earth-fixed coordinate system, which is used to reflect the characteristics of the observation information during the reinforcement learning process;
[0022] At each time step, the contact rocket performs an action and then enters the next state until it reaches the last moment, that is, it stops at the last state; the debris can choose ground maneuvers and different targets at each state state, and the action output set is obtained from x launch sites and y space debris
[0023] action = {0,,,xy - 1}
[0024] During the training process, when choosing an action at each state, there is a probability of ε to randomly select an action, and in other cases, the action with the largest Q is selected according to the Q-table. The debris agent randomly selects an action for neural network training:
[0025] action = random.randint(0, xy - 1)
[0026] Design the reward function as:
[0027] r = r jidong +r end
[0028] Where:
[0029]
[0030] r jidong is the designed integer reward, including the fuel consumption of the current maneuvering strategy and the current mission execution situation; r end is the reward for the rocket agent equipped with the contact debris cleaning tool to clean the debris;
[0031] The reward for debris is set as an immediate return reward. When a scene of training is completed, a reward signal is obtained, and the reward settings are as follows:
[0032]
[0033]
[0034] Among them, the cleaning judgment criterion is whether the debris enters within the limit deviation correction distance of 5 km from the starting position of the terminal guidance of the contact-type equipment. If the debris is within 5 km before the terminal guidance of the contact-type equipment starts, it means the debris is destroyed.
[0035] Furthermore, when the non-contact laser ablation deorbiting mission planning module based on reinforcement learning conducts space debris cleaning, the task constraints it is subject to include maneuvering speed, retraction time, deployment time, working duration, interval time, irradiation altitude, and output power;
[0036] The maneuvering speed is the speed at which the vehicle carrying the non-contact laser executes the task maneuvers on the ground; this constraint does not exceed 30 m / s.
[0037] The retraction time is the time required for the vehicle to retract back to the station after the non-contact laser is launched; in the task planning problem, this constraint is fixed at 64 seconds;
[0038] The deployment time is the time required for the equipment carrying the non-contact laser to perform the necessary pre-launch work after maneuvering to the launch site; this constraint does not exceed 45 minutes;
[0039] The working duration is the time from launch to stop of irradiation when the non-contact laser executes the task, which is about 20 seconds;
[0040] The interval time is the time required to switch different irradiation targets when the non-contact laser debris cleaning tool executes the task, which is set according to the specific scenario;
[0041] The irradiation altitude is the altitude that the non-contact laser can irradiate when the non-contact laser debris cleaning tool executes the task. In the space debris cleaning task, it is about the general altitude of low-orbit space debris;
[0042] The output power refers to the output power of the non-contact laser, which is set to 100 kw by default.
[0043] Furthermore, the non-contact laser requires reinforcement learning before mission planning: the state space of the non-contact laser debris removal tool agent includes its own spatial position information and the position information of the debris in the environment, and the position information of the debris is described by the earth-fixed coordinate system, and the earth-fixed coordinate system position information of the reference debris is added;
[0044] At each time step, the agent performs an action and then enters the next state. Until the last moment, that is, the last state stops, the agent can choose not to maneuver, maneuver (in four directions) in each state, and open the eyelids for a total of six actions, so the action output set is: action = {0, 1, 2, 3, 4, 5}
[0045] Action equals 0 to 5, which respectively indicates maneuvering directly upward under the system, maneuvering directly downward under the system, maneuvering directly to the left under the system, maneuvering directly to the right under the system, no maneuvering, and eyelid opening;
[0046] During the training process, when selecting an action in each state, there is a probability of ε to randomly select an action. In other cases, the action with the largest Q is selected according to the Q table. The agent randomly selects an action to train the neural network:
[0047] action = random.randint(0,6)
[0048] The reward for the fragment is set as an immediate reward. When one episode of training is completed, a reward signal is obtained. The reward is set as follows:
[0049] Fuel consumption: negative bonus is -10; eyelid opening bonus is 100;
[0050]
[0051] Among them, the cleaning judgment standard is whether the debris enters the 15km range of the predetermined irradiation position of the laser debris cleaning tool and whether the power density of the non-contact laser to the target reaches the threshold of space debris ablation and deorbit. If the debris does not enter the irradiation airspace of the laser debris cleaning tool or the power density of the non-contact laser to the target is insufficient, it means that the debris has not been cleaned.
[0052] Furthermore, the specific treatment of space debris is as follows: within a fixed time period, the launch vehicle is transported and maneuvered from the storage warehouse to the launch site at a certain maneuvering speed, and after being transported to the launch site, it is deployed within the specified time, that is, the necessary work before the launch is carried out;
[0053] When the contact capture deorbit removal vehicle performs a cleaning mission, it can clean up to a certain height. In the space debris cleaning mission, the probability of cleaning up space debris is generally 70%, which is determined by the initialization parameters of the contact capture deorbit removal vehicle.
[0054] When the non-contact ablation driven de-orbit laser performs a cleaning task, after the non-contact ablation driven de-orbit laser is emitted, the non-contact ablation driven de-orbit laser beam is focused on the target, and after sufficient target power density and time accumulation, the space debris is ablated and driven off-orbit after reaching the threshold of 0.2w / cm2;
[0055] The non-contact ablation-driven de-orbiting laser and the contact capture de-orbiting removal vehicle mainly target space debris, and the mission completion calculation model is:
[0056]
[0057] Among them, Ks is the weight of the number of cleanups, which is set to 0.9; Kt is the timeliness weight, which is set to 0.1;
[0058] The laser ablation deorbiting judgment standard is that the target power density reaches the threshold value of 0.2w / cm2, and the contact capture judgment standard is whether the space debris enters the contact capture deorbit removal vehicle terminal guidance start position within the limit deviation correction distance of 5km; if the space debris is within 5km before the contact capture deorbit removal vehicle terminal guidance start, it means that the space debris is captured;
[0059] Finally, after the non-contact ablation-driven deorbiting laser or contact capture deorbiting removal vehicle is launched, the vehicle retreats back to its base.
[0060] A method for intelligent planning of emergency tasks for space debris cleaning, the method specifically comprising the following steps:
[0061] Step 1: Set up the initial space debris cleanup scenario in SpaceSim software, including the start time, ground cleanup tools, and debris;
[0062] Step 2: Set the agent action space, state space and reward function in the external Python algorithm, and set the end condition of training; determine the agent neural network structure, training hyperparameters and number of training rounds;
[0063] Step 3: Initialize the training neural network and target neural network parameters, and initialize the experience pool space;
[0064] Step 4: Perform agent reinforcement learning according to the designer's needs. Use the plug-in algorithm to call the scenes designed in SpaceSim for each round of training. The agent makes decisions according to the initialized neural network in each state, and has a certain probability of random exploration. The state, action, reward and transfer state of each scene will be stored in the experience pool for training neural network learning;
[0065] Step 5: When the experience pool meets the maximum storage space, sample learning is performed according to the set neural network gradient, and the training neural network parameters are continuously updated; after the set number of learning times, the training neural network parameters are copied to the target neural network;
[0066] Step 6: For each round of training, when all the fragments are cleared or the set scene end time is reached, the round ends and step 4 is repeated; when the number of training times reaches the predetermined requirement, the training ends and the neural network and parameters are saved;
[0067] Step 7: The stored data can be monitored in real time using Python and displayed in curve tables; SpaceSim data can be played back to display the two-dimensional and three-dimensional status;
[0068] Step 8: Perform mission planning data statistics based on the data calculated by Python and SpaceSim to complete the planning of the mission platform.
[0069] An electronic device comprises a memory and a processor, wherein the memory stores a computer program, and the processor implements the steps of the above method when executing the computer program.
[0070] A computer-readable storage medium is used to store computer instructions, and when the computer instructions are executed by a processor, the steps of the above method are implemented.
[0071] Beneficial effects of the present invention
[0072] The commonly used contact-type capture and de-orbit removal method can complete the cleaning task, but due to its high cost, the present invention adopts a solution of combined cleaning of contact-type capture and de-orbit removal and non-contact laser ablation driven de-orbit removal to complete the cleaning task at a lower cost.
[0073] Relying on the SpaceSim reinforcement learning algorithm research and training platform, the present invention builds an intelligent emergency space mission planning platform for multi-dimensional space junk cleaning styles, conducts mission planning training and decision-making deduction, designs the operating logic of the task modules within the software according to the mission requirements, functional requirements and performance requirements, divides the modules and functions required for tasks such as active space debris cleaning, and significantly shortens the spacecraft reinforcement learning training time.
[0074] The present invention is based on empirical learning, with high sample utilization rate, more suitable for solving engineering problems, capable of handling complex environments and autonomous learning; reinforcement learning can also handle long-term reward problems, and the obtained results can be quickly reused and have transferability.
[0075] The space debris treatment method of the present invention has a larger number than the space debris method, has the processing ability of "many-to-one", and can better ensure the completion degree of the task. BRIEF DESCRIPTION OF THE DRAWINGS
[0076] Figure 1 It is a composition diagram of the SpaceSim reinforcement learning algorithm research and training platform of the present invention.
[0077] Figure 2 It is a flow chart of the reinforcement learning of the present invention.
[0078] Figure 3 It is a working flow chart of the space debris cleaning task planning platform of the present invention.
[0079] Figure 4 It is the loss value.
[0080] Figure 5 It is a reward curve graph.
[0081] Figure 6 It is an example diagram of a partial task sequence. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0082] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0083] Combined with Figures 1 to 6 。
[0084] An intelligent planning platform for emergency tasks of space debris cleaning:
[0085] The planning platform includes a contact debris cleaning tool task planning module based on reinforcement learning and a non-contact laser ablation deorbiting task planning module based on reinforcement learning, which are used to clean space debris in the scene (five space debris are set as space garbage, that is, space debris threatening spacecraft and flying in space orbits);
[0086] The contact debris cleaning tool task planning module based on reinforcement learning includes several (set 6) contact capture deorbiting removal aircraft and launch vehicles;
[0087] The contact capture and off-orbit removal vehicle is specifically a rocket equipped with a contact debris cleaning tool. The rocket is carried on a launch vehicle. First, it maneuvers from the station to the launch site. Then, each contact capture and off-orbit removal vehicle selects a space debris target for launch, calculates the launch window, and cleans the space debris.
[0088] The non-contact laser ablation off-orbit mission planning module based on reinforcement learning includes several (set to 6) non-contact ablation-driven off-orbit lasers and non-contact ablation-driven off-orbit laser vehicles.
[0089] The non-contact ablation-driven off-orbit laser is specifically a laser equipped with a non-contact laser debris cleaning tool. The non-contact laser is carried on a non-contact ablation-driven off-orbit laser vehicle. First, it moves from the station to the launch site. Then, each non-contact ablation-driven off-orbit laser selects a space debris target for emission, calculates the emission window, and cleans the space debris.
[0090] When the contact debris cleaning tool task planning module based on reinforcement learning conducts space debris cleaning, the task constraints it is subject to include maneuvering speed, retraction time, deployment time, action altitude, cleaning probability, and action object.
[0091] The maneuvering speed refers to the speed of the launch vehicle (the vehicle for the rocket equipped with the contact debris cleaning tool to perform tasks) during ground maneuvering, and this constraint does not exceed 30 m / s.
[0092] The retraction time refers to the time required for the vehicle to retract back to the station after the rocket equipped with the contact debris cleaning tool is launched; in the task planning problem, it will be fixed at 64 seconds.
[0093] The deployment time refers to the time required for the equipment of the rocket equipped with the contact debris cleaning tool to perform tasks to conduct necessary pre-launch work after maneuvering to the launch site, and this constraint does not exceed 45 minutes.
[0094] The action altitude refers to the altitude that can be affected when the rocket equipped with the contact debris cleaning tool performs tasks. In the space junk cleaning task, it is determined by the initial parameters of the rocket equipped with the contact debris cleaning tool.
[0095] The cleaning probability refers to the probability of cleaning up the debris when the rocket equipped with the contact debris cleaning tool flies near the debris, generally 70%.
[0096] The action object is the debris that has not been cleaned by other task equipment in the task.
[0097] The contact debris cleaning tool needs reinforcement learning before mission planning: Since the relative speed of both sides is high and the duration is very short, it is assumed that all aircraft in the entire system are not disturbed by external forces; the rocket carrying the contact debris cleaning tool is regarded as an independent intelligent agent, and the working process of the rocket carrying the contact debris cleaning tool is abstracted into a mission planning problem, and thus the problem is transformed into a task sequence for training a contact interceptor.
[0098] The state space of the rocket intelligent agent carrying the contact debris cleaning tool includes its own spatial position information and the position information of debris in the environment; the position information of the rocket intelligent agent carrying the contact debris cleaning tool and the debris is described using the Earth-fixed coordinate system, which is used to reflect the characteristics of the observation information during the reinforcement learning process.
[0099] At each time step, the contact rocket performs an action and then enters the next state until it reaches the last moment, that is, it stops at the last state; the debris can choose ground maneuvers and different targets at each state state, and the action output set is obtained from x launch sites and y space debris.
[0100] action = {0,,, xy - 1}
[0101] If there are a total of fifteen actions for three launch sites and five space debris, the action output set is:
[0102] action = {0, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14}
[0103]
[0104]
[0105] During the training process, when choosing an action at each state, there is a probability of ε to randomly select an action, and in other cases, the action with the largest Q is selected according to the Q table. The debris intelligent agent randomly selects an action for neural network training:
[0106] action = random.randint(0, xy - 1)
[0107] Designing a suitable reward function is crucial for the reinforcement learning algorithm, which will directly affect the training speed and even the feasibility of the algorithm. To solve problems such as poor algorithm convergence and slow learning speed caused by sparse rewards, reward function shaping is introduced. The designed reward function is: r = r jidong + r end
[0108] Where:
[0109]
[0110] Considering that the completion of the task includes both the task completion status and the task completion quality. r jidong For the designed shaping reward, including the fuel consumption of the current maneuvering strategy and the current task execution status; r end The reward for the rocket agent equipped with a contact debris cleaning tool to clean debris;
[0111] The reward for debris is set as an immediate return reward. When a scene of training is completed, a reward signal is obtained, and the reward is set as follows:
[0112]
[0113]
[0114] Among them, the cleaning judgment criterion is whether the debris enters within the limit deviation correction distance of 5 km from the end guidance start position of the contact equipment. If the debris is within 5 km before the end guidance start of the contact equipment, it means the debris has been destroyed. The ultimate optimization goal of the rocket agent equipped with a contact debris cleaning tool is to clean as many space debris as possible while consuming less maneuvering fuel.
[0115] When the non-contact laser ablation deorbiting mission planning module based on reinforcement learning conducts space debris cleaning, the task constraints it is subject to include maneuvering speed, retraction time, deployment time, working duration, interval time, irradiation altitude, and light output power;
[0116] The maneuvering speed is the speed of the vehicle equipped with a non-contact laser when maneuvering on the ground to perform the task; this constraint does not exceed 30 m / s.
[0117] The retraction time is the time required for the vehicle to retract back to the station after the non-contact laser is launched; in the task planning problem, this constraint is fixed at 64 seconds;
[0118] The deployment time is the time required for the equipment of the non-contact laser to perform necessary pre-launch work after maneuvering to the launch site to perform the task. This constraint does not exceed 45 minutes;
[0119] The working duration is the time from launch to stop of irradiation when the non-contact laser performs the task, which is about 20 seconds;
[0120] The interval time is the time required to switch different irradiation targets when the non-contact laser debris cleaning tool performs the task, which is set according to the specific scenario;
[0121] The irradiation height is the height that the non-contact laser debris cleaning tool can irradiate when performing tasks. In the space debris cleaning task, it is approximately the average height of low-orbit space debris.
[0122] The output power refers to the output power of the non-contact laser, which is set to 100 kw by default.
[0123] Before task planning, the non-contact laser needs to perform reinforcement learning. Similar to the rocket agent equipped with a contact debris cleaning tool, the state space of the non-contact laser debris cleaning tool agent includes its own spatial position information and the position information of debris in the environment. The position information of the debris is described using the Earth-fixed coordinate system, and the position information of the reference debris in the Earth-fixed coordinate system is added.
[0124] At each time step, the agent takes an action and then enters the next state. Until the last moment, that is, the last state stops. In each state state, the agent can choose not to maneuver, maneuver in four directions, or open the eyelid, a total of six actions. Then the action output set is: action={0,1,2,3,4,5}
[0125] When Action is equal to 0 to 5, it represents maneuvering directly above the system, maneuvering directly below the system, maneuvering directly to the left of the system, maneuvering directly to the right of the system, not maneuvering, and opening the eyelid, respectively.
[0126] During the training process, when choosing an action in each state, there is a probability of ε to randomly select an action, and in other cases, the action with the maximum Q value is selected according to the Q table. The agent randomly selects an action for neural network training: action = random.randint(0,6)
[0127] The reward for the debris is set as an immediate return reward. After completing a scene of training, a reward signal is obtained. The reward is set as follows:
[0128] Fuel consumption: The negative reward is -10; the reward for opening the eyelid is 100.
[0129]
[0130] Among them, the cleaning judgment criterion is whether the debris enters within 15 km of the predetermined irradiation position of the laser debris cleaning tool and the on-target power density of the non-contact laser reaches the ablation and deorbiting threshold of space debris. If the debris does not enter the irradiation airspace of the laser debris cleaning tool or the on-target power density of the non-contact laser is insufficient, it means that the debris has not been cleaned. The ultimate optimization goal of the agent is to clean as many target debris as possible while minimizing maneuvers.
[0131] An intelligent planning method for emergency tasks of space debris cleaning:
[0132] Relying on the SpaceSim reinforcement learning algorithm research and training platform, an intelligent emergency space mission planning platform for multi-dimensional space debris cleaning styles is built to conduct mission planning training and decision-making deduction. Based on several preset mission scenarios and cleaning work rules, it supports users to customize the number of cleaning tools (the number of agents) and the types of cleaning tools (the types of agents) called. It can customize the functions and attributes of the cleaning tools, support the customization of the agent behavior reward rules, etc., and specifically opens the RL-API interface for multi-agent deep reinforcement learning. The environment supports the implementation of algorithms using the Python language and supports the integrated call of common deep learning frameworks such as TensorFlow and PyTorch. The module composition is as Figure 1 shown. The SpaceSim software can support the whole-process simulation and analysis of space missions and provides relatively complete tool functions related to space calculations.
[0133] Design the operation logic of the internal task modules of the software according to the task requirements, functional requirements, and performance requirements, and divide and implement the modules and functions required for tasks such as active space debris cleaning. The reinforcement learning process requires a large amount of training. To improve the operation speed of the software, Python and C++ joint programming based on Pybind is used. The complex reinforcement learning calculation process is implemented on the Python platform, and the spacecraft orbit recursion, the solution process of relevant debris sub-components or ground maneuver strategies that require a large number of matrix operations and loops are implemented using the C++ language. This avoids the disadvantage of low calculation efficiency caused by the lack of compilation optimization of interpreted languages such as Python in matrix calculations and large-scale orbit operations. Through the above means, the reinforcement learning training time of the spacecraft is significantly shortened.
[0134] In the application scenario of space debris cleaning strategy generation, the cleaning tool is equivalent to an agent, and it is in a fixed space environment (Environment) with space debris. At each step in the simulation process, the state information of the debris and the cleaning tool will change. Therefore, at each step, the debris and the cleaning tool will have a specific position and state (State). The cleaning tool will choose to perform an action (Action) in each state in the environment, and these actions will bring corresponding rewards (Reward), and then enter the next state. In the new state, the debris will take new actions, and so on. As Figure 2 shown.
[0135] Therefore, based on the generation of task planning strategies for forced learning training, it is first necessary to establish a space debris cleaning scenario and set up a space debris cleaning task. Consider the cleaning tool as an agent for reinforcement learning training, and based on the cleaning strategy generation model constructed above, consider the different tools carried by the agent and the different task objectives, and design different action spaces for the agent. Each time step is a scene. The agent selects an action to execute, transfers the state according to the selected action, and enters the next scene until the simulation ends or the agent's task fails as a termination condition. Based on the preset mission scenarios and engagement rules of the SpaceSim software platform according to the mission requirements, the plug-in task planning strategy generation algorithm written in Python is used for training. The specific strategy generation process is as follows: Figure 3 shown.
[0136] The method specifically comprises the following steps:
[0137] Step 1: Set up the initial space junk cleanup scenario in the SpaceSim software, including the start time, ground cleanup tools, and debris.
[0138] Step 2: Set the agent action space, state space and reward function in the external Python algorithm, and set the end condition of training. Determine the agent neural network structure, training hyperparameters and number of training rounds.
[0139] Step 3: Initialize the training neural network and target neural network parameters, and initialize the experience pool space.
[0140] Step 4: Perform agent reinforcement learning according to the designer's needs. Use the plug-in algorithm to call the designed scenes in SpaceSim for each round of training. The agent makes decisions according to the initialized neural network in each state, and has a certain probability of random exploration. The state, action, reward and transfer state of each scene will be stored in the experience pool for training neural network learning.
[0141] Step 5: When the experience pool meets the maximum storage space, sample learning is performed according to the set neural network gradient, and the training neural network parameters are continuously updated; after the set number of learning times, the training neural network parameters are copied to the target neural network.
[0142] Step 6: For each round of training, when all the fragments are cleared or the set scene end time is reached, the round ends and step 4 is repeated; when the number of training times reaches the predetermined requirement, the training ends and the neural network and parameters are saved.
[0143] Step 7: The stored data can be used for real-time status monitoring and curve table display using Python; SpaceSim data can be played back for two-dimensional and three-dimensional status display.
[0144] Step 8: Perform task planning data statistics based on the data calculated by Python and SpaceSim to complete the planning of the task platform.
[0145] Example: A typical scenario was constructed. There are two parties in the scenario, namely the space debris disposal party and the space debris party. Among them, the space debris disposal party protects important space spacecraft, and the space debris poses a certain danger to the spacecraft. The space debris disposal party has 6 contact-type capture and deorbit removal payload launch vehicles, 6 non-contact ablation-driven deorbit laser launch vehicles, 3 storage warehouses, and 8 ground launch sites. The space debris party has 5 space debris, all of which are distributed in low Earth orbit.
[0146] Deployment of the space debris disposal party: The launch parameters of the contact-type capture and deorbit removal vehicle are set as follows:
[0147] Table 1 Launch parameters of the contact-type capture and deorbit removal vehicle
[0148]
[0149]
[0150] The corresponding loading relationship between the contact-type capture and deorbit removal vehicle and the contact-type capture and deorbit removal payload launch vehicle is as follows:
[0151] Table 2 Corresponding loading relationship between the contact-type capture and deorbit removal vehicle and the contact-type capture and deorbit removal payload launch vehicle
[0152]
[0153] The parameters of the non-contact ablation-driven deorbit laser are as follows:
[0154] Table 3 Parameters of the non-contact ablation-driven deorbit laser
[0155]
[0156] The storage warehouses are as follows:
[0157] Table 4 Storage longitude and latitude
[0158]
[0159] The longitude and latitude of the launch sites are as follows:
[0160] Table 5 Longitude and latitude of the launch sites
[0161]
[0162]
[0163] Scene rule settings:
[0164] Clean up the scene and set five space debris as space junk, that is, space debris that poses a threat to spacecraft. They are flying in space orbits. At this time, six contact capture and deorbit removal vehicles and six non-contact ablation-driven deorbit lasers are used to clean up these five space debris. The non-contact ablation-driven deorbit laser is carried on a non-contact ablation-driven deorbit laser vehicle. First, it moves from the station to the launch site. Then each non-contact ablation-driven deorbit laser selects a space debris target to execute the emission, calculates the emission window, and cleans up the space debris. The contact capture and deorbit removal vehicle is carried on a launch vehicle. First, it maneuvers from the station to the launch site. Then each contact capture and deorbit removal vehicle selects a space debris target to execute the launch, calculates the launch window, and cleans up the space debris.
[0165] Scene time and step size settings are as follows:
[0166] Scene start time: 2023 / 5 / 20_12:00:00
[0167] Scene end time: 2023 / 5 / 21_12:00:00
[0168] Simulation step size: 1s; Space junk situation:
[0169] Space junk situation: From 2023 / 5 / 20_12:00:00 to 2023 / 5 / 21_12:00:00, it is not cleaned by itself and poses a threat to high-value spacecraft at 2023 / 5 / 21_12:00:00..
[0170] Subdivide the process according to the time sequence, which can be specifically divided into:
[0171] 1) Flight segment: Continuously fly in orbit and need to reach the time point of 2023 / 5 / 21_12:00:00. It cannot be cleaned by space junk treatment equipment during the operation.
[0172] 2) Target imaging segment: After running to the time point of 2023 / 5 / 21_12:00:00, without being cleaned, it poses a threat to high-value spacecraft.
[0173] Target of space junk treatment equipment:
[0174] Target of space junk treatment equipment: Protect high-value spacecraft from being threatened and damaged by space junk.
[0175] Subdivide the process according to the time sequence, which can be specifically divided into:
[0176] 1) Ground Maneuver: After the ground detected space debris at 12:00:00 on May 20, 2023, it is necessary to transport the non-contact ablation-driven deorbiting laser and the contact capture and deorbiting removal vehicle from the storage warehouse to the launch site by using the non-contact ablation-driven deorbiting laser vehicle and the contact capture and deorbiting removal payload launch vehicle.
[0177] 2) Launch and Cleanup Phase: The space debris treatment equipment selects the non-contact ablation-driven deorbiting laser and the contact capture and deorbiting removal vehicle to perform a multi-to-one interception. Before the space debris poses a threat to high-value spacecraft, making it deorbit or be captured is considered a successful cleanup.
[0178] The start response time of the space debris treatment equipment is: 12:00:00 on May 20, 2023 (UTC). The space debris treatment equipment is deployed in the manner described above. The space debris treatment equipment calls multiple optimal non-contact ablation-driven deorbiting lasers and contact capture and deorbiting removal vehicles to perform the cleanup task.
[0179] In summary, the start time of the scenario is: 12:00:00 on May 20, 2023, the simulation step size is 1 (s), and the end time of the scenario is: 12:00:00 on May 21, 2023 (UTC).
[0180] Specifically, the way for the space debris treatment equipment to achieve the goal is that within the above time period, the launch vehicle is transported from the storage warehouse to the launch site at a certain maneuvering speed. The maneuvering speed refers to the speed of the vehicle carrying the non-contact ablation-driven deorbiting laser and the contact capture and deorbiting removal vehicle during ground maneuvering. This speed does not exceed 30 m / s. After being transported to the launch point, it is deployed within a certain time, that is, the necessary pre-launch work is carried out, and this time does not exceed 45 minutes. When the non-contact ablation-driven deorbiting laser performs the cleanup task, after the non-contact ablation-driven deorbiting laser is launched, the non-contact ablation-driven deorbiting laser beam is focused on the target. After having sufficient target power density and time accumulation and reaching the threshold of 0.2 w / cm2, the space debris is ablated and driven to deorbit. When the contact capture and deorbiting removal vehicle performs the cleanup task, it can clean up to a certain height, which is determined by the initialization parameters of the contact capture and deorbiting removal vehicle in the space debris cleanup task. When the contact capture and deorbiting removal vehicle cleans up space debris, the probability of cleaning up the space debris is generally 70%. In this task, the objects of action of the non-contact ablation-driven deorbiting laser and the contact capture and deorbiting removal vehicle are mainly space debris. The task completion calculation model is:
[0181]
[0182] Among them, Ks is the weight of the number of cleanups. Since it is the fundamental issue for task completion, it has the highest importance and is set to 0.9. Kt is the timeliness weight, which is an effect evaluation factor and is set to 0.1.
[0183] In this mission, the mission completion rate must be guaranteed to exceed 80%, that is, at least 4 pieces of space junk or space debris must be deorbited or destroyed.
[0184] The laser ablation deorbiting judgment standard is that the target power density reaches the threshold value of 0.2w / cm2, and the contact capture judgment standard is whether the space debris enters the contact capture deorbit removal vehicle terminal guidance start position within the limit deviation range of 5km. If the space debris is within 5km before the contact capture deorbit removal vehicle terminal guidance start, it means that the space debris is captured. The ultimate optimization goal of all intelligent agents is to clean up as much space junk and space debris as possible while reducing maneuvering consumption.
[0185] Finally, after the non-contact ablation-driven deorbiting laser or contact capture deorbiting removal vehicle is launched, it takes 64 seconds for the vehicle to withdraw to its base.
[0186] Simulation experiment analysis:
[0187] Based on the above experimental background, a simulation experiment environment was built, and the task planning training model based on the reinforcement learning algorithm was used to conduct simulation experiment analysis. A total of 1,300 trainings were conducted on the above simulation experiment environment. The loss value and profit curve of the algorithm's expected reward and current reward are shown in the figure below. Figure 4 , 5 .
[0188] After training, the neural network decision model has converged. The average profit curve is shown in the figure. After a certain number of training times, it gradually converges.
[0189] The best model was brought into the environment for testing, and 100 sets of test results were obtained, as shown in the following table.
[0190] Table 6 Test results
[0191]
[0192]
[0193] It can be seen that the task completion rate of group 100 is the highest. The strategy of group 100 is taken and input into the support platform. Finally, the task planning task sequence obtained by the agent decision is shown in the following part of the sequence in SpaceSim: Figure 6 shown.
[0194] As can be seen from the results, six non-contact ablation-driven deorbiting lasers and six contact capture and deorbiting removal vehicles have all maneuvered to relatively appropriate launch points and completed the launch. First, the non-contact ablation-driven deorbiting lasers _4, _5, and _6 ablated the space debris _1 out of orbit, the non-contact ablation-driven deorbiting laser _3 ablated the space debris _5 out of orbit, and the non-contact ablation-driven deorbiting laser _1 ablated the space debris _2 out of orbit. Then, the contact capture and deorbiting removal vehicle 2 captured the space debris _3, and the contact capture and deorbiting removal vehicle 1 captured the space debris _4. All the space debris has been ablated out of orbit or captured, and the mission completion rate is 92.1%, achieving the mission objective.
[0195] The present invention has the following advantages:
[0196] 1) Based on empirical learning, high sample utilization rate: Traditional optimization decision algorithms usually require a mathematical model to describe the system and then use optimization techniques to solve for the optimal solution. Reinforcement learning, on the other hand, does not require prior knowledge of the system model but learns the behavior of the system through interaction with the environment. Compared with traditional algorithms such as particle swarm and differential evolution algorithms, the advantage of reinforcement learning is that it makes full use of historical samples. Any optimization algorithm needs to find a direction of descent or a direction of descent with a certain probability. Gradient optimization algorithms choose the negative gradient, while particle swarm and differential evolution obtain a high-probability descent direction through random hybridization with the optimal solution. The above algorithms only use the current optimal sample to update, while reinforcement learning obtains the descent direction based on the evaluation function. Since the evaluation function is based on all existing samples, the information is utilized more fully.
[0197] 2) Can be more suitable for solving engineering problems, that is, it can handle complex environments and can learn adaptively: For contact capture and deorbiting removal vehicles, non-contact ablation-driven deorbiting lasers, etc., after taking an action, the environment changes due to the action. For example, in the scenario of space debris cleaning, the maneuvering destination of the vehicle affects the subsequent mission environment, and at this time, the previous decisions for a specific environment lose their meaning. Secondly, it is to solve the problem of model-free dynamic programming. Reinforcement learning only requires the agent to define the reward function and then, through continuous interaction with the environment, optimize the strategy according to the reward value to find the relative optimum. That is, the reinforcement learning algorithm is adaptive and can continuously improve its decision-making strategy in the process of continuous interaction with the environment and gradually approach the optimal strategy. Generally speaking, the advantage of reinforcement learning is that it can learn the optimal decision through interaction with the environment, adapt to environmental changes, and is adaptive without knowing the environment model and long-term rewards.
[0198] 3) Capable of handling long-term return problems: Reinforcement learning can handle long-term return problems, that is, the situation where the decision made at a certain point in time may only receive a return at a future point in time. Reinforcement learning can consider long-term impacts. For example, in the space debris cleaning scenario, the impact of vehicle maneuvering on the final cleaning result is a long-term return problem.
[0199] 4) Quick reuse of results: What reinforcement learning obtains is a Policy, while what traditional optimization algorithms obtain is only a solution. Strictly speaking, the output of reinforcement learning is not an optimal Action, but an optimal Policy.
[0200] 5) Result stability, i.e., result transferability: Reinforcement learning is more likely to achieve transfer optimization of similar optimization problems in optimization problems. This feature is an advantage brought by the Learning-based method. Similar optimization problems include similarities in the number of preset parameters of the problem or similarities in the nature of the target parameters of the problem research. In reinforcement learning, using machine learning models such as neural networks in reinforcement learning naturally has the possibility of achieving transfer between similar optimization problems.
[0201] An electronic device includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the steps of the above method are implemented.
[0202] A computer-readable storage medium is used to store computer instructions, and when the computer instructions are executed by a processor, the steps of the above method are implemented.
[0203] The memory in the embodiments of the present application may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. Among them, the non-volatile memory may be a read only memory (ROM), a programmable ROM (PROM), an erasable PROM (EPROM), an electrically erasable PROM (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of RAM are available, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchlink DRAM (SLDRAM), and direct rambus RAM (DR RAM). It should be noted that the memory of the method described in the present invention is intended to include but not limited to these and any other suitable types of memory.
[0204] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired means such as coaxial cables, optical fibers, digital subscriber line (DSL), or wireless means such as infrared, wireless, microwave, etc. The computer-readable storage medium can be any available medium that can be accessed by a computer, or a data storage device such as a server or data center that includes one or more integrated available media. The available medium can be a magnetic medium, such as a floppy disk, hard disk, magnetic tape, an optical medium, such as a high-definition digital video disc (DVD), or a semiconductor medium, such as a solid state disc (SSD), etc.
[0205] In the implementation process, the steps of the above method can be completed by the integrated logic circuit of the hardware in the processor or the instructions in the form of software. The steps of the method disclosed in the embodiments of the present application can be directly embodied as being executed by the hardware processor, or executed by the combination of the hardware and software modules in the processor. The software module can be located in a mature storage medium in the art, such as a random access memory, flash memory, read-only memory, programmable read-only memory, or electrically erasable programmable memory, register, etc. This storage medium is located in the memory, and the processor reads the information in the memory and combines its hardware to complete the steps of the above method. To avoid repetition, it will not be described in detail here.
[0206] It should be noted that the processor in the embodiments of the present application may be an integrated circuit chip with signal processing capabilities. In the implementation process, the steps of the above method embodiments can be completed by the integrated logic circuit in the hardware of the processor or instructions in the form of software. The above-mentioned processor may be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. It can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present application. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The steps of the method disclosed in combination with the embodiments of the present application can be directly embodied as being executed by the hardware decoding processor, or executed by a combination of the hardware and software modules in the decoding processor. The software module may be located in a mature storage medium in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, or an electrically erasable programmable memory, a register, etc. This storage medium is located in the memory, and the processor reads the information in the memory and combines its hardware to complete the steps of the above method.
[0207] The above has introduced in detail an intelligent planning method for emergency tasks of space debris cleaning proposed by the present invention, and expounded the principle and implementation manner of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention; at the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present invention.
Claims
1. An intelligent planning platform for emergency tasks of space debris cleaning, characterized in that: The planning platform includes a contact debris cleaning tool task planning module based on reinforcement learning and a non-contact laser ablation deorbiting task planning module based on reinforcement learning, which are used to clean space debris in the scenario; The contact debris cleaning tool task planning module based on reinforcement learning includes several contact capture and deorbiting removal aircraft and launch vehicles; The contact capture and deorbiting removal aircraft is specifically a rocket carrying a contact debris cleaning tool. The rocket is carried on the launch vehicle. First, it maneuvers from the station to the launch site, and then each contact capture and deorbiting removal aircraft selects a space debris target to execute the launch, calculates the launch window, and cleans the space debris; The non-contact laser ablation deorbiting task planning module based on reinforcement learning includes several non-contact ablation-driven deorbiting lasers and non-contact ablation-driven deorbiting laser vehicles The non-contact ablation-driven deorbiting laser is specifically a laser carrying a non-contact laser debris cleaning tool. The non-contact laser is carried on the non-contact ablation-driven deorbiting laser vehicle. First, it moves from the station to the launch site, and then each non-contact ablation-driven deorbiting laser selects a space debris target to execute the emission, calculates the emission window, and cleans the space debris; Before the contact debris cleaning tool conducts task planning, reinforcement learning is required: assume that all aircraft in the entire system are not affected by external forces; regard the rocket carrying the contact debris cleaning tool as an independent intelligent agent, and abstract the working process of the rocket carrying the contact debris cleaning tool as a task planning problem, thereby transforming the problem into a task sequence for training a contact interceptor; The state space of the rocket intelligent agent carrying the contact debris cleaning tool includes its own space position information and the position information of debris in the environment; the position information of the rocket intelligent agent carrying the contact debris cleaning tool and the debris is described using the earth-fixed coordinate system, which is used to reflect the characteristics of the observation information during the reinforcement learning process; At each time step, the contact rocket performs an action and then enters the next state until it reaches the last moment, that is, it stops at the last state; at each state state, the debris can choose ground maneuvering and different targets, and the action output set is obtained from x launch sites and y space debris action = {0,..., xy - 1} During the training process, when choosing an action at each state, there is a probability of ε to randomly select an action. In other cases, the action with the largest Q value is selected according to the Q table. The debris intelligent agent randomly selects an action for the training of the neural network: design the reward function; formulate the cleaning judgment criterion; Before the non-contact laser conducts task planning, reinforcement learning is required: similar to the rocket intelligent agent carrying the contact debris cleaning tool, the state space of the non-contact laser debris cleaning tool intelligent agent includes its own space position information and the position information of debris in the environment; The position information of the debris is described using the earth-fixed coordinate system, and the earth-fixed coordinate system position information of the reference debris is added; At each time step, the agent takes an action and then enters the next state; it continues until the last moment, i.e., it stops at the last state. In each state, the agent can choose not to perform a maneuver, maneuvers in four directions, or open the eyelids, for a total of six actions. Thus, the action output set is: action={0,1,2,3,4,5} When Action is equal to 0 to 5, it represents maneuvering directly upward in the system, maneuvering directly downward in the system, maneuvering directly to the left in the system, maneuvering directly to the right in the system, not performing a maneuver, and opening the eyelids, respectively; During the training process, when choosing an action in each state, there is a probability of ε of randomly selecting an action, and in other cases, the action with the maximum Q is selected according to the Q-table; the agent randomly selects an action to train the neural network: design the reward function; formulate the cleaning judgment criterion.
2. The planning platform according to claim 1, characterized in that: When the contact debris cleaning tool task planning module based on reinforcement learning performs space debris cleaning, the task constraints it is subject to include maneuvering speed, retraction time, deployment time, action height, cleaning probability, and action object; The maneuvering speed refers to the speed of the launch vehicle during ground maneuvering, and this constraint does not exceed 30 m / s; The retraction time refers to the time required for the vehicle carrying the contact debris cleaning tool to retract to the station after the rocket is launched; In the task planning problem, it is fixed at 64 seconds; The deployment time refers to the time required for the equipment of the rocket carrying the contact debris cleaning tool to perform necessary pre-launch work after maneuvering to the launch site, and this constraint does not exceed 45 minutes; The action height refers to the height that can be affected when the rocket carrying the contact debris cleaning tool performs the task. In the space junk cleaning task, it is determined by the initial parameters of the rocket carrying the contact debris cleaning tool; The cleaning probability refers to the probability of cleaning up the debris when the rocket carrying the contact debris cleaning tool flies near the debris; The action object is the debris that has not been cleaned by other task equipment in the task.
3. The planning platform according to claim 2, characterized in that: The reward function of the contact debris cleaning tool task planning module based on reinforcement learning is: r = r ji d ong + r en d Where: r jidong is the designed shaping reward, including the fuel consumption of the current maneuver strategy and the current mission execution situation; r end is the reward for the rocket agent equipped with a contact debris cleaning tool to clean debris; The reward for the debris is set as an immediate return reward. When a reward signal is obtained after completing a scene of training, the reward is set as follows: Among them, the cleaning judgment criterion of the contact debris cleaning tool task planning module based on reinforcement learning is whether the debris enters within the limit deviation correction distance of 5 km from the end guidance start position of the contact equipment. If the debris is within 5 km before the end guidance of the contact equipment starts, it means the debris has been destroyed.
4. The planning platform according to claim 3, characterized in that: When the non-contact laser ablation deorbiting task planning module based on reinforcement learning performs space debris cleaning, the task constraints it is subject to include maneuvering speed, retraction time, deployment time, working duration, interval time, irradiation height, and light output power; The maneuvering speed is the speed of the vehicle carrying the non-contact laser during task execution; this constraint does not exceed 30 m / s; The withdrawal time is the time required for the vehicle to withdraw and return to the station after the non-contact laser is emitted; in the mission planning problem, this constraint is fixed at 64 seconds; The deployment time is the time required for the equipment of the non-contact laser to perform the mission to move to the launch site and carry out the necessary work before launch, and this constraint does not exceed 45 minutes; The working duration is about 20 seconds from the launch to the stop of irradiation when the non-contact laser performs the mission; The interval time is the time required to switch different irradiation targets when the debris cleaning tool of the non-contact laser performs the mission, which is set according to the specific scenario; The irradiation height is the height that the non-contact laser can irradiate when the debris cleaning tool of the non-contact laser performs the mission. In the space debris cleaning mission, it is about the height of low-orbit space debris; The output power refers to the output power of the non-contact laser, and the default setting is 100 kw.
5. The planning platform according to claim 4, wherein: The reward of the non-contact laser ablation deorbiting mission planning module based on reinforcement learning is set as follows: The reward for debris is set as an immediate return reward. After completing a scene of training, a reward signal is obtained, and the reward is set as follows: Fuel consumption: The negative reward is -10; the eyelid opening reward is 100; Among them, the cleaning judgment criterion of the non-contact laser ablation deorbiting mission planning module based on reinforcement learning is whether the debris enters within 15 km of the predetermined irradiation position of the laser debris cleaning tool and the on-target power density of the non-contact laser reaches the space debris ablation deorbiting threshold. If the debris does not enter the irradiation airspace of the laser debris cleaning tool or the on-target power density of the non-contact laser is insufficient, it means that the debris has not been cleaned.
6. The planning platform according to claim 5, wherein: The specific space debris treatment is as follows: within a fixed time period, the launch vehicle is transported from the storage warehouse to the launch site at a certain maneuvering speed, and after being transported to the launch point, the deployment is completed within the specified time, that is, the necessary work before launch is carried out; When the contact capture deorbiting removal vehicle performs the cleaning mission, it can clean to a certain height. In the space debris cleaning mission, it is determined by the initial parameters of the contact capture deorbiting removal vehicle. When the contact capture deorbiting removal vehicle cleans space debris, the probability of cleaning up space debris; When the non-contact ablation-driven deorbiting laser performs the cleaning mission, after the non-contact ablation-driven deorbiting laser is emitted, the non-contact ablation-driven deorbiting laser beam is focused on the target. After having sufficient on-target power density and time accumulation, when the threshold of 0.2 w / cm2 is reached, the space debris is ablated and driven out of orbit; The action objects of the non-contact ablation-driven deorbiting laser and the contact capture deorbiting removal vehicle are space debris, and the mission completion calculation model is: Among them, Ks is the cleaning number weight, set to 0.9; Kt is the timeliness weight, set to 0.1; The laser ablation deorbiting judgment standard is that the target power density reaches the threshold value of 0.2w / cm2, and the contact capture judgment standard is whether the space debris enters the contact capture deorbit removal vehicle terminal guidance start position within the limit deviation correction distance of 5km; if the space debris is within 5km before the contact capture deorbit removal vehicle terminal guidance start, it means that the space debris is captured; Finally, after the non-contact ablation-driven deorbiting laser or contact capture deorbiting removal vehicle is launched, the vehicle retreats back to its base.
7. An intelligent planning method for emergency tasks of space debris cleaning, which is applied to the intelligent planning platform for emergency tasks of space debris cleaning according to any one of claims 1-6, wherein: The method specifically comprises the following steps: Step 1: Set up the initial space debris cleanup scenario in SpaceSim software, including the start time, ground cleanup tools, and debris; Step 2: Set the agent action space, state space and reward function in the external Python algorithm, and set the end condition of training; determine the agent neural network structure, training hyperparameters and number of training rounds; Step 3: Initialize the training neural network and target neural network parameters, and initialize the experience pool space; Step 4: Perform agent reinforcement learning according to the designer's needs. Use the plug-in algorithm to call the scenes designed in SpaceSim for each round of training. The agent makes decisions according to the initialized neural network in each state, and has a certain probability of random exploration. The state, action, reward and transfer state of each scene will be stored in the experience pool for training neural network learning; Step 5: When the experience pool meets the maximum storage space, sample learning is performed according to the set neural network gradient, and the training neural network parameters are continuously updated; after the set number of learning times, the training neural network parameters are copied to the target neural network; Step 6: For each round of training, when all the fragments are cleared or the set scene end time is reached, the round ends and step 4 is repeated; when the number of training times reaches the predetermined requirement, the training ends and the neural network and parameters are saved; Step 7: The stored data can be monitored in real time using Python and displayed in curve tables; SpaceSim data can be played back to display the two-dimensional and three-dimensional status; Step 8: Perform mission planning data statistics based on the data calculated by Python and SpaceSim to complete the planning of the mission platform.
8. An electronic device, comprising a memory and a processor, the memory storing a computer program, wherein, When the processor executes the computer program, the steps of the method described in claim 7 are implemented.
9. A computer-readable storage medium for storing computer instructions, wherein, When the computer instructions are executed by a processor, the steps of the method described in claim 7 are implemented.
Citation Information
Patent Citations
Satellite real-time guidance task planning method and system based on deep reinforcement learning
CN111950873A
Space debris removal method, device and system and storage medium
CN114750982A