A space multi-objective task planning method fusing reinforcement learning and curriculum learning

CN122595828APending Publication Date: 2026-08-18DALIAN UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610829925.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-10
Publication Date
2026-08-18

AI Technical Summary

Technical Problem

[0005]本发明主要解决空间多目标任务规划中环境高维、动态性强、奖励稀疏等技术问题,提出一种融合强化学习与课程学习的空间多目标任务规划方法,在空间目标密集、高风险的空间区域,训练出稳定、高效且泛化能力强的任务序列决策策略

Benefits of technology

[0101] First, a systematic decomposition of complex tasks improves the efficiency and stability of policy training. Addressing the difficulties in reinforcement learning training caused by the high dimensionality, strong dynamism, and sparse rewards of the space task environment, a four-stage progressive curriculum is designed, including static scenarios, simple perturbations, highly dynamic scenarios, and scenarios with some observable information. Complex problems are broken down into sub-tasks of increasing difficulty. The agent first learns basic task planning logic in a deterministic environment, then gradually introduces orbital perturbations, random events, and observation noise, enabling it to progressively master dynamic prediction, replanning, and decision-making capabilities under uncertainty. This effectively alleviates the problems of low exploration efficiency and unstable policy convergence.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122595828A_ABST
    Figure CN122595828A_ABST
Patent Text Reader

Abstract

This invention relates to the field of space mission planning technology, and provides a space multi-objective mission planning method that integrates reinforcement learning and curriculum learning. The method includes: constructing a space multi-objective mission simulation environment comprising an orbital dynamics model, an objective characteristic model, and a spacecraft constraint model; designing a progressive task sequence for agent training based on a curriculum learning method; and training the agent in the simulation environment using the progressive task sequence based on a deep reinforcement learning framework, updating network parameters by collecting experience data, and finally obtaining the optimal sequence decision-making strategy. This invention solves the technical problems of low decision quality, poor training efficiency, and weak policy adaptability caused by complex environments and sparse rewards in space multi-objective mission planning by guiding reinforcement learning through curriculum learning for stable and efficient training.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of space mission planning technology, and in particular to a space multi-objective mission planning method that integrates reinforcement learning and curriculum learning. Background Technology

[0002] Space missions are becoming increasingly complex, including multi-target on-orbit servicing, asset inspection, target clearance, close-range observation, and target acquisition. The core of these missions lies in planning an efficient sequence of tasks that meets various constraints related to fuel and time within the dynamic and uncertain space environment. This is crucial for achieving autonomy and intelligence in space mission planning. In space mission planning, close-range observation is the foundation and prerequisite for multiple tasks, including target acquisition and clearance.

[0003] Currently, relevant planning methods mainly rely on optimization algorithms based on accurate models. However, in real space environments, target trajectories are constantly changing due to various perturbations, and their state information often exhibits uncertainty or partial observability. This leads to insufficient robustness and high computational cost of optimization methods based on deterministic models, making them difficult to apply online. In recent years, deep reinforcement learning, as a method capable of learning optimal decision sequences from interactions, has shown potential in this field. However, its direct application to the planning of complex space multi-objective missions also faces significant challenges: the space multi-objective mission environment is characterized by high state space dimensionality, sparse reward signals, and strong dynamics. Especially in debris-dense areas or densely packed spacecraft formations in low Earth orbit, the numerous targets and complex relative motions significantly increase the risk of collisions, placing extremely high demands on the real-time performance and safety of the planning.

[0004] Therefore, how to overcome the above difficulties and stably and efficiently train intelligent agents with strong generalization capabilities to perform autonomous task planning in complex and uncertain spatial environments is a key problem that urgently needs to be solved in this field. Summary of the Invention

[0005] This invention primarily addresses the technical challenges of high-dimensional, dynamic, and reward-sparse environments in space multi-objective mission planning. It proposes a space multi-objective mission planning method that integrates reinforcement learning and curriculum learning. In space regions with dense space targets and high risk, this method trains a stable, efficient, and highly generalizable mission sequence decision-making strategy. This strategy can be used for sequence decision-making in close-range observations of space targets in low Earth orbit regions, providing crucial support for subsequent operations such as on-orbit servicing and target clearance.

[0006] This invention provides a spatial multi-objective task planning method that integrates reinforcement learning and curriculum learning, including:

[0007] Step 1: Construct a space multi-objective mission simulation environment that includes orbital dynamics models, target characteristic models, and spacecraft constraint models;

[0008] Step 2: Based on the course learning method, design a progressive task sequence for agent training; the progressive task sequence includes a basic planning stage, a simple perturbation decision stage, a high-dynamic scene response stage, and a partial observable information processing stage.

[0009] Step 3: Based on the deep reinforcement learning framework, the agent is trained in the simulation environment using the progressive task sequence. The network parameters are updated by collecting experience data, and finally the optimal sequence decision-making strategy is obtained.

[0010] Furthermore, step 1 includes the following steps 101 to 105:

[0011] Step 101: Construct an orbital dynamics model;

[0012] Step 102: Construct the target characteristic model;

[0013] Step 103: Construct a spacecraft constraint model;

[0014] Step 104: Key elements for constructing a space multi-objective mission simulation environment;

[0015] Step 105: Construct an agent network based on the Actor-Critic framework.

[0016] Furthermore, the orbital dynamics model is used to simulate the orbital evolution of a space target, and the simulated perturbations include atmospheric drag perturbations and Earth's non-spherical gravity. Perturbations and gravitational perturbations from the Sun and Moon correspond to atmospheric perturbation accelerations, respectively. Earth's non-spherical gravitational perturbation acceleration Solar gravitational perturbation acceleration and lunar gravitational acceleration ;

[0017] Let the acceleration of the space target be... Central gravitational acceleration The perturbed motion of a spacecraft or space target in a low Earth orbit environment is described by formulas (1) to (7):

[0018] (1);

[0019] (2);

[0020] (3);

[0021] (4);

[0022] (5);

[0023] (6);

[0024] in, This represents the Earth's oblateness harmonic coefficient, indicating the main influence of the Earth's equatorial uplift. ; This represents the position vector of a spacecraft or space target in the geocentric inertial coordinate system. The magnitude of the position vector, i.e., the distance from the center of the earth; Represents the Earth's gravitational constant. ; This represents the drag coefficient, which is dimensionless and depends on the shape and surface properties of the object; it is typically taken as 2.0 to 2.2. This represents the cross-sectional area of ​​a space target in the direction of its velocity. Indicates the mass of space targets; This represents atmospheric density, which is the distance from the Earth's center. and time Complex functions; This represents the velocity vector of a spacecraft or space target relative to the rotating atmosphere. , It is the velocity of a spacecraft or space target in a geocentric inertial coordinate system. It is the Earth's rotational angular velocity vector; Indicates the Earth's reference equatorial radius. ; Represents position vector Z-axis component in the geocentric inertial coordinate system; The unit vector representing the Z-axis of the geocentric inertial coordinate system; Represents the solar gravitational constant; Represents the lunar gravitational constant; This represents the position vector of the Sun in the geocentric inertial coordinate system; This represents the position vector of the Moon in the geocentric inertial coordinate system; This represents the position vector pointing from the sun to the spacecraft or space target. ; This represents the position vector pointing from the Moon to the spacecraft or space target. ;

[0025] Integrating equation (1) from the initial state Deducing any time Target orbital state .

[0026] Furthermore, the target characteristic model is used to define the physical attributes, mission attributes, and risk attributes of space targets such as abandoned satellites, space debris, and in-orbit satellites, including:

[0027] For each space target Its complete characteristic state is , Indicates physical properties, Indicates risk attributes, The formula represents the task attribute as follows:

[0028] (7);

[0029] (8);

[0030] (9);

[0031] in, Indicates the mass of a space target; This represents the average cross-sectional area of ​​a space target, used to calculate atmospheric drag. The drag coefficient, representing a space target, depends on its shape and surface characteristics; Material density of space debris (unit: (), used to estimate the relationship between its size and mass; Indicates space debris at time The instantaneous orbital element vector includes elements such as semi-major axis, eccentricity, inclination, right ascension of ascending node, argument of perigee, and true anomaly. Indicates space debris at time Probability of collision with key assets; Indicates target priority;

[0032] Calculated based on the target's collision probability, mass, and cross-sectional area. :

[0033] (10);

[0034] in, , , The coefficients are in the range of 0 to 1. It is a normalization function.

[0035] Furthermore, the spacecraft constraint model is used to define the mission constraints of the spacecraft, including fuel budget constraints, propulsion system capability constraints, payload operation constraints, and total mission duration constraints.

[0036] The fuel budget constraint is:

[0037] (11);

[0038] in, Indicates the first The speed increment consumed by the second maneuver; This indicates the upper limit of the total speed increment corresponding to the total fuel budget for the mission;

[0039] The capability constraints of the propulsion system are:

[0040] (12);

[0041] in, Indicates the thrust vector. This indicates the engine's maximum thrust;

[0042] The load operation constraint is:

[0043] , (13);

[0044] in, This indicates the time required for each close-up observation; Indicates the shortest time required for a single operation; and These represent the positions of the observed space target and the spacecraft, respectively. Indicates the effective range of the observation payload;

[0045] The total task duration constraint is as follows:

[0046] (14);

[0047] in, Indicates the maximum allowed time window from the start of the task until it must be completed; Indicates the time taken for the task. , These represent the task end time and the task start time, respectively.

[0048] Furthermore, step 104 includes the following steps 1041 to 1043:

[0049] Step 1041: Determine the state of the simulation environment, the state of the spacecraft itself, and the mission progress and time status;

[0050] In the simulation environment, the environmental state changes with time and the agent's actions; at time... The state of the simulation environment :

[0051] (15);

[0052] in, express A set of states for a spatial target. The state of each target is the joint output of the target characteristic model and the dynamic model;

[0053] Indicates the spacecraft's own status:

[0054] (16);

[0055] in, These represent the spacecraft's position and velocity vectors in the inertial frame, respectively. Indicates the remaining available speed increment. Indicates propellant mass;

[0056] Indicates task progress and time status:

[0057] (17);

[0058] in, This indicates the time elapsed for the task. Indicates the maximum allowed task duration. Indicates the remaining time window for the task;

[0059] Step 1042: Design the actions of the intelligent agent;

[0060] At every decision-making moment The intelligent agent from the action space Choose an action : ; Indicates at time It does not perform any operations, but only conducts observations and orbit maintenance; Indicates at time Select one with a unique identifier goal Perform a close approach maneuver;

[0061] Step 1043: Design the reward function for the simulation environment;

[0062] The reward function integrates basic rewards, penalty items, and task completion rewards in a weighted manner to guide the agent to balance the constraints of fuel consumption, collision risk, and task time limit.

[0063] The reward function is key to guiding the agent's learning, and the reward function is designed as follows:

[0064]

[0065] (18);

[0066] in, time The rewards received by the intelligent agent. It is the basic reward for successfully completing a close-up observation of the target. It is a reward based on the target's priority. The earlier a high-priority target is observed, the higher the reward value. It is a penalty for the fuel consumed during the maneuver; It is a severe punishment for a spacecraft colliding with debris; It's a penalty for exceeding the task timeout; For positive values and and It is a negative value; These are adjustable weighting coefficients.

[0067] Furthermore, the intelligent agent network is based on the Actor-Critic framework, including an Actor network and a Critic network;

[0068] The Actor network is a parameterized network. Deep neural networks, which will state Mapping to action space A probability distribution on the above, to implement a random strategy : (19);

[0069] The Critic network is a parameterized network. A deep neural network that evaluates the state Next action The long-term expected cumulative reward, i.e., the state-action value function. : (20);

[0070] in, It is a discount factor used to weigh the importance of current and future rewards; The execution action is shown in formula (18). Later time The immediate environmental rewards obtained.

[0071] Furthermore, step 2 includes the following steps 201 to 204:

[0072] Step 201: The basic planning phase is designed to keep the orbital states of the target and spacecraft constant, and is used to train the agent to predict motion trends and make dynamic decisions.

[0073] In the orbital dynamics model, all orbital dynamics perturbations are turned off, and the orbital states of the target and spacecraft are completely fixed; that is, all perturbation acceleration terms except for the central gravitational force are set to 0. ;

[0074] In the target characteristic model, orbital elements Constant, physical properties and target priority coefficient It is the primary basis for decision-making, collision probability. Calculations based on fixed geometric relationships;

[0075] In the spacecraft constraint model, the upper limit of the total velocity increment corresponding to the spacecraft's fuel budget. and task duration This constitutes a core constraint;

[0076] Step 202: Configure the simple perturbation decision stage by introducing orbital perturbation forces to cause the orbital patterns of the space target to evolve, which is used to train the agent to predict motion trends and make dynamic decisions.

[0077] In the orbital dynamics model, atmospheric drag and the Earth's non-spherical gravity are introduced. Perturbation and gravitational perturbation by the Sun and Moon, i.e., in formula (1) ;

[0078] In the target characteristic model, orbital elements The time-varying beginnings, the cross-sectional area of ​​the fragments and drag coefficient pass It directly affects the orbital decay rate and has become an important factor;

[0079] In spacecraft constraint models, Constraints remain key, but agents need to learn to reserve fuel for predicted approach windows;

[0080] Step 203: Configure the high-dynamic scenario response stage to simulate scenarios of multi-target convergence and trajectory change, which is used to train the agent to perform concurrent event processing and rapid replanning;

[0081] In the orbital dynamics model, similar to the simple perturbation decision stage in step 202, all perturbation forces are maintained simultaneously;

[0082] In the target characteristic model, the target maneuver probability is defined such that the target... With constant probability Perform a single orbital maneuver to instantly change the orbital elements. Furthermore, with constant probability New targets to be observed are randomly generated in the simulation environment; , All are constants;

[0083] Step 204, the partial observable information processing stage, is configured to add observation noise to the simulation environment. It is used to train agents to perform state estimation and robust decision-making under uncertainty;

[0084] All model output states are influenced by the observed models, and the real state received by the agent is disturbed by noise:

[0085] (twenty one);

[0086] At this point, in the target characteristic model, the observed orbital elements are increased with Gaussian noise. :

[0087] , (twenty two);

[0088] Observed spatial target attributes There is an error.

[0089] Furthermore, step 3 includes the following process:

[0090] The training begins with the basic planning phase and then executes four task sequences in sequence: the simple perturbation decision-making phase, the high-dynamic scenario response phase, and the partial observable information processing phase.

[0091] In the basic stage of static scene planning If at any time , , Compared to Update only time progress The remaining state components remain unchanged; if at time , That is, the target After performing a close-range observation operation, the required velocity increment needs to be calculated based on the orbital maneuvering model. Update spacecraft status: , Accordingly, the spacecraft orbital elements are updated to match the target. Intersection; Update Target The status is set to 'Observed', and it is moved from the pending task set to the completed list;

[0092] During the simple perturbation decision phase, the orbital elements of the target and spacecraft evolve according to formula (23):

[0093] (twenty three);

[0094] In the simple perturbation decision-making stage For the new states of all remaining targets and spacecraft after the action, a numerical integrator needs to be used to advance the time step. Calculate the orbital elements of all objects at the new moment;

[0095] During the high-dynamic scenario response phase, the orbital elements of the target and spacecraft still evolve according to formula (23); After one time step, all target trajectories are first updated according to the dynamic formula; then, sudden maneuvering collision events are checked and handled: for each target, according to probability... Determine if a maneuver has occurred; if so, modify the target's trajectory elements according to a preset maneuver model. Next, process the new target event: using probability. Add a new target to the target list, with its attributes. And the initial orbital elements are randomly generated; accordingly, Target set components Update the target's status;

[0096] During the partial observable information processing stage, the orbital elements of the target and spacecraft still evolve according to formula (23), with each target evolving probabilistically. A maneuver occurs, and the target's orbital elements are modified according to a pre-defined maneuver model. At the same time, with probability Add a new target to the target list, with its attributes. The initial orbital elements are generated randomly;

[0097] The most significant characteristic of this stage is that the state received by the agent is noisy and incomplete observation. ; , ;

[0098] In each course phase, the agent interacts with the environment within the simulation environment of that phase; at each time step... The policy network is based on the current state of the environment. Select Action The environmental state changes according to the state transition function of the current task stage. The agent receives a reward based on the reward function designed in step 1043. Collect experience data And use this data to update the parameters of the Actor network and the Critic network;

[0099] In recent times, surveillance intelligent agents The average cumulative reward over a training round, when it is consecutive ( The average value of the rounds exceeded the preset threshold. And the task success rate simultaneously reached the threshold. When the current training phase is deemed complete, the system automatically switches to the next task phase.

[0100] Compared with existing technologies, this invention proposes a spatial multi-objective task planning method that integrates reinforcement learning and curriculum learning, which has the following beneficial effects.

[0101] First, a systematic decomposition of complex tasks improves the efficiency and stability of policy training. Addressing the difficulties in reinforcement learning training caused by the high dimensionality, strong dynamism, and sparse rewards of the space task environment, a four-stage progressive curriculum is designed, including static scenarios, simple perturbations, highly dynamic scenarios, and scenarios with some observable information. Complex problems are broken down into sub-tasks of increasing difficulty. The agent first learns basic task planning logic in a deterministic environment, then gradually introduces orbital perturbations, random events, and observation noise, enabling it to progressively master dynamic prediction, replanning, and decision-making capabilities under uncertainty. This effectively alleviates the problems of low exploration efficiency and unstable policy convergence.

[0102] Second, it enables high-fidelity capability transfer from simulation to real-world tasks. The course phase is deeply coupled with the physical model parameters in the high-fidelity simulation environment. By progressively activating models such as perturbation forces, event probabilities, and observation noise, the task complexity is increased, ensuring that the evolution of the training environment conforms to real physical laws and task characteristics. The capability structure developed by the agent in the course remains consistent with the requirements of real tasks, achieving a smooth transfer of strategies from idealized simulation to a high-fidelity simulation environment. This narrows the gap between simulation and practical application, improving the practicality and deployment reliability of the strategies.

[0103] Third, establish a standardized and reproducible automated training process. Course phase transitions are linked to quantitative indicators such as average cumulative reward and task success rate, enabling objective evaluation and automated progression of the training process, reducing reliance on human experience. This mechanism ensures the agent fully grasps its current capabilities before entering higher difficulty levels, preventing training failures, while also improving the engineering efficiency of the training process and the consistency of policy output.

[0104] In summary, this invention significantly improves training efficiency, policy robustness, and generalization ability in complex spatial multi-objective programming through the deep integration of course learning and reinforcement learning, providing an effective technical solution for the efficient training of autonomous planning agents in dynamic and uncertain environments. Attached Figure Description

[0105] Figure 1 This is a flowchart illustrating the implementation of the spatial multi-objective task planning method that integrates reinforcement learning and curriculum learning provided by this invention.

[0106] Figure 2 is a flowchart of a reinforcement learning simulation environment for low Earth orbit provided by an embodiment of the present invention;

[0107] Figure 3 This is a flowchart of a deep reinforcement learning algorithm used in an embodiment of the present invention;

[0108] Figure 4 is a flowchart of a progressive task sequence for agent training based on a course learning method provided in an embodiment of the present invention.

[0109] Figure 5 This is a flowchart of a deep reinforcement learning algorithm training method that integrates course learning, provided by an embodiment of the present invention. Detailed Implementation

[0110] To make the technical problems solved by this invention, the technical solutions adopted, and the technical effects achieved clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative of the invention and not intended to limit it. Furthermore, it should be noted that, for ease of description, only the parts relevant to the invention are shown in the accompanying drawings, not all of them.

[0111] like Figure 1 As shown in the figure, an embodiment of the present invention provides a spatial multi-objective task planning method that integrates reinforcement learning and curriculum learning, comprising:

[0112] Step 1: Construct a space multi-objective mission simulation environment that includes orbital dynamics model, target characteristic model, and spacecraft constraint model.

[0113] A high-fidelity simulation environment is constructed to model a multi-target observation mission in low Earth orbit. This simulation environment serves as a virtual platform for intelligent agents to interact and learn. The intelligent agents refer to the spacecraft used for movement and observation within the simulation environment. The simulation environment includes orbital dynamics models, target characteristic models, and spacecraft constraint models. These models simulate the physical constraints of real space, ensuring the fidelity and reliability of the simulation environment.

[0114] like Figure 2 As shown, step 1 includes the following steps 101 to 105:

[0115] Step 101: Construct an orbital dynamics model.

[0116] The orbital dynamics model is used to simulate the orbital evolution of a space target. The perturbations simulated include atmospheric drag perturbations and Earth's non-spherical gravity. Perturbations and gravitational perturbations from the Sun and Moon correspond to atmospheric perturbation accelerations, respectively. Earth's non-spherical gravitational perturbation acceleration Solar gravitational perturbation acceleration and lunar gravitational acceleration .

[0117] Let the acceleration of the space target be... Central gravitational acceleration The perturbed motion of a spacecraft (intelligent agent) or space target in a low Earth orbit environment can be described by formulas (1) to (7):

[0118] (1);

[0119] (2);

[0120] (3);

[0121] (4);

[0122] (5);

[0123] (6);

[0124] in, This represents the Earth's oblateness harmonic coefficient, indicating the main influence of the Earth's equatorial uplift. ; This represents the position vector of a spacecraft or space target in the geocentric inertial coordinate system. The magnitude of the position vector, i.e., the distance from the center of the earth; Represents the Earth's gravitational constant. ; This represents the drag coefficient, which is dimensionless and depends on the shape and surface properties of the object; it is typically taken as 2.0 to 2.2. This represents the cross-sectional area of ​​a space target in the direction of its velocity. Indicates the mass of space targets; This represents atmospheric density, which is the distance from the Earth's center. and time Complex functions; This represents the velocity vector of a spacecraft or space target relative to the rotating atmosphere. , It is the velocity of a spacecraft or space target in a geocentric inertial coordinate system. It is the Earth's rotational angular velocity vector; Indicates the Earth's reference equatorial radius. ; Represents position vector Z-axis component in the geocentric inertial coordinate system; The unit vector representing the Z-axis of the geocentric inertial coordinate system; Represents the solar gravitational constant; Represents the lunar gravitational constant; This represents the position vector of the Sun in the geocentric inertial coordinate system; This represents the position vector of the Moon in the geocentric inertial coordinate system; This represents the position vector pointing from the sun to the spacecraft or space target. ; This represents the position vector pointing from the Moon to the spacecraft or space target. .

[0125] The physical parameters of space targets are provided by the target characteristic model, while the physical and constraint parameters of spacecraft are provided by the spacecraft constraint model.

[0126] When integrating the equation (1) during the application, starting from the initial state... Deducing any time Target orbital state .

[0127] Step 102: Construct the target characteristic model.

[0128] The target characteristic model is used to define the physical attributes, mission attributes, and risk attributes of space targets such as abandoned satellites, space debris, and in-orbit satellites, including:

[0129] For each space target Its complete characteristic state is , Indicates physical properties, Indicates risk attributes, This indicates the task attribute. The formula is as follows:

[0130] (7);

[0131] (8);

[0132] (9);

[0133] in, Indicates the mass of a space target; This represents the average cross-sectional area of ​​a space target, used to calculate atmospheric drag. The drag coefficient, representing a space target, depends on its shape and surface characteristics; Material density of space debris (unit: (), used to estimate the relationship between its size and mass; Indicates space debris at time The instantaneous orbital element vector includes elements such as semi-major axis, eccentricity, inclination, right ascension of ascending node, argument of perigee, and true anomaly. Indicates space debris at time Probability of collision with key assets; Indicates target priority;

[0134] Calculated based on the target's collision probability, mass, and cross-sectional area. :

[0135] (10);

[0136] in, , , The coefficients are in the range of 0 to 1. It is a normalization function. Target priority. The higher the value, the higher the observation priority. This design allows spacecraft to prioritize avoiding targets with a high probability of collision, while simultaneously observing high-threat space targets with large mass and cross-sectional area.

[0137] The output of the target characteristic model is an important part of the space target and is used to construct the simulation environment state. An important component.

[0138] Step 103: Construct a spacecraft constraint model.

[0139] The spacecraft constraint model is used to define the mission constraints of the spacecraft, including fuel budget constraints, propulsion system capability constraints, payload operation constraints, and total mission duration constraints.

[0140] The fuel budget constraint is:

[0141] (11);

[0142] in, Indicates the first The speed increment consumed by the second maneuver; This indicates the upper limit of the total speed increment corresponding to the total fuel budget of the mission.

[0143] The capability constraints of the propulsion system are:

[0144] (12);

[0145] in, Indicates the thrust vector. This indicates the engine's maximum thrust.

[0146] The load operation constraint is:

[0147] , (13);

[0148] in, This indicates the time required for each close-up observation; Indicates the shortest time required for a single operation; and These represent the positions of the observed space target and the spacecraft, respectively. Indicates the effective range of the observation payload (such as a high-definition camera).

[0149] The total task duration constraint is as follows:

[0150] (14);

[0151] in, This represents the maximum allowed time window from the start of the task until it must be completed. Indicates the time taken for the task. , These represent the task end time and the task start time, respectively.

[0152] The fuel budget constraint, propulsion system capability constraint, payload operation constraint, and total mission duration constraint together limit the agent's action space. The action / decision sequence mentioned in subsequent steps must satisfy all the conditions defined in this model.

[0153] Step 104: Key elements for constructing a space multi-objective mission simulation environment.

[0154] In addition to the models mentioned above, within the reinforcement learning framework, the simulation environment possesses key elements such as state, action, policy, and reward function, and the process is as follows: Figure 2 Step 104 includes the following steps 1041 to 1043:

[0155] Step 1041: Determine the state of the simulation environment, the state of the spacecraft itself, and the mission progress and time status.

[0156] In the simulation environment, the environmental state changes with time and the actions of the agent. At time... The state of the simulation environment :

[0157] (15);

[0158] in, express A set of states for a spatial target. The state of each target is the joint output of the target characteristic model and the dynamic model.

[0159] Indicates the spacecraft's own status:

[0160] (16);

[0161] in, These represent the spacecraft's position and velocity vectors in the inertial frame, respectively. Indicates the remaining available speed increment. Indicates the mass of the propellant.

[0162] Indicates task progress and time status:

[0163] (17);

[0164] in, This indicates the time elapsed for the task. Indicates the maximum allowed task duration. Indicates the remaining time window for the task.

[0165] Step 1042: Design the actions of the intelligent agent.

[0166] The simulation environment applies discrete actions. At each decision-making moment... The intelligent agent from the action space Choose an action : . Indicates at time It does not perform any operations, but only conducts observations and orbit maintenance; Indicates at time Select one with a unique identifier goal Perform a close approach maneuver.

[0167] Step 1043: Design the reward function for the simulation environment.

[0168] The reward function integrates basic rewards, penalty items, and task completion rewards in a weighted manner to guide the agent to balance the constraints of fuel consumption, collision risk, and task time limit.

[0169] The reward function is key to guiding the agent's learning. In this embodiment, the reward function is designed as follows:

[0170]

[0171] (18);

[0172] in, time The rewards received by the intelligent agent. It is the basic reward for successfully completing a close-up observation of the target. It is a reward based on the target's priority. The earlier a high-priority target is observed, the higher the reward value. It is a penalty for the fuel consumed during the maneuver; It is a severe punishment for a spacecraft colliding with debris; It's a penalty for exceeding the task timeout; For positive values and and It is a negative value. The weighting coefficients are adjustable. This design transforms multiple objectives such as fuel consumption, risk reduction, and mission time limits into a single scalar reward signal.

[0173] Step 105: Construct an agent network based on the Actor-Critic framework.

[0174] The intelligent agent network is based on the Actor-Critic framework, including an Actor network and a Critic network. The Actor network is a policy network used to output action instructions based on the environmental state. The Critic network is a value network used to evaluate the long-term expected cumulative reward of performing the action suggested by the Actor network in the current state. In this embodiment, both the policy network (Actor network) and the value network (Critic network) are deep neural networks.

[0175] Specifically, such as Figure 3 An Actor network is a parameterized network. Deep neural networks, which will state Mapping to action space A probability distribution on the above implements a random policy. : (19);

[0176] The Critic network is a parameterized network. A deep neural network that evaluates the state Next action The long-term expected cumulative reward, i.e., the state-action value function. : (20);

[0177] in, It is a discount factor used to weigh the importance of current and future rewards; The execution action is shown in formula (18). Later time The immediate environmental rewards obtained.

[0178] Step 2: Based on the course learning method, design a progressive task sequence for agent training.

[0179] To address the issues of high difficulty and sparse rewards in one-time training, this embodiment designs a progressive task sequence based on a course-based learning method.

[0180] The progressive task sequence includes a basic planning phase, a simple perturbation decision-making phase, a high-dynamic scenario response phase, and a partial observable information processing phase. For example... Figure 4 The course, consisting of four progressively more difficult stages, systematically trains the agent by gradually "activating" or "modifying" specific model parameters to increase environmental complexity and uncertainty. Step 2 includes steps 201 to 204 as follows:

[0181] Step 201: The basic planning phase is designed to keep the orbital states of the target and spacecraft constant, and is used to train the agent to predict motion trends and make dynamic decisions.

[0182] In the orbital dynamics model, all orbital dynamics perturbations are disabled, and the orbital states of the target and spacecraft are completely fixed. That is, all perturbation acceleration terms except for the central gravitational force are set to 0. ;

[0183] In the target characteristic model, orbital elements Constant, physical properties and target priority coefficient It is the primary basis for decision-making, collision probability. Calculations based on fixed geometric relationships;

[0184] In the spacecraft constraint model, the upper limit of the total velocity increment corresponding to the spacecraft's fuel budget. and task duration It constitutes the core constraint.

[0185] The goal of this stage is to enable the agent to learn to evaluate the task value and maneuver costs of each objective in a completely deterministic environment. The trade-offs are focused on the logic of basic sequence planning, without having to deal with dynamic predictions.

[0186] Step 202: Configure the simple perturbation decision stage by introducing orbital perturbation forces to cause the orbital patterns of the space target to evolve, which is used to train the agent to predict motion trends and make dynamic decisions.

[0187] In the orbital dynamics model, atmospheric drag and the Earth's non-spherical gravity are introduced. Perturbation and gravitational perturbation by the Sun and Moon, i.e., in formula (1) .

[0188] In the target characteristic model, orbital elements The time-varying beginnings, the cross-sectional area of ​​the fragments and drag coefficient pass It directly affects the orbital decay rate and has become an important factor.

[0189] In spacecraft constraint models, Constraints remain key, but agents need to learn to reserve fuel for predicted approach windows (the time period during which orbital rendezvous conditions must be met to achieve close-in observations).

[0190] At this stage, the agent learns the laws of orbital dynamics, understands the long-term drift caused by perturbation, and begins to make time-series decisions, with its prediction and planning capabilities receiving initial training.

[0191] Step 203: Configure the high-dynamic scenario response stage to simulate scenarios of multi-target convergence and trajectory abrupt changes, which is used to train the agent to perform concurrent event processing and rapid replanning.

[0192] This stage simulates highly dynamic and complex environments in densely populated space target areas or high-risk orbital regions.

[0193] In the orbital dynamics model, similar to the simple perturbation decision stage in step 202, all perturbation forces are maintained simultaneously;

[0194] In the target characteristic model, the target maneuver probability is defined such that the target... With constant probability Perform a single orbital maneuver to instantly change the orbital elements. Furthermore, with constant probability New targets to be observed are randomly generated in the simulation environment. , All of these are constants. Therefore, target maneuvering collisions and emerging events are introduced.

[0195] During this stage, the intelligent agent faces non-stationary, sudden, and high-collision-risk environments. It must learn to dynamically replan, balancing the original task with emergency collision avoidance and rapid replanning for newly emerging high-value targets. Its reaction speed and multi-task coordination ability in dangerous environments are trained.

[0196] Step 204, the partial observable information processing stage, is configured to add observation noise to the simulation environment. It is used to train agents to perform state estimation and robust decision-making under uncertainty.

[0197] All model output states are influenced by the observed models, and the real state received by the agent is disturbed by noise:

[0198] (twenty one);

[0199] At this point, in the target characteristic model, the observed orbital elements are increased with Gaussian noise. :

[0200] , (twenty two);

[0201] Observed spatial target attributes There is an error.

[0202] This final stage simulates the observational uncertainties of real-world tasks. The agent cannot obtain the true and complete state of the environment. It must learn implicit state estimates based on noisy historical observation sequences and make robust decisions under uncertainty, ultimately training a policy that can adapt to the limitations of real-world observations.

[0203] The above four stages, by gradually introducing dynamic perturbation parameters, adding random events, and superimposing observation noise, constitute a progressive course sequence from complete determinism to high uncertainty, and from static to highly dynamic.

[0204] Step 3: Based on the deep reinforcement learning framework, the agent is trained in the simulation environment using the progressive task sequence. The network parameters are updated by collecting experience data, and finally the optimal sequence decision-making strategy is obtained.

[0205] Deep reinforcement learning frameworks such as Figure 5 .

[0206] The training begins with the basic planning phase and then executes four task sequences in sequence: the simple perturbation decision-making phase, the high-dynamic scenario response phase, and the partial observable information processing phase.

[0207] At each stage of the task, the state changes of the agent after taking an action vary due to different environmental scenarios.

[0208] In the basic stage of static scene planning If at time , , Compared to Update only time progress The remaining state components remain unchanged; if at time , That is, the target After performing a close-range observation operation, the required velocity increment needs to be calculated based on the orbital maneuvering model. Update spacecraft status: , Accordingly, the spacecraft orbital elements are updated to match the target. Meeting. Update target. The status is set to 'observed', and it is moved from the pending task set to the completed list.

[0209] During the simple perturbation decision phase, the orbital elements of the target and spacecraft evolve according to formula (23):

[0210] (twenty three);

[0211] In the simple perturbation decision-making stage After the action, the new states of all remaining targets and spacecraft need to be advanced by one time step using a numerical integrator. Calculate the orbital elements of all objects at the new moment.

[0212] During the high-dynamic scenario response phase, the orbital elements of the target and spacecraft still evolve according to formula (23). After one time step, all target trajectories are first updated according to the dynamic formula. Then, sudden maneuvering collision events are checked and handled: for each target, probabilistically... Determine if a maneuver has occurred; if so, modify the target's trajectory elements according to a preset maneuver model. Next, process the new target event: using probability. Add a new target to the target list, with its attributes. And the initial orbital elements are randomly generated. Correspondingly, Target set components Update the target's status.

[0213] During the partial observable information processing stage, the orbital elements of the target and spacecraft still evolve according to formula (23), with each target evolving probabilistically. A maneuver occurs, and the target's orbital elements are modified according to a pre-defined maneuver model. At the same time, with probability Add a new target to the target list, with its attributes. The initial orbital elements are generated randomly;

[0214] The most significant characteristic of this stage is that the state received by the agent is noisy and incomplete observation. . , .

[0215] In each of the above-described course phases, the agent follows the simulation environment of that phase. Figure 5 The interaction with the environment is shown in the manner described. Each time step... The policy network is based on the current state of the environment. Select Action The environmental state changes according to the state transition function of the current task stage. The agent receives a reward based on the reward function designed in step 1043. Collect empirical data And use this data to update the parameters of the Actor network and the Critic network.

[0216] In recent times, surveillance intelligent agents The average cumulative reward over a training round, when it is consecutive ( The average value of the rounds exceeded the preset threshold. And the task success rate simultaneously reached the threshold. When the current training phase is deemed complete, the system automatically switches to the next task phase. This multi-metric evaluation mechanism ensures that the agent not only learns to obtain rewards but also effectively masters the robust ability to complete phased task objectives, laying a solid foundation for learning in more complex phases. Through this learning mechanism, the agent learns stably and efficiently in a sequence of progressively more difficult tasks, ultimately acquiring optimal observation sequence decision-making strategies that can be directly applied to complex real-world scenarios.

[0217] In summary, this embodiment effectively solves the problems of difficult agent training, slow policy convergence, and weak generalization ability caused by the high dimensionality, strong dynamism, and sparse rewards in low Earth orbit space multi-target observation missions by constructing a high-fidelity simulation environment, designing a structured course learning stage, and implementing a training process based on a deep reinforcement learning framework. It provides a feasible technical approach for realizing autonomous sequence planning.

[0218] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications to the technical solutions described in the foregoing embodiments, or equivalent substitutions for some or all of the technical features, do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A spatial multi-objective task planning method integrating reinforcement learning and curriculum learning, characterized in that, include: Step 1: Construct a space multi-objective mission simulation environment that includes orbital dynamics models, target characteristic models, and spacecraft constraint models; Step 2: Based on the course learning method, design a progressive task sequence for agent training; the progressive task sequence includes a basic planning stage, a simple perturbation decision stage, a high-dynamic scene response stage, and a partial observable information processing stage. Step 3: Based on the deep reinforcement learning framework, the agent is trained in the simulation environment using the progressive task sequence. The network parameters are updated by collecting experience data, and finally the optimal sequence decision-making strategy is obtained.

2. The spatial multi-objective task planning method integrating reinforcement learning and curriculum learning according to claim 1, characterized in that, Step 1 includes the following steps 101 to 105: Step 101: Construct an orbital dynamics model; Step 102: Construct the target characteristic model; Step 103: Construct a spacecraft constraint model; Step 104: Key elements for constructing a space multi-objective mission simulation environment; Step 105: Construct an agent network based on the Actor-Critic framework.

3. The spatial multi-objective task planning method integrating reinforcement learning and curriculum learning according to claim 2, characterized in that, The orbital dynamics model is used to simulate the orbital evolution of a space target. The perturbations simulated include atmospheric drag perturbations and Earth's non-spherical gravity. Perturbations and gravitational perturbations from the Sun and Moon correspond to atmospheric perturbation accelerations, respectively. Earth's non-spherical gravitational perturbation acceleration Solar gravitational perturbation acceleration and lunar gravitational acceleration ; Let the acceleration of the space target be... Central gravitational acceleration The perturbed motion of a spacecraft or space target in a low Earth orbit environment is described by formulas (1) to (7): (1); (2); (3); (4); (5); (6); in, This represents the Earth's oblateness harmonic coefficient, indicating the main influence of the Earth's equatorial uplift. ; This represents the position vector of a spacecraft or space target in the geocentric inertial coordinate system. The magnitude of the position vector, i.e., the distance from the center of the earth; Represents the Earth's gravitational constant. ; This represents the drag coefficient, which is dimensionless and depends on the shape and surface properties of the object; it is typically taken as 2.0 to 2.

2. This represents the cross-sectional area of ​​a space target in the direction of its velocity. Indicates the mass of space targets; This represents atmospheric density, which is the distance from the Earth's center. and time Complex functions; This represents the velocity vector of a spacecraft or space target relative to the rotating atmosphere. , It is the velocity of a spacecraft or space target in a geocentric inertial coordinate system. It is the Earth's rotational angular velocity vector; Indicates the Earth's reference equatorial radius. ; Represents position vector Z-axis component in the geocentric inertial coordinate system; The unit vector representing the Z-axis of the geocentric inertial coordinate system; Represents the solar gravitational constant; Represents the lunar gravitational constant; This represents the position vector of the Sun in the geocentric inertial coordinate system; This represents the position vector of the Moon in the geocentric inertial coordinate system; This represents the position vector pointing from the sun to the spacecraft or space target. ; This represents the position vector pointing from the Moon to the spacecraft or space target. ; Integrating equation (1) from the initial state Deducing any time Target orbital state .

4. The spatial multi-objective task planning method integrating reinforcement learning and curriculum learning according to claim 3, characterized in that, The target characteristic model is used to define the physical attributes, mission attributes, and risk attributes of space targets such as abandoned satellites, space debris, and in-orbit satellites, including: For each space target Its complete characteristic state is , Indicates physical properties, Indicates risk attributes, The formula represents the task attribute as follows: (7); (8); (9); in, Indicates the mass of a space target; This represents the average cross-sectional area of ​​a space target, used to calculate atmospheric drag. The drag coefficient, representing a space target, depends on its shape and surface characteristics; Material density of space debris (unit: (), used to estimate the relationship between its size and mass; Indicates space debris at time The instantaneous orbital element vector includes elements such as semi-major axis, eccentricity, inclination, right ascension of ascending node, argument of perigee, and true anomaly. Indicates space debris at time Probability of collision with key assets; Indicates target priority; Calculated based on the target's collision probability, mass, and cross-sectional area. : (10); in, , , The coefficients are in the range of 0 to 1. It is a normalization function.

5. The spatial multi-objective task planning method integrating reinforcement learning and curriculum learning according to claim 4, characterized in that, The spacecraft constraint model is used to define the mission constraints of the spacecraft, including fuel budget constraints, propulsion system capability constraints, payload operation constraints, and total mission duration constraints. The fuel budget constraint is: (11); in, Indicates the first The speed increment consumed by the second maneuver; This indicates the upper limit of the total speed increment corresponding to the total fuel budget for the mission; The capability constraints of the propulsion system are: (12); in, Indicates the thrust vector. This indicates the engine's maximum thrust; The load operation constraint is: , (13); in, This indicates the time required for each close-up observation; Indicates the shortest time required for a single operation; and These represent the positions of the observed space target and the spacecraft, respectively. Indicates the effective range of the observation payload; The total task duration constraint is as follows: (14); in, Indicates the maximum allowed time window from the start of the task until it must be completed; Indicates the time taken for the task. , These represent the task end time and the task start time, respectively.

6. The spatial multi-objective task planning method integrating reinforcement learning and curriculum learning according to claim 5, characterized in that, Step 104 includes the following steps 1041 to 1043: Step 1041: Determine the state of the simulation environment, the state of the spacecraft itself, and the mission progress and time status; In the simulation environment, the environmental state changes with time and the agent's actions; at time... The state of the simulation environment : (15); in, express A set of states for a spatial target. The state of each target is the joint output of the target characteristic model and the dynamic model; Indicates the spacecraft's own status: (16); in, These represent the spacecraft's position and velocity vectors in the inertial frame, respectively. Indicates the remaining available speed increment. Indicates propellant mass; Indicates task progress and time status: (17); in, This indicates the time elapsed for the task. Indicates the maximum allowed task duration. Indicates the remaining time window for the task; Step 1042: Design the actions of the intelligent agent; At every decision-making moment The intelligent agent from the action space Choose an action : ; Indicates at time It does not perform any operations, but only conducts observations and orbit maintenance; Indicates at time Select one with a unique identifier goal Perform a close approach maneuver; Step 1043: Design the reward function for the simulation environment; The reward function integrates basic rewards, penalty items, and task completion rewards in a weighted manner to guide the agent to balance the constraints of fuel consumption, collision risk, and task time limit. The reward function is key to guiding the agent's learning, and the reward function is designed as follows: (18); in, time The rewards received by the intelligent agent. It is the basic reward for successfully completing a close-up observation of the target. It is a reward based on the target's priority. The earlier a high-priority target is observed, the higher the reward value. It is a penalty for the fuel consumed during the maneuver; It is a severe punishment for a spacecraft colliding with debris; It's a penalty for exceeding the task timeout; For positive values and and It is a negative value; These are adjustable weighting coefficients.

7. The spatial multi-objective task planning method integrating reinforcement learning and curriculum learning according to claim 6, characterized in that, The agent network is based on the Actor-Critic framework, including an Actor network and a Critic network; The Actor network is a parameterized network. Deep neural networks, which will state Mapping to action space A probability distribution on the above, to implement a random strategy : (19); The Critic network is a parameterized network. A deep neural network that evaluates the state Next action The long-term expected cumulative reward, i.e., the state-action value function. : (20); in, It is a discount factor used to weigh the importance of current and future rewards; The execution action is shown in formula (18). Later time The immediate environmental rewards obtained.

8. The spatial multi-objective task planning method integrating reinforcement learning and curriculum learning according to claim 7, characterized in that, Step 2 includes the following steps 201 to 204: Step 201: The basic planning phase is designed to keep the orbital states of the target and spacecraft constant, and is used to train the agent to predict motion trends and make dynamic decisions. In the orbital dynamics model, all orbital dynamics perturbations are turned off, and the orbital states of the target and spacecraft are completely fixed; that is, all perturbation acceleration terms except for the central gravitational force are set to 0. ; In the target characteristic model, orbital elements Constant, physical properties and target priority coefficient It is the primary basis for decision-making, collision probability. Calculations based on fixed geometric relationships; In the spacecraft constraint model, the upper limit of the total velocity increment corresponding to the spacecraft's fuel budget. and task duration This constitutes a core constraint; Step 202: Configure the simple perturbation decision stage by introducing orbital perturbation forces to cause the orbital patterns of the space target to evolve, which is used to train the agent to predict motion trends and make dynamic decisions. In the orbital dynamics model, atmospheric drag and the Earth's non-spherical gravity are introduced. Perturbation and gravitational perturbation by the Sun and Moon, i.e., in formula (1) ; In the target characteristic model, orbital elements The time-varying beginnings, the cross-sectional area of ​​the fragments and drag coefficient pass It directly affects the orbital decay rate and has become an important factor; In spacecraft constraint models, Constraints remain key, but agents need to learn to reserve fuel for predicted approach windows; Step 203: Configure the high-dynamic scenario response stage to simulate scenarios of multi-target convergence and trajectory change, which is used to train the agent to perform concurrent event processing and rapid replanning; In the orbital dynamics model, similar to the simple perturbation decision stage in step 202, all perturbation forces are maintained simultaneously; In the target characteristic model, the target maneuver probability is defined such that the target... With constant probability Perform a single orbital maneuver to instantly change the orbital elements. Furthermore, with constant probability New targets to be observed are randomly generated in the simulation environment; , All are constants; Step 204, the partial observable information processing stage, is configured to add observation noise to the simulation environment. It is used to train agents to perform state estimation and robust decision-making under uncertainty; All model output states are influenced by the observed models, and the real state received by the agent is disturbed by noise: (21); At this point, in the target characteristic model, the observed orbital elements are increased with Gaussian noise. : , (22); Observed spatial target attributes There is an error.

9. The spatial multi-objective task planning method integrating reinforcement learning and curriculum learning according to claim 8, characterized in that, Step 3 includes the following process: The training begins with the basic planning phase and then executes four task sequences in sequence: the simple perturbation decision-making phase, the high-dynamic scenario response phase, and the partial observable information processing phase. In the basic stage of static scene planning ; If time , , Compared to Update only time progress The remaining state components remain unchanged; if at time , That is, the target After performing a close-range observation operation, the required velocity increment needs to be calculated based on the orbital maneuvering model. Update spacecraft status: , Accordingly, the spacecraft orbital elements are updated to match the target. Intersection; Update Target The status is set to 'Observed', and it is moved from the pending task set to the completed list; During the simple perturbation decision phase, the orbital elements of the target and spacecraft evolve according to formula (23): (23); In the simple perturbation decision-making stage For the new states of all remaining targets and spacecraft after the action, a numerical integrator needs to be used to advance the time step. Calculate the orbital elements of all objects at the new moment; During the high-dynamic scenario response phase, the orbital elements of the target and spacecraft still evolve according to formula (23); After one time step, all target trajectories are first updated according to the dynamic formula; then, sudden maneuvering collision events are checked and handled: for each target, according to probability... Determine if a maneuver has occurred; if so, modify the target's trajectory elements according to a preset maneuver model. Next, process the new target event: using probability. Add a new target to the target list, with its attributes. And the initial orbital elements are randomly generated; accordingly, Target set components Update the target's status; During the partial observable information processing stage, the orbital elements of the target and spacecraft still evolve according to formula (23), with each target evolving probabilistically. A maneuver occurs, and the target's orbital elements are modified according to a pre-defined maneuver model. At the same time, with probability Add a new target to the target list, with its attributes. The initial orbital elements are generated randomly; The most significant characteristic of this stage is that the state received by the agent is noisy and incomplete observation. ; , ; In each course phase, the agent interacts with the environment within the simulation environment of that phase; at each time step... The policy network is based on the current state of the environment. Select Action The environmental state changes according to the state transition function of the current task stage. The agent receives a reward based on the reward function designed in step 1043. Collect experience data And use this data to update the parameters of the Actor network and the Critic network; In recent times, surveillance intelligent agents The average cumulative reward over a training round, when it is consecutive ( The average value of the rounds exceeded the preset threshold. And the task success rate simultaneously reached the threshold. When the current training phase is deemed complete, the system automatically switches to the next task phase.