Spacecraft trajectory planning method and device for many-to-one pursuit game task
By building a roundup mission model and using the MADDPG method to train the trajectory planning network, the problem of calculation time-consuming in the multi-to-one pursuit and fugitive game mission is solved, and the success rate and timeliness of spacecraft trajectory planning are improved. It is suitable for the trajectory planning of multiple roundup spacecraft on a single evasion spacecraft.
Patent Information
- Application Number
- CN202510811086.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-17
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2045-06-17
AI Technical Summary
The existing technology is difficult to effectively solve the problem of spacecraft trajectory planning in many-to-one pursuit and fugitive game missions, especially in clustered spacecraft scenarios, which take a long time to calculate, and the pulse propulsion method does not apply to traditional differential countermeasures theory.
A roundup task model is constructed, and the trajectory planning network is trained using the MADDPG method. By using the roundup party as an agent and the evasive party as an environment, the optimal roundup strategy is trained using the reward mechanism to output the spacecraft's pulse velocity change in the local horizontal and vertical coordinate system.
It improves the tracking success rate and timeliness of strategy generation in many-to-one pursuit and fugitive game missions, and is suitable for the trajectory planning of multiple round-up spacecrafts on a single evasion spacecraft.
Smart Images

Figure CN120333464A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of spacecraft trajectory planning, and in particular to a spacecraft trajectory planning method and device for a multi-to-one pursuit-evasion game task. Background Technique
[0002] With the gradual increase of space activities, the safety issue of spacecraft has attracted more and more attention from scholars in recent years. The factors threatening the safety of spacecraft include both traditional uncontrolled flying objects such as space debris and failed satellites, and non-cooperative spacecraft with active control capabilities. Usually, such problems are called spacecraft pursuit-evasion game problems, in which the two sides of the game have adversarial goals, and the pursuer and the evader respectively execute their respective control strategies to play the game so as to maximize their respective objective functions. Some researchers have carried out research work on this problem based on the theory of two-person zero-sum differential games. Some of them solve the game problem at a fixed time by deriving the linear Riccati equation, and some other researchers convert the original problem into a two-point boundary value problem based on the Pontryagin maximum principle, and then solve it by the indirect method, semi-direct method or direct method. In the indirect method, the state variables and co-state variables of the pursuer and the evader are combined as the state variables of the problem, and the shooting method is used for solving; in the semi-direct method, usually the co-state variable of one of the pursuer or the escapee is used as the state variable, so that based on the derived necessary conditions for optimality, the original bilateral problem is transformed into a unilateral optimal control problem, and then methods such as nonlinear programming algorithms are applied for solving. In the direct method, the original bilateral game problem is converted into two separate trajectory planning problems, and the approximate optimal strategy can be quickly obtained by iterative solution.
[0003] Most of the above methods are only applicable to one-to-one game problems, that is, the number of the pursuer and the evader is both one, and it is required that both sides of the game are completely rational. With the rapid development of cluster spacecraft in recent years, cluster spacecraft have received more and more attention. When the number of pursuer spacecraft increases, the above technical solutions are difficult to effectively solve, or the calculation time for a single scenario is relatively long, which is difficult to meet the requirements of future on-board autonomous real-time decision-making. On the other hand, at present, most spacecraft propulsion methods are impulse propulsion, and the above differential game theory is not applicable to such problems. Summary of the Invention
[0004] Based on this, it is necessary to provide a spacecraft trajectory planning method and device for a multi-to-one pursuit-evasion game task that can effectively improve the pursuit success rate and the timeliness of strategy generation for the above technical problems.
[0005] A spacecraft trajectory planning method for a multi-to-one pursuit-evasion game task, the method is applied to a capture task scenario, in which the scenario includes multiple spacecraft as the capturer and one spacecraft as the evader, and the method includes: Construct a pursuit mission model. The pursuit mission model uses the orbit where the evader is located as the reference orbit, introduces the local horizontal and local vertical coordinate system, takes the initial position of the evader as the origin of the coordinate system, and each pursuer is on the side length of a square environment with the initial position of the evader as the centroid; Based on the pursuit mission model, represent the state information of the evader and pursuers. The state information includes the positions and velocities of the evader and pursuers at a certain moment; Under the pursuit mission model, based on the MADDPG method, train a trajectory planning network that can output the optimal pursuit strategy for the evader. During the training process, each pursuer is regarded as an agent, and the evader is regarded as part of the environment. Each agent trains its own execution network by interacting with the environment to obtain a reward value, so as to output the optimal pursuit strategy for the evader. The pursuit strategy is the change in impulse velocity in two directions in the local horizontal and local vertical coordinate system; Obtain the current state information of the evader to be pursued and multiple pursuers implementing the pursuit. Each pursuer uses the trajectory planning network to output the next execution strategy according to its own state information and the current state information of the evader, and then obtains the trajectories of each pursuer.
[0006] In one embodiment, in the pursuit mission scenario, the evader evades the pursuit of multiple pursuers according to the state information of each pursuer and the set evasion strategy.
[0007] In one embodiment, the set evasion strategy is expressed as: ; In the above formula, represents the importance vector containing all pursuers obtained by the evader based on the state information of each pursuer, represents the evasion strategy for dealing with the i-th pursuer. Among them, the evasion strategy adopts the single-step optimal relative distance strategy based on RRDD.
[0008] In one embodiment, when training based on the MADDPG method, each agent consists of four neural networks, including an evaluation target network, an execution target network, an evaluation online network, and an execution online network; The execution online network obtains an execution action according to the current local observation data. After the execution action, it obtains the next state by interacting with the environment, and calculates the reward through a preset reward function; The execution target network generates a target action according to the next state, and the evaluation target network calculates the target Q value according to the next state and the target action; The evaluation online network calculates the current Q value based on the current global observation data and the execution actions of all agents, calculates the loss function according to the current Q value and the target Q value, and updates the parameters in the evaluation online network and the execution online network through the loss function until convergence, so as to obtain the trained evaluation online network and the execution online network; The trained evaluation online network and the execution online network are used as the trajectory planning network.
[0009] In one embodiment, during each iterative training, after updating the parameters in the evaluation online network and the execution online network, the updated parameters are also used to softly update the parameters of the evaluation target network and the execution target network at a preset frequency.
[0010] In one embodiment, the current local observation data is the current evader state information and the current pursuer's own state information; The current global observation data is the current evader state information and the states of all current pursuers.
[0011] In one embodiment, the reward function includes an approximation dense reward, an end-task reward, and an out-of-bounds penalty.
[0012] In one embodiment, the approximation dense reward is expressed as: ; In the above formula, and are the position coordinates of the agent and the evader agent in the environment respectively; The end-task reward is expressed as: ; In the above formula, represents the end-task reward value when a single agent completes an interception, represents the total number of agents that have completed the interception, is a boolean variable indicating whether the agent has reached the end condition, represents the contribution degree of the current agent in the group; The out-of-bounds penalty is expressed as: ; In the above formula, represents the penalty value when a single agent exceeds the boundary, represents the position coordinates of the agent that has gone out of bounds, It is a Boolean variable representing the agent whether the end condition has been reached.
[0013] In one embodiment, in the encirclement task model, when the distance between a certain encircler and the evader is less than or equal to 2.5% of the side length of the square environment, the encirclement task technology is involved.
[0014] This application also proposes a spacecraft trajectory planning device for the multi-to-one pursuit-evasion game task, and the device includes: An encirclement task model construction module, which is used to construct an encirclement task model. The encirclement task model takes the orbit where the evader is located as the reference orbit, introduces the local horizontal and local vertical coordinate system, takes the initial position of the evader as the origin of the coordinate system, and multiple encirclers are on the side length of the square environment with the initial position of the evader as the centroid; A state information representation module, which is used to represent the state information of the evader and the encirclers based on the encirclement task model. The state information includes the positions and velocities of the evader and the encirclers at a certain moment; A MADDPG method training module, which is used to train a trajectory planning network capable of outputting an optimal encirclement strategy for the evader based on the MADDPG method under the encirclement task model. Among them, during the training process, each encircler is regarded as an agent, the evader is regarded as a part of the environment, and each agent trains its own execution network by interacting with the environment to obtain a reward value, so as to output an optimal encirclement strategy for the evader. Among them, the encirclement strategy is the pulse velocity change amount in two directions in the local horizontal and local vertical coordinate system; A multi-encircler trajectory planning module, which is used to obtain the current state information of the evader to be encircled and multiple encirclers for implementing the encirclement. Each encircler uses the trajectory planning network to output the next execution strategy according to its own state information and the current state information of the evader, and then obtains the trajectories of each encircler.
[0015] A computer device includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, it implements the steps in the above-mentioned spacecraft trajectory planning method for the multi-to-one pursuit-evasion game task.
[0016] A computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements the steps in the above-mentioned spacecraft trajectory planning method for the multi-to-one pursuit-evasion game task.
[0017] The above spacecraft trajectory planning method and device for the multi-to-one pursuit-evasion game task, under the constructed encirclement task model, based on the MADDPG method, trains a trajectory planning network capable of outputting an optimal encirclement strategy for the evader. Among them, during the training process, each pursuer is used as an agent, and the evader is regarded as a part of the environment. Each agent trains its own execution network by interacting with the environment to obtain a reward value, so as to output an optimal encirclement strategy for the evader. Among them, the encirclement strategy is the pulse velocity change amount in two directions in the local horizontal and local vertical coordinate systems. According to the state information of each pursuer and the current state information of the evader, the trajectory planning network is used to output the next execution strategy, and then the trajectories of each pursuer are obtained. Using this method can effectively improve the tracking success rate and the timeliness of strategy generation. Description of the Drawings
[0018] Figure 1 It is a schematic flowchart of the spacecraft trajectory planning method for the multi-to-one pursuit-evasion game task in an embodiment; Figure 2 It is a schematic diagram of the multi-to-one pursuit-evasion game problem scenario of the spacecraft in an embodiment; Figure 3 It is a schematic diagram of the change of the return value during training and evaluation in a simulation experiment; Figure 4 It is a bar chart of the return value distribution of 100 random trials in a simulation experiment; Figure 5 It is a box plot of the return value distribution of 100 random trials in a simulation experiment; Figure 6 It is a schematic diagram of the trajectory of four pursuers surrounding a single evader in a simulation experiment; Figure 7 It is a schematic diagram of the X-axis pulse maneuver when four pursuers surround a single evader and execute a decision in a simulation experiment; Figure 8 It is a schematic diagram of the Y-axis pulse maneuver intention when four pursuers surround a single evader and execute a decision in a simulation experiment; Figure 9 It is a schematic diagram of the agent return value in a simulation experiment; Figure 10 It is a structural block diagram of the spacecraft trajectory planning device for the multi-to-one pursuit-evasion game task in an embodiment. Detailed Embodiments
[0019] In order to make the purpose, technical solutions and advantages of this application clearer, the following further describes this application in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described here are only used to explain this application and are not used to limit this application.
[0020] For the problem of multiple spacecrafts surrounding one spacecraft, in this application, as Figure 1 shown, a spacecraft trajectory planning method for the multi - to - one pursuit - evasion game task is provided. This method is applied to the surrounding task scenario including multiple spacecrafts as pursuers and one spacecraft as the evader, and specifically includes the following steps: Step S100: Construct a surrounding task model. The surrounding task model takes the orbit where the evader is located as the reference orbit, introduces the local - horizontal - local - vertical coordinate system, takes the initial position of the evader as the origin of the coordinate system, and each pursuer is on the side length of a square environment with the initial position of the evader as the centroid.
[0021] Step S110: Based on the surrounding task model, represent the state information of the evader and the pursuers. The state information includes the positions and velocities of the evader and the pursuers at a certain moment.
[0022] Step S120: Under the surrounding task model, based on the MADDPG method, train a trajectory planning network that can output the optimal surrounding strategy for the evader. Among them, during the training process, each pursuer is regarded as an agent, and the evader is regarded as part of the environment. Each agent interacts with the environment to obtain a reward value to train its own execution network, so as to output the optimal surrounding strategy for the evader. The surrounding strategy is the pulse velocity change amount in two directions in the local - horizontal - local - vertical coordinate system.
[0023] Step S130: Obtain the current state information of the evader to be surrounded and the multiple pursuers that implement the surrounding. Each pursuer, according to its own state information and the current state information of the evader, uses the trajectory planning network to output the next execution strategy, and then obtains the trajectories of each pursuer.
[0024] In step S100, for the surrounding task scenario of multiple - to - one spacecrafts, a model is built. Among them, the party being surrounded also evades according to the set evasion strategy. At this time, this surrounding problem is actually a surrounding game problem.
[0025] In this embodiment, without loss of generality, the problem scenario is set near the GEO orbit. The circular orbit of the two - body model at the position of the evader at the initial moment is used as the problem reference orbit, and the local - horizontal - local - vertical coordinate system (LocalVertical Local Horizontal reference frame, LVLH) is introduced. The origin of the coordinate system is located at the centroid of the evader at the initial moment, as Figure 2 shown.
[0026] As Figure 2As shown in the figure, the red circle on the left side of the figure represents the evader at the initial moment, and the dashed line is the problem reference orbit, as shown on the right side. and respectively represent the x-axis and y-axis of the LVLH coordinate system. Assume that at the initial moment, the pursuer P is located on the side of a square with a side length of centered at the evader E1, as shown by the blue circle on the right side of Figure 2. In this paper, taking the number of pursuers set to 4 as an example to illustrate this method, they are denoted as P1, P2, P3, and P4 respectively, and the number of evaders is one, denoted as E1.
[0027] In step S110, the state variables of a single spacecraft in the plane of the LVLH coordinate system are expressed as: (1) In formula (1), is the position variable, as Figure 2 shown, is the corresponding velocity variable. When the spacecraft acting as an evader is uncontrolled, in this coordinate system, its state transition equation is: (2) In formula (2), is the state transition matrix, expressed as: (3) When the spacecraft applies an impulsive maneuver, then formula (2) can be written as: (4) In formula (4), , is the acceleration vector applied at time
[0028] Meanwhile, in this method, there are also assumed conditions, including: both the pursuer and the evader obtain the position and velocity information of the other party at their respective fixed information acquisition frequencies and the observation information is accurate. The information acquisition frequencies of the two game parties are the same as the decision-making frequencies, that is, they immediately perform impulsive maneuvers after obtaining the information. Since the spacecraft cannot obtain environmental information at other times, the evaluation of the game results is only carried out at the information acquisition moment. When the distance between at least one pursuer spacecraft and the evader spacecraft is less than or equal to 2.5% of the side length of the square environment, it is considered that the interception is completed.
[0029] In this embodiment, since the spacecraft to be pursued, that is, the evader, may be an out-of-control spacecraft, in this method, it is assumed that an evasion strategy is set inside it to avoid the pursuit of multiple pursuers.
[0030] In this embodiment, in fact, the evasion strategy can adopt the existing one-to-one evasion strategy, and on this basis, expand it to multiple pursuers. Here, the evader has stronger observation ability than the pursuer agent. Its observation input includes not only its own accurate state information, but also the accurate state information of all pursuers. Therefore, the set evasion strategy is expressed as: (5) In formula (5), represents the importance vector containing all pursuers obtained by the evader based on the state information of each pursuer, represents the evasion strategy for dealing with the i-th pursuer.
[0031] Furthermore, calculate using the following formula: (6) In formula (6), represents the velocity change under the single-step optimal relative distance condition of the pursuer based on RRDD, represents the number of pursuers.
[0032] In one of the embodiments, the evasion strategy adopts the single-step optimal relative distance strategy based on RRDD. Since the CW equation is used as the dynamic equation in this method, the difference between the state variables of the pursuer and the evader can be obtained as: (7) In formula (7), the subscript p represents the pursuer, the subscript e represents the evader, and the subscript d represents the difference between the states of the two.
[0033] Assume that at time , the single-step maximum velocity changes of the pursuer and the evader are and . Therefore At any moment in the time domain , the reachable domain is an ellipse , In the time domain , the complete reachable domain is the envelope formed by the area swept by the ellipse over time, which can be called the relative reachable domain difference (RRDD). The precise calculation of RRDD is a more complex problem. In this method, only the termination moment That is, the RRDD at each perception moment. Such an assumption is because the spacecraft does not estimate the opponent during the powered - off flight phase and cannot receive external information. For any party, the Euclidean distance from the origin to the center of the RRDD at the subsequent situation awareness moment is selected as the basis for threat assessment. The center of the RRDD is: (8) Then, when one party applies thrust and attempts to change the relative position in formula (8), we can obtain: (9) In formula (9), is the change amount of in formula (7).
[0034] When the evader aims to minimize the threat of the opponent, then: (10) In formula (10), is a positive number greater than zero. Therefore, we can obtain: (11) In formula (11), is expressed as the maximum value of, and . Therefore, for the evader, its evasion strategy is: (12) Meanwhile, for the pursuer: (13) So far, formula (12) represents the evasion strategy of a single pursuer based on the single - step optimal relative distance strategy of RRDD.
[0035] In step S120, for the multi - to - one pursuit problem, each pursuer can be regarded as an agent, and the evader can be regarded as part of the environment. Multiple agents obtain reward values by interacting with the environment to train their own actor networks, so as to obtain the optimal strategy in this environment.
[0036] The MADDPG (Multi - Agent Deep Deterministic Policy Gradient) algorithm belongs to centralized training with decentralized execution (CTDE). Among them, centralized training means that the global information is used for training and evaluating the network during training, while when each agent executes an action, the input of its execution network is only its own local perception information.
[0037] In this embodiment, the agent is composed of N pursuers (taking N = 4 as an example), and the pursuer 's own local observation information is , and the output action is denoted as . The state space of each agent consists of four-dimensional information of its position and velocity in the orbital plane, and the action space is the change in pulse velocity that can be applied in the X-axis and Y-axis directions. The action components in each axis can be applied independently without affecting each other.
[0038] Specifically, the maneuvering capabilities of each spacecraft are also normalized according to their specific maneuvering capabilities, that is, when the maximum maneuvering capability is executed in the current period, the action is recorded as 1.0, and when no maneuver is executed in the current period, the action is recorded as 0.0.
[0039] In this embodiment, the current local observation data is the state information of the current evader and the state information of the current pursuer itself, and the global observation data is the state information of the current evader and the state information of all current pursuers.
[0040] In this embodiment, the observation space of a single agent consists of the state of the evader and its own state and their extrapolation information perceived at each maneuver decision moment. Specifically, at the decision moment during the operation of the scenario , the state information of the pursuer agent is , and the state information of the evader that it can observe is , then its local observation information is: (14) In formula (14), is the relative motion transfer matrix, is the decision moment, is the time vector for the agent to extrapolate based on the current state.
[0041] In this embodiment, when training based on the MADDPG method, each agent consists of four neural networks, including the evaluation target network , the execution target network , the evaluation online network and the execution online network . The execution online network obtains the execution action according to the current local observation data. After executing this action, it obtains the next state through interaction with the environment and calculates the reward through a preset reward function. The execution target network generates the target action according to the next state, and the evaluation target network Calculate the target Q value according to the next state and the target action, and evaluate the online network Calculate the current Q value according to the current global observation data and the execution actions of all agents, calculate the loss function according to the current Q value and the target Q value, and perform gradient updates on the parameters in the evaluation online network and the execution online network through the loss function until convergence, obtaining the trained evaluation online network and the execution online network, and using the trained evaluation online network and the execution online network as the trajectory planning network.
[0042] Specifically, according to the idea of centralized training, input into the network to obtain , and then according to the data in the sample pool obtained by sampling the environment, the network can obtain , so that the loss function of the network can be constructed as follows: (15) In formula (15), indicates whether the agent has ended.
[0043] In this embodiment, in each iterative training, after updating the parameters in the evaluation online network and the execution online network, the updated parameters are also used to softly update the parameters of the evaluation target network and the execution target network at a preset frequency, so that the training process is more stable, and its update process is expressed as: (16) In this embodiment, a pseudocode that can implement the above training process is also provided, as shown in Table 1: Table 1 Pseudocode of the multi-agent deep deterministic policy gradient algorithm for implementing the above training process
[0044] It should be noted that since at the initial stage of algorithm iteration, the tracker has no experience, a large number of experiences without rewards are placed in the experience replay pool, resulting in a slow convergence speed of the algorithm. Therefore, a set of curriculum learning parameters can be set at the initial stage of algorithm training, where can adjust the proportion of the expert policy when the agent makes action decisions. According to the above formula (13), the expert policy of the pursuer agent is , introducing the expert policy parameter as , then its action at this time is . The curriculum learning parameter Then the maneuverability of the evader can be adjusted , where is the true maneuverability value of the evader. At the initial stage of training, the maneuverability of the evader is set to a small value and gradually increased to its true maneuverability.
[0045] The setting of the reward function has an important impact on the convergence of the reinforcement learning algorithm. According to the above problem description, only when the reinforcement learning environment runs to the end, the pursuer agent 's reward function . Obviously, the reward function in this environment is a sparse reward, which brings challenges to the algorithm convergence. To improve the algorithm convergence, in this embodiment, the reward function of each agent is set as a linear combination of three sub-reward functions, and the three sub-reward functions are respectively the approximate dense reward , the end-task reward and the out-of-bounds penalty .
[0046] Specifically, according to the agreement on the scene end condition in the previous text, only when at least one pursuer spacecraft is less than or equal to 2.5% of the side length of the square environment from the evader spacecraft, the tracker will obtain a reward, which is not conducive to the algorithm convergence. Therefore, after each agent executes an action, an immediate dense reward can be set up to reflect the effect of the current agent's action in real time, so as to improve the algorithm convergence. Therefore, in this embodiment, for the agent , the dense reward for each of its actions is set as follows: (17) In formula (17), and are the position coordinates of the agent and the evader agent in the environment respectively.
[0047] Furthermore, according to the above scene termination condition, the end-task reward can be set as: (18) In formula (18), represents the end-task reward value when a single agent completes the interception, represents the total number of agents that have completed the interception, is a boolean variable indicating whether the agent has reached the end condition, represents the contribution degree of the current agent in the group, which is expressed as: (19) In this embodiment, to prevent the agent from frequently exceeding the scene boundary during action exploration, which may lead to a large number of invalid explorations, it is necessary to impose a penalty on the agent when it exceeds the problem boundary. The penalty for exceeding the boundary is expressed as: (20) In formula (20), represents the penalty value for a single agent when it exceeds the boundary, represents the position coordinates of the agent that has gone out of bounds, is a boolean variable indicating whether the agent has reached the end condition.
[0048] In step S130, the trained trajectory planning network is set in each chasing spacecraft, enabling it to give the optimal decision based on the current state information. When the chasing spacecraft implements according to the current optimal decision, it can reach a certain position. After executing the decision multiple times, the trajectory of the chasing party during the encirclement can be obtained.
[0049] In this article, the effectiveness of this method is also demonstrated through simulation experiments. In the simulation experiment, the orbit where the evader is located at the initial moment is used as the reference orbit, and the average angular velocity of the reference orbit is . The initial positions of the four chasers are randomly generated on , , , the four boundaries. To increase the task difficulty of the chasers as much as possible, the initial velocities of all chasers in all directions are set to 0, and the maximum velocity change of each chasing spacecraft is less than that of the evader, as shown in Table 2. Other parameter settings for the multi-agent one-versus-many game environment and the MADDPG algorithm can be seen in Tables 2 to 4. At the beginning of training, the curriculum learning parameters and are set to 1.0 and 0.2 respectively. However, as the training process progresses, every 2E4 steps, and decrease and increase by 0.2 respectively. Therefore, after 1E6 steps, the multi-agent training process will no longer involve the expert strategy and the maneuverability of the evader is set to its actual maneuverability.
[0050] Table 2 Environmental scene parameters
[0051] Table 3 Hyperparameters of the MADDPG algorithm
[0052] Table 4 Network structure parameters
[0053] According to the above hyperparameter settings, the total number of steps in the training process is 9E6 steps, that is, when the total number of steps in the environment exceeds 9E6, the algorithm stops. Figure 3 The figure shows the change of the sum of the reward values of the four tracking parties during the training process. It can be seen that starting from step 1E6, the reward values in the training and evaluation process show a gradually increasing trend. After step 7E6, the reward values obtained by training gradually stabilize around 0.0, and the reward values obtained by evaluation also gradually stabilize around 50, indicating that the algorithm has basically converged. It should be noted that the difference in reward values obtained during training and evaluation is due to the influence of exploration noise during the training process, which makes the action selection of the agent not only depend on the input of the actor network, but also increases the influence of action noise.
[0054] To further demonstrate the effect of the trained multi-agent model in the scenario, the trained model was subjected to 100 random tests. The reward function results obtained from the random tests are as follows: Figure 4 and Figure 5 As shown. Figure 4 In the figure, the horizontal axis represents the number of trials, and the vertical axis represents the sum of the agent's reward value function. Blue means that the evading party was successfully intercepted in this trial, while red means that the evading party was not successfully intercepted in this trial. Among them, the tracking party successfully captured the evading party in 85 trials, and the overall success rate was 85%. Figure 5 The corresponding box plot is shown in , where the mean return value of 100 trials is approximately 50.
[0055] In order to further illustrate the trapping effect of the trained intelligent agent, the test results of Episode=38 in the above 100 tests are selected for display. Figure 6 The figure shows the complete process of four pursuit parties encircling the evading party. Figure 7 and Figure 8 The action choices of the agent on the X-axis and Y-axis at each decision moment are shown in Figure 1, where the Y-axis is the normalized speed change. The reward value function of each agent is as follows: Figure 9 shown.
[0056] In the above-mentioned spacecraft trajectory planning method for many-to-one pursuit and escape game tasks, a multi-agent reinforcement learning environment for many-to-one pursuit and escape game is established based on a partially observable Markov decision process, and the control strategy of the pursuer is trained based on the MADDPG algorithm to obtain the pursuit decision model, i.e., the trajectory planning network. Compared with the existing technology, the success rate of many-to-one capture is improved under the condition that the pursuer has no communication.
[0057] It should be understood that although Figure 1The steps in the flowchart are shown in sequence according to the arrows, but these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order restriction for the execution of these steps, and these steps can be executed in other orders. Moreover, Figure 1 at least a part of the steps in Figure 1 may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential, but can be executed alternately or in turn with at least a part of other steps or sub-steps or stages of other steps.
[0058] In one embodiment, as Figure 10 shown, a spacecraft trajectory planning device for a multi-to-one pursuit-evasion game task is provided, including: a pursuit task model construction module 200, a state information representation module 210, a MADDPG method training module 220, and a multi-pursuer trajectory planning module 230, where: The pursuit task model construction module 200 is configured to construct a pursuit task model. The pursuit task model uses the orbit where the evader is located as a reference orbit, introduces a local horizontal and local vertical coordinate system, takes the initial position of the evader as the origin of the coordinate system, and multiple pursuers are located on the side length of a square environment with the initial position of the evader as the centroid; The state information representation module 210 is configured to represent the state information of the evader and the pursuers based on the pursuit task model. The state information includes the positions and velocities of the evader and the pursuers at a certain moment; The MADDPG method training module 220 is configured to train a trajectory planning network capable of outputting an optimal pursuit strategy for the evader based on the MADDPG method under the pursuit task model. During the training process, each of the pursuers is regarded as an agent, and the evader is regarded as a part of the environment. Each agent trains its own execution network by interacting with the environment to obtain a reward value, so as to output an optimal pursuit strategy for the evader. The pursuit strategy is the pulse velocity change amount in two directions in the local horizontal and local vertical coordinate system; The multi-pursuer trajectory planning module 230 is configured to obtain the current state information of the evader to be pursued and multiple pursuers for implementing the pursuit. Each pursuer uses the trajectory planning network to output the next execution strategy according to its own state information and the current state information of the evader, and then obtains the trajectories of each pursuer.
[0059] For the specific limitations of the spacecraft trajectory planning device for the multi-to-one pursuit-evasion game task, reference can be made to the limitations of the spacecraft trajectory planning method for the multi-to-one pursuit-evasion game task in the foregoing text, which will not be elaborated here. Each module in the above-mentioned spacecraft trajectory planning device for the multi-to-one pursuit-evasion game task can be implemented in whole or in part by software, hardware, and their combination. The above-mentioned modules can be embedded in the processor of the computer device in hardware form or be independent of it, or can be stored in the memory of the computer device in software form, so as to facilitate the processor to call and execute the operations corresponding to each of the above modules.
[0060] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.
[0061] The above-described embodiments merely represent several implementation manners of the present application. The description is relatively specific and detailed, but it should not be construed as a limitation on the scope of the invention. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the appended claims.
Claims
1. A spacecraft trajectory planning method for a multi-to-one pursuit-evasion game task, characterized in that The method is applied to the scenario of a capture mission. In the scenario, there are multiple spacecrafts acting as the capturing party and one spacecraft acting as the evading party. The method includes: Construct a capture mission model. The capture mission model uses the orbit where the evading party is located as the reference orbit, introduces the local horizontal and local vertical coordinate system, takes the initial position of the evading party as the origin of the coordinate system, and each of the capturing parties is on the side length of a square environment with the initial position of the evading party as the centroid; Based on the capture mission model, represent the state information of the evading party and the capturing parties. The state information includes the positions and velocities of the evading party and the capturing parties at a certain moment; Under the capture mission model, based on the MADDPG method, train a trajectory planning network that can output the optimal capture strategy for the evading party. Among them, during the training process, each of the capturing parties is regarded as an agent, and the evading party is regarded as a part of the environment. Each agent trains its own execution network by interacting with the environment to obtain a reward value, so as to output the optimal capture strategy for the evading party. The capture strategy is the change in impulse velocity in two directions in the local horizontal and local vertical coordinate system; Obtain the current state information of the evading party to be captured and the multiple capturing parties implementing the capture. Each of the capturing parties uses the trajectory planning network to output the next execution strategy according to its own state information and the current state information of the evading party, and then obtains the trajectories of each capturing party.
2. The spacecraft trajectory planning method for the multi-to-one pursuit-evasion game task according to claim 1, wherein In the capture mission scenario, the evading party evades the capture of multiple capturing parties according to the state information of each capturing party and the set evasion strategy.
3. The spacecraft trajectory planning method for the one-to-many pursuit-evasion game task according to claim 2, wherein The set evasion strategy is expressed as: In the above formula, represents the importance vector containing all the encircling parties obtained by the evading party based on the status information of each encircling party. represents the evasion strategy for dealing with the i-th encircling party, where the evasion strategy adopts the single-step optimal relative distance strategy based on RRDD.
4. The spacecraft trajectory planning method for the one-to-many pursuit-evasion game task according to any one of claims 1-3, characterized in that When training based on the MADDPG method, each agent consists of four neural networks, including an evaluation target network, an execution target network, an evaluation online network, and an execution online network; The execution online network obtains an execution action according to the current local observation data. After the execution action, it obtains the next state by interacting with the environment, and calculates the reward through a preset reward function; The execution target network generates a target action according to the next state, and the evaluation target network calculates the target Q value according to the next state and the target action; The evaluation online network calculates the current Q value according to the current global observation data and the execution actions of all agents, calculates the loss function according to the current Q value and the target Q value, and updates the parameters in the evaluation online network and the execution online network through the loss function until convergence, and obtains the trained evaluation online network and execution online network; Use the trained evaluation online network and execution online network as the trajectory planning network.
5. The spacecraft trajectory planning method for the one-to-many pursuit-evasion game task according to claim 4, wherein During each iteration training, after updating the parameters in the evaluation online network and the execution online network, the updated parameters are also used to softly update the parameters of the evaluation target network and the execution target network at a preset frequency.
6. The spacecraft trajectory planning method for the one-to-many pursuit-evasion game task according to claim 5, wherein The current local observation data is the current state information of the evading party and the current state information of the capturing party itself; The current global observation data is the current evader state information and the state information of all current pursuers.
7. The spacecraft trajectory planning method for the one-to-many pursuit-evasion game task according to claim 6, wherein The reward function includes an approaching dense reward, an end-task reward, and an out-of-bounds penalty.
8. The spacecraft trajectory planning method for the multi-to-one pursuit-evasion game task according to claim 7, wherein The approaching dense reward is expressed as: In the above formula, and are the position coordinates of the agent and the avoidance agent in the environment, respectively; The end-task reward is expressed as: In the above formula, represents the end - mission reward value when a single agent completes the interception, represents the total number of agents that have completed the interception, is a boolean variable indicating whether the agent has reached the end condition, represents the contribution degree of the current agent in the group; The out-of-bounds penalty is expressed as: In the above formula, represents the penalty value when a single agent exceeds the boundary, represents the position coordinates of the out-of-bounds agent, and is a boolean variable indicating whether the agent has reached the end condition.
9. The spacecraft trajectory planning method for the one-to-many pursuit-evasion game task according to claim 8, wherein In the pursuer-evader task model, when the distance between a certain pursuer and the evader is less than or equal to 2.5% of the side length of the square environment, the pursuer-evader task technology is achieved.
10. A spacecraft trajectory planning device for a multi-to-one pursuit-evasion game task, characterized in that The device includes: A pursuer-evader task model construction module for constructing a pursuer-evader task model. The pursuer-evader task model uses the orbit where the evader is located as the reference orbit, introduces a local horizontal and local vertical coordinate system, takes the initial position of the evader as the origin of the coordinate system, and multiple pursuers are located on the side length of a square environment with the initial position of the evader as the centroid; A state information representation module for representing the state information of the evader and pursuers based on the pursuer-evader task model. The state information includes the positions and velocities of the evader and pursuers at a certain moment; A MADDPG method training module for training, under the pursuer-evader task model, a trajectory planning network capable of outputting an optimal pursuer strategy for the evader based on the MADDPG method. During the training process, each pursuer is regarded as an agent, and the evader is regarded as a part of the environment. Each agent trains its own execution network by interacting with the environment to obtain a reward value, so as to output an optimal pursuer strategy for the evader. The pursuer strategy is the change in impulse velocity in two directions in the local horizontal and local vertical coordinate system; A multi-pursuer trajectory planning module for obtaining the current state information of the evader to be pursued and multiple pursuers for implementing the pursuit. Each pursuer uses the trajectory planning network to output the next execution strategy according to its own state information and the current state information of the evader, and then obtains the trajectories of each pursuer.
Citation Information
Patent Citations
Pulse type track pursuit game method based on PRD-MADDPG algorithm
CN115320890A
Pulse type track pursuit barrier cooperative game intelligent decision control method
CN116991067A
Real-time path optimization and maneuvering avoidance action planning method and device
CN118089768A
Control method, device and equipment for cooperative game of multiple spacecrafts and medium
CN119439708A
Star group orbit pursuit decision-making method based on multi-near-end reinforcement learning
CN119962403A