A Pulse-Type Orbital Pursuit-Evasion Game Method Based on the PRD-MADDPG Algorithm
Through the PRD-MADDPG algorithm, a spacecraft pulsed orbit pursuit game model is established and reward functions are designed, and an intelligent control strategy network is trained, which solves the pulsed maneuver control problem in the spacecraft orbit pursuit game, and achieves efficient intelligent control effect.
Patent Information
- Application Number
- CN202211000653.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-08-19
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2042-08-19
AI Technical Summary
The existing technology is difficult to effectively solve the problem of pulsed maneuver control of spacecraft in space orbit pursuit and escape game, especially intelligent control under constraints such as orbital dynamics, spacecraft maneuver mode and capabilities.
The pulsed orbital pursuit and escape game method based on the PRD-MADDPG algorithm is adopted. Through modeling, designing reward functions and predictive reward detection training framework, combining the MADDPG algorithm to train the spacecraft's intelligent control strategy network, and output control instructions to realize the spacecraft's pulsed orbital pursuit and escape game.
It realizes effective intelligent control under the constraints of orbital dynamics and spacecraft motion characteristics, improves the efficiency and success rate of the spacecraft's pulsed orbital pursuit and escape game, and meets the actual needs of space missions.
Smart Images

Figure CN115320890B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of aerospace technology, in particular to the application in space orbit game, and specifically to a pulsed orbit pursuit-evasion game method based on the PRD-MADDPG algorithm. Background Technique
[0002] With the continuous development of space technology, more and more countries and institutions have carried out and participated in space activities. In 2021 alone, 40 countries and regions around the world carried out a total of 144 space launch activities, sending a total of 1,816 spacecraft into space. Facing the limited orbital resources in space, the increasing number of spacecraft also means that the orbital game in space is becoming increasingly fierce. In fact, the research on space orbit game has continuously received extensive attention from scholars. Among them, the spacecraft orbital pursuit-evasion game (OPEG) is the most common type of orbital game and is a major research hotspot in the aerospace field. Since the 1990s, a large number of scholars have carried out research on it.
[0003] Differential game theory has been widely studied and applied in the pursuit-evasion game problems of unmanned aerial vehicles and missiles. However, for the OPEG problem of spacecraft, the constraints of orbital dynamics make the solution process of traditional methods complex and the computational amount increase. In addition, because in actual space engineering tasks, the maneuverability of continuous thrust engines is very small, and currently pulsed orbital maneuvers are still the mainstream, while differential games are applicable to continuous control systems, resulting in few studies on the OPEG problem for pulsed maneuvers. Therefore, it is necessary to find a more suitable, efficient and intelligent method. The multi-agent deep deterministic policy gradient (MADDPG) algorithm can solve the competition, cooperation and cooperative game scenarios among multiple agents. In recent years, the MADDPG algorithm has been applied in tasks such as cooperative task decision-making, path planning, and cooperative hunting of multiple unmanned aerial vehicles. However, in space missions, different from other task scenarios, due to the constraints of orbital dynamics, the maneuvering mode and ability of spacecraft itself, etc., in the space orbit pursuit-evasion game problem, it is necessary to propose a targeted algorithm model and training environment construction method for these characteristics of OPEG. However, in the existing research on the pursuit-evasion problem of spacecraft in space, the vast majority of the research assumes that the maneuvering mode of spacecraft is based on continuous control, and at the same time, the control constraints of spacecraft are rarely considered. Therefore, for the space pulsed orbit pursuit-evasion game problem, effective intelligent control still cannot be achieved at present. Summary of the Invention
[0004] Aiming at the problem of spatial pulse-type orbital pursuit-evasion game in the existing technology, the present invention provides a pulse-type orbital pursuit-evasion game method based on the PRD-MADDPG algorithm, which effectively realizes the control of the pulse-type orbital pursuit-evasion game of spacecraft.
[0005] The present invention is realized through the following technical solutions:
[0006] A pulse-type orbital pursuit-evasion game method based on the PRD-MADDPG algorithm, comprising the following steps:
[0007] S1. Model the pulse-type orbital pursuit-evasion game problem to obtain a game model, and obtain the reward functions of both sides in the pulse-type orbital pursuit-evasion game according to the mission objectives of the two spacecraft in the pulse-type orbital pursuit-evasion game;
[0008] S2. Design a prediction reward detection training framework according to the game model and the reward functions of both sides in the pulse-type orbital pursuit-evasion game;
[0009] S3. Combine the prediction reward detection training framework with the MADDPG algorithm to train the pursuit-evasion game intelligent control strategy network;
[0010] S4. The pursuit-evasion game intelligent control strategy network receives the observation information of the spacecraft itself about the environment and outputs control instructions to complete the control of the pulse-type orbital pursuit-evasion game of the spacecraft.
[0011] Preferably, the process of modeling the pulse-type orbital pursuit-evasion game problem is as follows:
[0012] Design a pulse-type orbital pursuit-evasion game scenario, and select two circular orbits near the two spacecraft as reference orbits according to the relative distance between the spacecrafts relative to the orbital radius, and perform CW equation calculation.
[0013] Preferably, a spacecraft pulse-type orbital maneuver model is established under the CW equation, and the CW equation calculation formula is as follows:
[0014]
[0015] φ(t,t0)=[φ1(Δt) φ2(Δt)];
[0016] φ v (t,t i )=φ2(t-t i )=φ2(Δt);
[0017] Δv i =[Δv i,x Δv i,y Δv i,z T ;
[0018]
[0019] Among them, φ(t, t0) is the state transition matrix from time t0 to time t obtained by organizing the analytical solution of the C-W equation; Δv i represents the velocity increment vector of spacecraft i; φ v (t, t i ) represents the state transition matrix of the velocity increment part of the spacecraft from time t i to time t; N represents the total number of pulse maneuvers of the spacecraft; φ1(Δt) represents; φ2(Δt) represents; Δv i,x represents the velocity increment of spacecraft i in the x direction; Δv i,y represents the velocity increment of spacecraft i in the y direction; Δv i,z represents the velocity increment of spacecraft i in the z direction; μ is the gravitational constant, a is the orbital radius of the reference orbit; Δt represents the time interval between pulses.
[0020] Preferably, the rewards for both sides in the pulsed orbit pursuit-evasion game include a distance guidance term reward, a time reward term, a fuel consumption reward term, and a result reward term.
[0021] Furthermore, the reward functions for both sides in the pulsed orbit pursuit-evasion game are the weighted sum of the distance guidance term reward, the time reward term, the fuel consumption reward term, and the result reward term.
[0022] Preferably, the process of the prediction reward detection training framework is as follows:
[0023] S2.1. At time t i , the two spacecraft respectively make decisions based on the state information feedback by the environment and their own current policy network Actor, output the pulse control taken by the spacecraft, and change the state of the pursuer and evader spacecraft before applying the pulse control to the state of the pursuer and evader spacecraft after applying the pulse control;
[0024] S2.2. Define the time t i when the pulse control is applied as the decision point, and set up a detection point every ΔT i time between two decision points t i+1 to t d . A total of σ detection points are set. Define as the m-th detection point between the decision points [t i , t i+1 , then m ∈ [1, 2..., σ], and the size of σ is designed according to the length of the natural transfer time, the maneuverability of the spacecraft, and the size of the orbit transfer range;
[0025] S2.3. According to the CW equation, through ti The states of the chasing and evading spacecrafts before and after applying pulse control at a certain moment are calculated to obtain t i The m-th detection point after the decision point at time t state and
[0026] S2.4. Calculate the immediate reward at the detection point according to the reward functions of both sides of the pulse-type orbital chasing and evading game and the state of the predicted detection point, and calculate the cumulative predicted rewards of both spacecrafts;
[0027] S2.5. Determine whether the chasing and evading task terminates according to the state of the predicted detection point. If the chasing and evading task terminates, directly store the current environmental information, the cumulative predicted rewards of both sides, and the task termination signal into the experience pool, and the process of this task ends; if the chasing and evading task does not terminate, then determine whether this detection point is the last detection point. If this detection point is the last detection point, then transmit the current environmental information, the cumulative predicted rewards of both sides, and the signal for the task to continue to the policy networks of each spacecraft for the next decision-making. If this detection point is not the last detection point, then enter the next detection point, and repeat S2.3 to S2.5.
[0028] Preferably, the training process of the intelligent control strategy network for the chasing and evading game is as follows:
[0029] S3.1. Initialize the parameters of the policy network Actor and the evaluation network Critic of the chasing and evading spacecrafts and the state space of the spacecrafts;
[0030] S3.2. The two spacecrafts take actions according to their own observation information according to the designed prediction detection reward training framework, interact with the environment model, obtain the training data of rewards, actions, and the state space at the next moment, and store them in the replay experience pool;
[0031] S3.3. Update the parameters of the policy network Actor and the evaluation network Critic according to the MADDPG method;
[0032] S3.4. When the return reward remains within a certain range and no longer rises for a long time, stop updating and the training is completed.
[0033] Preferably, the respective policy networks Actor of the chasing and evading spacecrafts are obtained through the training of the intelligent control strategy network for the chasing and evading game. The spacecraft uses its own observation information of the environment as the input of the policy network Actor, and the output is the control instruction to be taken by the spacecraft.
[0034] Compared with the prior art, the present invention has the following beneficial technical effects:
[0035] The present invention provides a pulsed orbital pursuit-evasion game method based on the PRD-MADDPG algorithm. By modeling the pulsed orbital pursuit-evasion game problem and aiming at the mission objectives of both spacecraft in the pulsed orbital pursuit-evasion game, a reward function for both sides of the pulsed orbital pursuit-evasion game is designed. Based on the designed game model and reward function, a prediction reward detection training framework is designed. Based on the designed prediction reward detection training framework, combined with the MADDPG algorithm, the training of the pursuit-evasion game intelligent control strategy network is completed. The spacecraft uses the trained strategy network to output control commands according to its own observation information of the environment, realizing the intelligent control of the pulsed orbital pursuit-evasion game of the spacecraft. The present invention fully combines the orbital dynamics constraints and the motion characteristics of the spacecraft, establishes a pulsed orbital pursuit-evasion game model, and successfully combines the multi-agent reinforcement learning theory with the motion characteristics of the spacecraft to design the prediction reward detection multi-agent deep deterministic policy gradient algorithm (PRD-MADDPG) to solve the pulsed orbital pursuit-evasion game problem under the constraints of considering orbital dynamics, maneuvering methods and capabilities, fuel consumption, etc. This has important value in the aspect of spacecraft space orbital pursuit-evasion games. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] Figure 1 It is a flow chart of the pulsed orbital pursuit-evasion game method in the present invention;
[0037] Figure 2 It is a schematic diagram of the LVLH and ECI coordinate systems in the present invention;
[0038] Figure 3 It is a schematic diagram of the process of the space pulsed orbital pursuit-evasion game in the present invention;
[0039] Figure 4 It is a multi-agent deep reinforcement learning training framework based on prediction reward detection in the present invention;
[0040] Figure 5 It is a graph showing the change of the reward of the pursuing spacecraft with the number of training times in the present invention;
[0041] Figure 6 It is a graph showing the change of the reward of the escaping spacecraft with the number of training times in the present invention;
[0042] Figure 7 It is a graph showing the change of the pursuit success rate (per thousand times) with the number of training times in the present invention;
[0043] Figure 8 It is a pursuit-evasion trajectory graph of the pursuing spacecraft and the escaping spacecraft in the LVLH coordinate system in the present invention;
[0044] Figure 9 It is a graph showing the change of the relative distance with the mission time in the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0045] To enable those skilled in the art to better understand the solution of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0046] It should be noted that the terms "first", "second", etc. in the specification and claims of the present invention and the above-mentioned accompanying drawings are used to distinguish similar objects, and do not necessarily need to describe a specific order or sequence. It should be understood that such data used can be interchanged under appropriate circumstances so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device including a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0047] The present invention will be further described in detail below in conjunction with the accompanying drawings:
[0048] The present invention provides a pulsed orbital pursuit-evasion game method based on the PRD-MADDPG algorithm, which effectively realizes the pulsed orbital pursuit-evasion game control of spacecraft.
[0049] Specifically, the pulsed orbital pursuit-evasion game method includes the following steps:
[0050] S1. Model the pulsed orbital pursuit-evasion game problem to obtain a game model, and obtain the reward functions of both sides of the pulsed orbital pursuit-evasion game according to the mission objectives of the two spacecraft in the pulsed orbital pursuit-evasion game;
[0051] Specifically, the pulsed orbital maneuver model of the spacecraft:
[0052] The space orbital pursuit-evasion game belongs to a special relative motion between spacecraft, and the distance between the pursuer and the evader is a small quantity relative to the orbital radius. Therefore, a pulsed orbital maneuver model of the spacecraft is established under the CW equation. Since the time for applying the control by pulsed maneuver is very short compared with the orbital transfer time, it is generally considered that the pulsed control instantaneously obtains a velocity increment at the maneuver point, and the motion process between the pulsed maneuver points is the natural orbital drift of the spacecraft. The control model of the pulsed maneuver of the spacecraft can be established, and the formula is as follows:
[0053]
[0054] φ(t, t0) = [φ1(Δt) φ2(Δt)];
[0055] φ v (t, t i ) = φ2(t - t i ) = φ2(Δt);
[0056] Δv i = [Δv i,x Δv i,y T ;
[0057]
[0058] where φ(t, t0) is the state transition matrix from time t0 to time t obtained by organizing the analytical solution of the C-W equation; Δv i represents the velocity increment vector of spacecraft i; φ v (t, t i ) represents the state transition matrix of the velocity increment part of the spacecraft from time t i to time t; N represents the total number of impulsive maneuvers of the spacecraft; φ1(Δt) represents; φ2(Δt) represents; Δv i,x represents the velocity increment of spacecraft i in the x direction; Δv i,y represents the velocity increment of spacecraft i in the y direction; μ is the gravitational constant, a is the orbital radius of the reference orbit; Δt represents the time interval between impulses.
[0059] Design of impulsive orbital pursuit-evasion game scenario:
[0060] In the impulsive orbital pursuit-evasion game scenario, the relative distance between spacecraft is usually relatively close compared to the orbital radius. Circular orbits near the two spacecraft can be selected as reference orbits to satisfy the CW equation conditions. In the impulsive orbital pursuit-evasion game problem studied in this patent, the set of game participants is denoted as N = {P, E}, and the motion states of the pursuer and evader at time t in the LVLH coordinate system are respectively: and where respectively represent the position and velocity of the pursuer spacecraft in the x and y directions at time t; respectively represent the position and velocity of the evader spacecraft in the x and y directions at time t; Denote the velocity increments obtained by the pursuer and evader spacecraft at time t by applying impulsive control as and where respectively represent the velocity increments of the pursuer spacecraft in the x, y, and three directions at time t; respectively represent the velocity increments of the escaping spacecraft in the x, y, and z directions at time t. Considering that there is a time interval Δt between two pulse controls of the spacecraft in actual engineering tasks, and in the survival pursuit-evasion game, both sides will perform control to the best of their abilities. Therefore, as Figure 2 shown, it is assumed that the pulse time intervals of both the pursuer and the evader are the same, and both sides will simultaneously perform a pulse control every Δt time interval and where t i represents the time when both sides apply the i-th pulse control, i = 1, 2... n. In addition, considering the actual engineering background, it is also necessary to consider the limited maneuverability of the spacecraft, that is, there is an upper limit constraint on the velocity increment obtained by the spacecraft's single pulse control. Define the upper limits of the single velocity increments of the pursuer and the evader as that is and satisfy
[0061]
[0062] where, respectively represent the absolute values of the velocity increments of the pursuing spacecraft in the x and y directions at time t i ; respectively represent the absolute values of the velocity increments of the escaping spacecraft in the x and y directions at time t i ;
[0063] In the spacecraft pursuit-evasion game, the goal of the pursuing spacecraft is to catch up with the target in the shortest time, while the goal of the escaping spacecraft is to stay as far away from the pursuing spacecraft as possible, avoid being captured, or maximize its own survival time. Therefore, the goals of both sides in the spacecraft pursuit-evasion game can be described by the following equations:
[0064]
[0065]
[0066] In the formula, T c is the time required for the pursuing spacecraft to successfully catch up with the escaping spacecraft, that is, the pursuit time. The above formula means that the goal of the pursuing spacecraft is to find a pulse control sequence that can minimize the pursuit time for itself On the contrary, the goal of the escaping spacecraft is to find a pulse control sequence that can maximize the pursuit time for itself
[0067] When the distance between the pursuing and escaping spacecrafts first satisfies the following relationship, it is determined that the pursuing task is successful:
[0068] ||r E - r P || ≤ Δr max (5)
[0069] where r P = [x P , y P , z P are the coordinates of the chasing spacecraft in the LVLH, and r E = [x E , y E , z E are the position coordinates of the escaping spacecraft, and Δr max is the maximum relative distance for determining the success of the chasing mission. The condition for determining the failure of the mission is that the chasing time exceeds the maximum time set for the mission, that is, when the following condition holds, the mission is determined to fail:
[0070]
[0071] where r P (t) represents the position coordinates of the chasing spacecraft at time t; r E (t) represents the position coordinates of the escaping spacecraft at time t, and ||r P (t) - r E (t)|| represents the distance between the chasing spacecraft and the escaping spacecraft at time t.
[0072] Specifically, the rewards for both sides in the pulsed orbital pursuit - evasion game include a distance - guiding term reward, a time reward term, a fuel - consumption reward term, and a result reward term.
[0073] Among them, the distance - guiding term r L :
[0074] Define the relative distance between the two spacecraft at time t as ΔL(t), ΔL = ||r P (t) - r E (t)||. The goal of the chasing spacecraft is to shorten the relative distance, while the escaping spacecraft is the opposite. The distance - guiding rewards for both the chasing and escaping sides are designed as follows:
[0075]
[0076]
[0077] where α l is the distance - reward coefficient, which is used to control the distance reward within a reasonable range.
[0078] The time - reward term r t ;
[0079] Since in the pursuit - evasion game task, the chasing spacecraft needs to catch up with the target as soon as possible, while the goal of the escaping spacecraft is to maximize its own survival time, the time - reward term is designed as follows:
[0080]
[0081]
[0082] Where ρ is a positive constant representing the time reward value. For the pursuing spacecraft, as long as the pursuit mission has not ended, a fixed negative reward return -ρ will be obtained at each monitoring point, while for the escaping spacecraft, on the contrary, a positive reward return will be obtained at each detection point.
[0083] Fuel consumption reward term r Δv ;
[0084] Both spacecraft also need to consider minimizing their own fuel consumption as much as possible while achieving their own goals. Therefore, the fuel consumption reward terms for both sides are designed in the reward return function as follows:
[0085]
[0086]
[0087] Where α Δv is the fuel consumption reward coefficient. By adding a negative reward return at the moment when the spacecraft applies a pulse maneuver each time, the purpose of minimizing the fuel consumption of the spacecraft is achieved.
[0088] Result reward term r done :
[0089] This term belongs to the sparse reward type. In the pursuit-evasion game task scenario of the spacecraft, different task results will correspond to different reward returns for both spacecraft. In the scenario of this paper, there are three conditions for task termination: a. Successful pursuit; b. Exceeding the maximum task duration. Next, the result reward term expressions for the pursuing and escaping sides are given:
[0090]
[0091]
[0092] Where are all positive constants, representing the result reward values of the pursuing and escaping spacecraft under different results. The positive and negative of the coefficients represent the positive and negative of the rewards. A positive reward return represents the incentive for this result, while a negative reward return represents the punishment for this result.
[0093] Reward function design
[0094] Combining the above four types of reward terms, the reward function of the spacecraft at time t is the weighted sum of these four terms, and the formula is as follows:
[0095]
[0096]
[0097] where are the reward weighting coefficients of the spacecrafts of the two sides in the pursuit and escape respectively, satisfying The focus of each spacecraft of the two sides in the mission can be changed by adjusting the weighting coefficients.
[0098] S2. Design a predicted reward detection training framework according to the game model and the reward functions of the two sides in the pulsed orbital pursuit and escape game;
[0099] Specifically, the process of the predicted reward detection training framework is as follows:
[0100] As Figure 3 and Figure 4 shown, the process of this training framework will be explained below by taking the i-th pulse control to the (i + 1)-th pulse as an example:
[0101] S2.1, State change
[0102] First, at time t i , the spacecrafts of the two sides respectively make decisions based on the state information feedback by the environment and their current policy network Actor, and output the pulse control taken by the spacecraft Then, the states of the pursuing and escaping spacecrafts respectively change from to where and respectively represent the states before and after the pulse control is applied at time t i , that is The subscripts P and E respectively represent the pursuer and the escapee.
[0103] S2.2, Set detection points
[0104] Define the time t i when the pulse control is applied as the decision point. Set a detection point every ΔT i to t i+1 between two decision points, and a total of σ detection points are set. Define d as the m-th detection point between the decision points [t , t i , t i+1 , then m ∈ [1, 2…, σ]. The size of σ needs to be designed considering factors such as the length of the natural transfer time, the maneuverability of the spacecraft, and the size of the orbital transfer range.
[0105] S2.3, Predict states
[0106] According to the analytical solution of the CW equation, only the state at time t i needs to be known t can be calculated i The m-th detection point after the decision point at a certain moment status
[0107]
[0108] where respectively represent the states of the pursuit and evasion spacecraft at the m-th detection point when, and the state transition matrix of the spacecraft from time t to time i to time, and respectively represent the states of the pursuit and evasion spacecraft after applying the velocity increment at time t i when
[0109] S2.4, Calculate the cumulative reward
[0110] First, calculate the immediate reward at the detection point according to the immediate reward formulas (15) and (16) of both spacecraft and the state of the predicted detection point Then give the cumulative predicted rewards of both spacecraft The calculation formula is as follows
[0111]
[0112] where respectively represent the cumulative predicted rewards of the pursuit and evasion spacecraft at time, γ represents the reward discount factor and respectively represent the immediate rewards of the pursuit and evasion spacecraft at time
[0113] S2.5, Determine whether the task terminates
[0114] According to the predicted state judge whether the pursuit-evasion task terminates. If it terminates, directly store the current environmental information, the cumulative predicted rewards of both sides, and the task termination signal in the experience pool, and the process of this task ends. If the task does not terminate, then judge whether this detection point is the last detection point: if it is, then pass the current environmental information, the cumulative predicted rewards of both sides, and the signal of task continuation to the policy networks of each spacecraft for the next decision-making; if not, then enter the next detection point (m = m + 1), and repeat the above steps S2.3 - S2.5
[0115] S3. Combine the predicted reward detection training framework with the MADDPG algorithm to train the pursuit-evasion game intelligent control policy network
[0116] Specifically, the training process of the intelligent control strategy network for the pursuit-evasion game is as follows:
[0117] S3.1, Initialize the parameters of the strategy network Actor and the evaluation network Critic for the spacecraft on both sides of the pursuit-evasion and the state space of the spacecraft;
[0118] S3.2, The spacecraft on both sides take actions according to the designed prediction detection reward training framework, interact with the environment model based on their own observation information, obtain the training data of rewards, actions, and the state space at the next moment, and store them in the replay experience pool;
[0119] S3.3, Update the parameters of the strategy network Actor and the evaluation network Critic according to the MADDPG method;
[0120] S3.4, When the return reward remains within a certain range and no longer rises, stop the update and complete the training.
[0121] S4. The intelligent control strategy network for the pursuit-evasion game receives the observation information of the spacecraft itself about the environment and outputs control instructions to complete the pulse-type orbital pursuit-evasion game control of the spacecraft.
[0122] Specifically, through the training of the intelligent control strategy network for the pursuit-evasion game, the respective strategy networks Actor of the spacecraft on both sides of the pursuit-evasion are obtained. The spacecraft uses its own observation information about the environment as the input of the strategy network Actor, and the output is the control instruction to be taken by the spacecraft.
[0123] Embodiment
[0124] To illustrate the effectiveness of the proposed algorithm, a 1V1 pulse-type pursuit-evasion game scenario occurring in the GEO orbital plane is used as an example to verify the effectiveness of the algorithm. First, the algorithm parameters used in the algorithm training and their physical meanings are introduced. The parameter settings of the pulse-type pursuit-evasion game scenario are shown in Table 1:
[0125]
[0126]
[0127] Table 1 Parameter table of the pulse-type 2v1 pursuit-evasion game scenario
[0128] Next, the reward function designs for two pursuing spacecraft and one evading spacecraft are given, as shown in Table 2:
[0129]
[0130] Table 2 Parameter table of the pulse-type 2v1 pursuit-evasion game scenario
[0131] The simulation environment used in the experiment is entirely written in the Python language, using the Spyder 5.05 and Anaconda 3 platforms. The deep learning environment uses TensorFlow 1.8.0 and gym 0.10.5. The computer configuration is CPU Inter i7-9700F@3.00GHz with 32GB of memory. The spacecraft observes the environmental state, obtains the control quantity according to the set control strategy, and then adjusts the control strategy using the feedback of the environment to form a closed-loop training process.
[0132] Through training, each spacecraft can obtain a set of Actor network parameters and can control according to its own observation of the environment. Next, the training results of the 1V1 pulsed pursuit-evasion game task will be presented. First is the training process, as Figure 5 and Figure 6 shown. As the number of training times of the PRD-MADDPG algorithm increases, the reward value of the pursuing spacecraft rises to about 30 and remains stable, while the reward of the escaping spacecraft decreases to about 100 and stabilizes. Combining with Figure 7 the graph of the pursuit success rate changing with the number of training times, it can be seen that the PRD-MADDPG algorithm can stabilize the pursuit success rate at about 97% after training, indicating the effectiveness and stability of the proposed algorithm.
[0133] After the PRD-MADDPG algorithm is trained, a set of game strategy networks can be obtained. Each spacecraft executes control through its own strategy network. To verify the effectiveness of the strategy network, the initial position coordinates of the pursuer are selected as m, and the initial position of the escapee is. The two sides carry out the pursuit-evasion game using the trained strategy network. The effect of the pursuit-evasion game is as Figure 8 、 Figure 9 shown, further verifying the effectiveness of the trained strategy network.
[0134] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the above embodiments, those of ordinary skill in the art should understand that: the specific implementation manners of the present invention can still be modified or equivalently replaced, and any modification or equivalent replacement that does not depart from the spirit and scope of the present invention shall be covered by the protection scope of the claims of the present invention.
Claims
1. A pulsed orbital pursuit-evasion game method based on the PRD-MADDPG algorithm, characterized in that, It includes the following steps: S1. Model the pulsed orbital pursuit-evasion game problem to obtain a game model, and obtain the reward functions for both sides of the pulsed orbital pursuit-evasion game according to the mission objectives of the two spacecraft in the pulsed orbital pursuit-evasion game; S2. Design a prediction reward detection training framework based on the game model and the reward functions for both sides of the pulsed orbital pursuit-evasion game; S3. Combine the prediction reward detection training framework with the MADDPG algorithm to train the pursuit-evasion game intelligent control strategy network; S4. The pursuit-evasion game intelligent control strategy network receives the observation information of the spacecraft itself about the environment and outputs control instructions to complete the pulsed orbital pursuit-evasion game control of the spacecraft; The process of modeling the pulsed orbital pursuit-evasion game problem is as follows: Design a pulsed orbital pursuit-evasion game scenario, and select two circular orbits near the two spacecraft as reference orbits according to the relative distance between the spacecrafts relative to the orbital radius, and perform CW equation calculations; Establish a pulsed orbital maneuver model for the spacecraft under the CW equation. The calculation formula of the CW equation is as follows: φ(t,t0)=[φ1(Δt)φ2(Δt)]; φ v (t, t i ) = φ2(t - t i ) = φ2(Δt); Δv i = [Δv i,x Δv i,y Δv i,z T ; Among them, φ(t, t0) is the state transition matrix from time t0 to time t obtained by organizing the analytical solution of the CW equation; Δv i represents the velocity increment vector of spacecraft i; φ v (t, t i ) represents the state transition matrix of the velocity increment part of the spacecraft from time t i to time t; N represents the total number of impulsive maneuvers of the spacecraft; Δv i,x represents the velocity increment of spacecraft i in the x direction; Δv i,y represents the velocity increment of spacecraft i in the y direction; Δv i,z represents the velocity increment of spacecraft i in z direction; μ is the gravitational constant, a is the orbital radius of the reference orbit; Δt represents the time interval between pulses; The process of the prediction reward detection training framework is as follows: S2.
1. At time t i The two spacecrafts respectively make decisions based on the state information feedback by the environment according to their current policy network Actor, output the impulse control adopted by the spacecrafts, and change the state of the chasing and escaping spacecrafts before applying the impulse control to the state of the chasing and escaping spacecrafts after applying the impulse control; S2.
2. Define the application time t of the pulse control i As decision points, two decision points t i to t i+1 Set up a detection point every ΔT d moments, a total of σ detection points are set, and define As the m-th detection point between the decision points [t i , t i+1 , then m ∈ [1, 2, …, σ], and the size of σ is designed according to the length of the natural transfer time, the maneuverability of the spacecraft, and the size of the orbit transfer range; S2.
3. Calculate, according to the CW equation, the states of the pursuer and evader spacecrafts before and after applying pulse control at time t i to obtain the state of the i m-th detection point after the decision point at time t and S2.
4. Calculate the immediate reward at the detection point according to the reward functions for both sides of the pulsed orbital pursuit-evasion game combined with the state of the prediction detection point, and calculate the cumulative predicted rewards of both spacecrafts; S2.
5. Judge whether the pursuit-evasion task terminates according to the state of the prediction detection point. If the pursuit-evasion task terminates, directly store the current environmental information, the cumulative predicted rewards of both sides, and the task termination signal in the experience pool, and the process of this task ends; If the pursuit-evasion task does not terminate, judge whether this detection point is the last detection point. If this detection point is the last detection point, then transmit the current environmental information, the cumulative predicted rewards of both sides, and the signal of task continuation to the policy networks of each spacecraft for the next decision-making. If this detection point is not the last detection point, then enter the next detection point, and repeat S2.3 to S2.5; The training process of the pursuit-evasion game intelligent control strategy network is as follows: S3.
1. Initialize the parameters of the policy network Actor and the evaluation network Critic of the pursuit-evasion spacecrafts and the state space of the spacecrafts; S3.
2. The two spacecrafts take actions according to the designed prediction reward detection training framework based on their own observation information, interact with the environment model, obtain training data of rewards, actions, and the state space at the next moment, and store them in the replay experience pool; S3.
3. Update the parameters of the policy network Actor and the evaluation network Critic according to the MADDPG method; S3.
4. When the return reward remains within a certain range and no longer rises for a long time, stop updating and the training is completed; The PRD-MADDPG algorithm: Prediction Reward Detection Multi-Agent Deep Deterministic Policy Gradient algorithm; The MADDPG: Multi-Agent Deep Deterministic Policy Gradient.
2. The pulsed orbital pursuit-evasion game method based on the PRD-MADDPG algorithm according to claim 1, wherein The rewards for both sides of the pulsed orbital pursuit-evasion game include distance guidance term rewards, time reward terms, fuel consumption reward terms, and result reward terms.
3. A pulsed orbital pursuit-evasion game method based on the PRD-MADDPG algorithm according to claim 2, characterized in that The reward functions for both sides of the pulsed orbital pursuit-evasion game are the weighted sum of the distance guidance term rewards, time reward terms, fuel consumption reward terms, and result reward terms.
4. A pulse-type orbital pursuit-evasion game method based on the PRD-MADDPG algorithm according to claim 1, characterized in that, Through the training of the intelligent control strategy network for the pursuit-evasion game, the respective strategy networks Actor of the spacecrafts on both sides of the pursuit and evasion are obtained. The spacecraft uses its own observation information of the environment as the input of the strategy network Actor, and the output is the control instruction to be taken by the spacecraft.
Citation Information
Patent Citations
Spacecraft pursuit intelligent orbit control method and device and storage medium
CN113311851A
Spacecraft anti-rendezvous escape pulse solving method based on deep learning
CN114115307A