A multi-spacecraft hunting and pursuit game decision-making method based on a group non-dependent learning strategy

By adopting a multi-spacecraft encirclement and pursuit game decision-making method based on a swarm-independent learning strategy, the problems of time constraints and limited fuel in multi-spacecraft encirclement and pursuit games are solved, realizing autonomous and intelligent maneuvering of spacecraft, improving mission success rate and reducing computational resource consumption.

CN118052290BActive Publication Date: 2026-07-21NORTHWESTERN POLYTECHNICAL UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
NORTHWESTERN POLYTECHNICAL UNIV
Filing Date
2024-03-15
Publication Date
2026-07-21

AI Technical Summary

Technical Problem

Existing technologies are insufficient to effectively address the issues of time constraints, consistent maneuverability between the two sides, and limited fuel in multi-spacecraft encirclement and pursuit games. Furthermore, reinforcement learning methods suffer from low stability and model interpretability.

Method used

By adopting a group-independent learning strategy, a multi-spacecraft encirclement and pursuit game model is constructed, an intelligent learning algorithm is designed, and behavioral rewards and acceleration/deceleration mechanisms are combined to achieve autonomous and intelligent maneuvering of spacecraft. The zero-sum game principle and deep reinforcement learning are used to optimize the game strategy of spacecraft.

Benefits of technology

It improved the success rate of multi-spacecraft pursuit and capture missions, reduced satellite learning and training time, saved computing resources, and enhanced the autonomous maneuverability of spacecraft.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118052290B_ABST
    Figure CN118052290B_ABST
Patent Text Reader

Abstract

The application discloses a kind of multi-spacecraft hunting pursuit game decision-making methods based on group non-dependent learning strategy, specifically includes following main process: with speed pulse as the basic strategy of pursuit and evasion, establish multi-spacecraft hunting pursuit game optimization mathematical model;Intelligent learning algorithm is designed based on proximal policy optimization framework, on this basis, three decision-making capabilities of pulse size selection, behavior switching and task allocation are fused;Design group non-dependent basic game behavior set, and establish reward function model with behavior reward core;Design three kinds of auxiliary game mechanisms of acceleration, semi-forced behavior switching and dynamic task allocation.The algorithm proposed in the application is guided by the bottom simple behavior, compared with the traditional intelligent learning strategy based on terminal distance, can improve the learning efficiency and quality of spacecraft, and the designed auxiliary mechanism can effectively improve the flexibility of cluster game.The application has the characteristics of simple training, strong adaptability and strong real-time performance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of spacecraft game planning technology, specifically involving a multi-spacecraft encirclement and pursuit game decision-making method based on a group-independent learning strategy. Background Technology

[0002] The increase in space debris and junk has dramatically increased the risk of space collisions. Due to limited fuel and fragile physical structures, high-value large spacecraft are not suitable for large-scale orbital maneuvers to evade these threats. Therefore, introducing chasing spacecraft to actively protect high-value spacecraft is crucial. Compared to individual spacecraft, spacecraft swarms can leverage their numerical advantage to gain a dynamic advantage, making them more suitable for space games. However, only a few studies have focused on orbital encirclement and pursuit game problems at the swarm level. Furthermore, the study of game strategies remains one of the challenges in orbital pursuit game research. Therefore, the research and design of spacecraft swarm encirclement and pursuit game strategies has significant practical implications.

[0003] Game theory strategy research can be mainly divided into two categories: differential game theory methods and reinforcement learning intelligent learning methods. Differential game theory methods are mainly applicable to games involving continuous thrust spacecraft. Reinforcement learning intelligent methods do not rely on complex mathematical models and are highly adaptable to dynamic game scenarios. However, reinforcement learning methods face challenges such as stability issues and low model interpretability. They lack the ability to conduct theoretical analysis to determine the advantages and disadvantages of the algorithm and often rely on a large number of trial-and-error experiments to evaluate the effectiveness of the model. Summary of the Invention

[0004] To address the above problems, this invention proposes a multi-spacecraft encirclement and pursuit game decision-making method based on a group-independent learning strategy, which solves the close-range encirclement and pursuit game problem with time constraints, consistent maneuverability of both sides, and limited fuel.

[0005] To solve the above-mentioned technical problems, the technical solution adopted by the present invention is as follows:

[0006] A multi-spacecraft encirclement and pursuit game decision-making method based on a group-independent learning strategy includes the following steps:

[0007] Step 1: In a multi-spacecraft pursuit and escape game scenario, establish motion models of the chasing spacecraft, monitoring spacecraft, and target spacecraft in the LVLH coordinate system of the orbital reference point to describe the relative motion process of the spacecraft; using velocity pulses as the basic strategies of both sides and relative distance as the optimization objective, consider the spacecraft velocity pulse amplitude, time constraints, and terminal game constraints, and construct a multi-spacecraft pursuit and escape game model based on the zero-sum game principle; the LVLH coordinate system represents a locally vertical and locally horizontal coordinate system.

[0008] Step 2: Design a spacecraft intelligent learning algorithm based on the near-end strategy optimization framework, construct an observation state space including the position and velocity of both the pursuer and the pursuer, construct an action state space including velocity pulse size, behavior type and target assignment sequence, and design the training process for both the pursuer and the pursuer based on the principle of left and right mutual fighting.

[0009] Step 3: Based on Step 2, design a reward function to decompose and transform the game strategies in spacecraft game, including forward flight, backward flight, rendezvous and formation flight, into six underlying behaviors that are conducive to the high dynamic game of spacecraft clusters. Combine these behaviors with terminal distance status rewards to form spacecraft training action incentives, with behavioral rewards accounting for more than 90%, ensuring that behavioral guidance is the core.

[0010] Step 4: Design an acceleration / deceleration mechanism to selectively correct the spacecraft's actions, ensuring that the chasing spacecraft can quickly reverse its disadvantage as the relative distance continues to increase; design a semi-mandatory behavior switching mechanism to ensure the adaptability of the chasing spacecraft in different game stages through flexible but infrequent behavior switching; design a dynamic task allocation mechanism to maintain a uniform distribution of resources while effectively allocating chasing tasks.

[0011] Step 5: Train and update the spacecraft pursuit and escape game decision algorithm constructed in Steps 1-4, and finally output a spacecraft encirclement and pursuit game strategy adapted to high dynamic game scenarios.

[0012] Furthermore, in step 1, the relative motion model of the spacecraft is represented as:

[0013] ;

[0014]

[0015] in, express The state of velocity and acceleration at any given moment. and It is the state transition matrix. These represent the real-time time and the initial time, respectively. These represent the sets of chasing spacecraft and the sets of escape spacecraft, respectively. The number of rounds in the game. It is the first Wheel speed increment, Indicates the current game round. Indicates the first Round number satellites The magnitude of the velocity pulse on the axis, with the superscript T indicating the transpose of the matrix; the equation describes the relative position of the spacecraft with respect to the origin of the LVLH coordinate system at a given time, without the need for control conditions;

[0016] Using a round-based game theory approach, an optimization mathematical model is established for the spacecraft encirclement and pursuit game problem, expressed as follows:

[0017] ;

[0018] in, This indicates optimizing the objective function value. This represents a game strategy for capturing spacecraft, specifically targeting escaped spacecraft. The spacecraft encirclement alliance stated that it was , and Representing sets Benefits and escape spacecraft The benefits, Indicates the maximum pulse limit. This indicates the minimum distance at which a chasing spacecraft can effectively capture an escaping spacecraft. and These represent the chasing spacecraft in the [number]th [year]. The position of the round and the escape spacecraft in the The position of the round, and These represent the sets of chasing spacecraft and the sets of escaping spacecraft, respectively. This represents the maximum number of rounds in the game. This represents the 2-norm of a matrix.

[0019] Furthermore, in step 2, a spacecraft intelligent learning algorithm following the zero-sum game principle is designed based on the near-end strategy optimization framework, and an observation state space including the position and velocity of both the pursuer and the pursuer is constructed. Construct velocity pulses including chasing and escaping spacecraft. Types of behavior in chasing spacecraft and target allocation sequence The action state space is used to solve the spacecraft's action strategy;

[0020] Spacecraft action state space Perform the following coding design: Set the velocity pulse index Types of behavior in chasing spacecraft and target allocation sequence To follow the decimals in the interval [0,1], where The behavior type chosen by the chasing spacecraft is represented and mapped to an integer in the interval [1,8] after processing the cumulative probability density. This represents the selection of one of six game behaviors. The three codes conform to Gaussian, Beta, and Beta distributions, respectively. The number of the defensive escape spacecraft assigned to the chasing spacecraft is mapped to an integer in the interval [1, m] after processing with the accumulated probability density; the impulse index This represents a pulse combination containing pulses along three directional axes, with a total of [number] pulses per directional axis. A velocity pulse increment scheme, The pulse combination consists of a constant determined by the magnitude of the pulse gradient. kind, After processing the accumulated probability density, it is mapped to the interval [1, ...]. [A set of integers, each integer corresponding to a velocity pulse increment scheme;]

[0021] Constructing the actor-critic neural network required for information transmission during the training process; the initial state of the spacecraft's actor network input. or Output action strategy or And the probability of selection, where include , include The action strategy is iterated through the spacecraft kinematics model to obtain the state space at the next moment. and and action rewards And store it in the experience pool; the spacecraft commentator network is based on the current global state. The output state value function predicts the long-term reward of the input action and calculates the time difference error, ultimately achieving a score for the selected action.

[0022] Furthermore, the behavioral reward in step 3 includes:

[0023] 1) Directly pursue behavioral rewards Encourage the chasing spacecraft to move in the same direction as the escaping spacecraft, applying pressure and directly rewarding the chasing behavior. Depends on the speed of the chasing spacecraft and escape spacecraft speed The similarity is expressed as ,in The dot product of two unit vectors is represented; the chasing spacecraft in direct pursuit behavior is approximately in a formation flight state with the escaping spacecraft;

[0024] 2) Tracking behavior rewards : Encourages the chasing spacecraft to move towards the previous position of the escaping spacecraft to achieve a tracking effect; rewards tracking behavior. Depends on the speed of the chasing spacecraft With relative distance vector The similarity in direction is represented as ;

[0025] 3) Targeted behavioral rewards : Encourage the chasing spacecraft to adjust its position to align with the escape spacecraft's direction of motion, thereby enhancing the tracking effect in subsequent rounds; pointing behavior reward Depends on the direction of the relative distance vector relative to the velocity direction of the escape spacecraft The similarity is expressed as ;

[0026] 4) Rewarding imitation behavior Encourage the imitation of the best-performing spacecraft in the group, and reward such behavior. Depends on the speed of the chasing spacecraft With the best chasing spacecraft;

[0027] 5) Rewarding the surrounding behavior The chasing spacecraft is encouraged to move in the direction in which the escaping spacecraft is most likely to escape. The velocity vectors of the chasing spacecraft and the escaping spacecraft are combined to obtain... ; Chasing spacecraft along The direction of movement reduces the angle and space for the escape spacecraft to evade; the encirclement behavior is rewarded. Depends on the speed of the chasing spacecraft direction and The similarity in direction is represented as ;

[0028] 6) Predictive behavioral rewards : Encourages chasing spacecraft to move towards the location where the escaping spacecraft might escape in the next moment; predictive behavior rewards Depends on the speed of the chasing spacecraft direction and The similarity in direction (of the spacecraft's velocity at the next moment) is expressed as: ;

[0029] In addition to behavioral rewards, terminal status rewards are added, specifically the following three settings:

[0030] 1) Distance maintenance bonus To guide the pursuing spacecraft to maintain a distance from the escaping spacecraft and prevent the pursuing spacecraft from becoming too far away from the escaping spacecraft during the pursuit process, it is represented as:

[0031] ;

[0032] Where D0 and k0 are parameters. , ;

[0033] 2) Closer Reward : Guiding the chasing spacecraft to approach the escaping spacecraft, inheriting the characteristics of traditional models to ensure training stability, is represented as:

[0034] ;

[0035] 3) Rewards for completing the mission Encouraging pursuing spacecraft to remain within the target's safe corridor until all remaining targets are pursued is represented as:

[0036] ;

[0037] Based on behavioral rewards and terminal state rewards, and combined with dynamic parameters, the spacecraft is chased. Game-theoretic escape spacecraft The reward function is constructed as follows:

[0038] ;

[0039] in, This signifies a reward for chasing spacecraft. Indicates the distance between the escaping spacecraft and the target. This indicates the distance between the escaping spacecraft and the pursuing spacecraft. Indicates chasing spacecraft Rewards for using behavior Represents logical values, It is a weighting coefficient, determined by the initial distance between the parties in the pursuit and capture; the reward for the escaping spacecraft is set as the negative of the reward for all corresponding pursuing spacecraft.

[0040] Furthermore, the acceleration / deceleration mechanism in step 4 includes:

[0041] 1) When the speed of the escaping spacecraft exceeds the speed of the pursuing spacecraft And the relative distance continues to increase. , The threshold constant is These represent the distances between the pursuer and the escapee in the current round and the previous round, respectively. A deceleration trend is used to suppress the continuous increase of the relative distance. The trend is determined by the deceleration pulse group (-M, -M, -M, -M, -M).

[0042] 2) When the speed of the escaping spacecraft exceeds the speed of the pursuing spacecraft And the relative distance continues to increase. An acceleration trend is used to suppress the continuous increase in relative distance, and the trend is determined by a deceleration pulse group (M, M, M, M, M).

[0043] Furthermore, the semi-mandatory behavior switching mechanism in step 4 is a semi-random behavior switching method, which guides the spacecraft to make reasonable behavior switching; it stipulates that the chasing spacecraft cannot make two behavior switchings within three consecutive game rounds, so that the spacecraft decides whether to change after accumulating its behavioral advantages and disadvantages over a period of time.

[0044] Furthermore, the dynamic task allocation method in step 4 includes:

[0045] If the escape spacecraft is not selected by any pursuing spacecraft, all pursuing spacecraft freely choose their targets based solely on distance. Once the escape spacecraft is selected, other pursuing spacecraft, when selecting that target, need to consider the current mission coalition. Has the quantity reached the expected value? If the target is not reached, then one is free to choose, and the specific free choice process is as follows: If chasing a spacecraft distance Smaller than the current alliance The chasing spacecraft Maximum distance from the target ,but Replace the alliance In , is represented as:

[0046] ;

[0047] When the target missions are of similar difficulty or the spacecraft are isomorphic, the expected number of mission coalitions is determined by an equal distribution principle. When the target mission difficulties are inconsistent, the expected number of mission coalitions is designed based on the spacecraft's game-theoretic efficiency, allocating more chasing spacecraft to spacecraft with stronger escape game-theoretic capabilities. Under the assumption of isomorphic spacecraft, an equal distribution mechanism is adopted, meaning the number of coalitions per mission is... In cases where the value cannot be divided evenly, the value of each task alliance is determined manually or through some mechanism. .

[0048] Beneficial effects:

[0049] The game-theoretic decision-making method in this invention, through a deep reinforcement learning process, can ensure that spacecraft have autonomous and intelligent maneuverability, effectively improving the success rate of multi-spacecraft encirclement and pursuit missions. The deep reinforcement learning reward function in this invention can effectively reduce the learning and training time and difficulty of satellites, and can effectively save computing resources. Attached Figure Description

[0050] Figure 1 The flowchart shows a multi-spacecraft encirclement and pursuit game decision-making method based on a group-independent learning strategy according to the present invention.

[0051] Figure 2 This is a schematic diagram of the learning and training based on the proximal policy optimization framework in this invention;

[0052] Figure 3 This is a schematic diagram of the maneuvering trajectory of three chasing spacecraft surrounding an escaping spacecraft, as planned in this invention. Specific implementation methods

[0053] like Figure 1 As shown, the present invention provides a multi-spacecraft encirclement and pursuit game decision-making method based on a group-independent learning strategy, comprising the following steps:

[0054] Step 1: Using velocity pulses as the basic strategy for both the pursuer and the pursuer, establish an optimization mathematical model for a multi-spacecraft encirclement and pursuit game.

[0055] Step 2: Design an intelligent learning algorithm based on the proximal policy optimization framework, integrating three decision-making capabilities: impulse size selection, behavior switching, and task allocation;

[0056] Step 3: Design eight types of group-independent basic game behaviors and establish a reward function model with behavior rewards as the core;

[0057] Step 4: Design three auxiliary game theory mechanisms: semi-deceleration, semi-forced behavior switching, and dynamic task allocation.

[0058] Specifically, step 1 includes:

[0059] In a multi-spacecraft encirclement and pursuit game scenario, a motion model of the pursuing spacecraft and the escaping spacecraft in the LVLH coordinate system of the orbital reference point is established to describe the relative motion process of the spacecraft.

[0060] The relative motion model of the spacecraft is represented as follows:

[0061] ;

[0062] ;

[0063] in, express The state of velocity and acceleration at any given moment. and It is the state transition matrix. These represent the real-time time and the initial time, respectively. These represent the sets of chasing spacecraft and the sets of escape spacecraft, respectively. The number of rounds in the game. It is the first Wheel speed increment, It is the first The velocity increment of the wheel, with the superscript T denoteing the transpose of the matrix. The equations can be used to describe the relative position of the spacecraft with respect to the origin of a Local Vertical-Local Horizontal (LVLH) coordinate system at a given moment, without the need for control conditions.

[0064] Using velocity pulses as the basic strategies of both sides in the game and relative distance as the optimization objective, this paper constructs a multi-spacecraft encirclement and pursuit game model based on the zero-sum game principle, taking into account the spacecraft velocity pulse amplitude, time constraints, and terminal game constraints.

[0065] Using a round-based game theory approach, an optimization mathematical model is established for the spacecraft encirclement and pursuit game problem, specifically expressed as follows:

[0066] ;

[0067] in, This indicates optimizing the objective function value. Indicates the current game round. Indicates the first Round number satellite Magnitude of velocity pulses on the axis This represents a game strategy for capturing spacecraft, specifically targeting escaped spacecraft. The spacecraft encirclement alliance stated that it was , and Representing sets Benefits and escape spacecraft The benefits, Indicates the maximum pulse limit. This indicates the minimum distance at which a chasing spacecraft can effectively capture an escaping spacecraft. and These represent the chasing spacecraft in the [number]th [year]. The position of the round and the escape spacecraft in the The position of the round, and These represent the sets of chasing spacecraft and the sets of escaping spacecraft, respectively. This represents the maximum number of rounds in the game. This represents the 2-norm of a matrix.

[0068] Specifically, step 2 includes:

[0069] First, an observation state space is constructed, including the three-axis positions and velocities of both the pursuing and fleeing parties. ;

[0070] Then, an action state space is constructed, including velocity pulse size, behavior type, and target assignment sequence. ;

[0071] The encoding method is as follows:

[0072] Set speed pulse Types of behavior in chasing spacecraft and target allocation sequence To follow the decimal range [0,1];

[0073] in, The behavior type chosen by the chasing spacecraft is represented and mapped to an integer in the interval [1,6] after processing the cumulative probability density. This represents the selection of one of six game behaviors. The three codes conform to Gaussian, Beta and Beta distributions, respectively.

[0074] This represents the escape spacecraft number assigned to the chasing spacecraft, which is then mapped to an integer in the interval [1, m] after processing with the accumulated probability density.

[0075] Pulse Amplitude Index This indicates a pulse combination containing pulses along three directional axes, with a total of [number] pulses per axis. A velocity pulse increment scheme, the pulse combination includes a total of kind, It is a constant determined by the magnitude of the pulse gradient. After processing the accumulated probability density, it is mapped to the interval [1, ...]. [] is an integer, and each integer corresponds to a velocity pulse increment scheme.

[0076] After constructing the spacecraft's observation state space and action state space, the spacecraft's actor network inputs the initial state. or Output action strategy or And the probability of selection, where include , include The action strategy is iterated through the spacecraft kinematics model to obtain the state space at the next moment. and and action rewards And store it in the experience pool;

[0077] Spacecraft commentator network based on the current global state The output state value function predicts the long-term reward of the input action and calculates the time difference error, ultimately achieving a score for the selected action.

[0078] Specifically, step 3 includes:

[0079] Inspired by classic game strategies in spacecraft swarm competition, such as forward circling, backward circling, rendezvous, and formation flying;

[0080] These strategies are decomposed and transformed into six underlying behaviors that are beneficial to the high-dynamic game of spacecraft clusters, and corresponding behavioral rewards are constructed as follows:

[0081] Directly pursuing behavioral rewards Encourage the chasing spacecraft to move in the same direction as the escaping spacecraft to create directional pressure;

[0082] Directly pursuing behavioral rewards Depends on the speed of the chasing spacecraft and escape spacecraft speed The similarity is expressed as ,in This represents the dot product of two unit vectors. A chasing spacecraft, under direct pursuit behavior, approximately behaves as a formation flight with an escaping spacecraft.

[0083] Tracking behavior rewards Encourage the pursuing spacecraft to move toward the escaped spacecraft's previous location to achieve tracking;

[0084] Tracking behavior rewards Depends on the speed of the chasing spacecraft With relative distance vector direction( This represents the position vector of the escaping spacecraft. The similarity (representing the position vector of the chasing spacecraft) is expressed as... ;

[0085] Pointing to behavioral rewards Encourage the pursuing spacecraft to adjust its position to align with the escape spacecraft's direction of motion to enhance the tracking effectiveness in subsequent rounds;

[0086] Pointing to behavioral rewards Depends on the direction of the relative distance vector relative to the velocity direction of the escape spacecraft The similarity is expressed as ;

[0087] Imitation behavior reward Encourage the pursuit of replicating the best-performing spacecraft within the spacecraft swarm;

[0088] Imitation behavior reward Depends on the speed of the chasing spacecraft Speed ​​of the best chasing spacecraft (the chasing spacecraft closest to the target) The similarity is expressed as ;

[0089] Encirclement behavior reward Encourage pursuing spacecraft to move in the direction in which the escaping spacecraft is most likely to flee, in order to form a containment mechanism;

[0090] Combining the velocity vectors of the chasing spacecraft and the escaping spacecraft, we obtain , This represents the composite vector. The chasing spacecraft follows... The directional movement reduces the escape angle and space available for the escaping spacecraft. As the number of pursuing spacecraft increases, the effective escape space for the escaping spacecraft decreases. Encirclement behavior reward. Depends on the speed of the chasing spacecraft direction and The similarity in direction is represented as ;

[0091] Predictive behavior rewards Encourage the chasing spacecraft to move towards the location where the escaping spacecraft might escape in the next moment;

[0092] Predictive behavior rewards Depends on the speed of the chasing spacecraft direction and The similarity of (chasing the spacecraft's velocity at the next moment) is expressed as ;

[0093] In addition to behavioral rewards, terminal state rewards are added. The function is set as follows:

[0094] Distance maintenance reward To guide the pursuing spacecraft to maintain a distance from the escaping spacecraft and prevent the pursuing spacecraft from becoming too far away from the escaping spacecraft during the pursuit process, it is represented as:

[0095] ;

[0096] Where D0 and k0 are parameters. , .

[0097] Closer reward The process guides the chasing spacecraft closer to the escaping spacecraft, inheriting characteristics of traditional models to ensure training stability, and is represented as follows:

[0098] ;

[0099] Mission completion reward Encouraging pursuing spacecraft to remain within the target's safe corridor until all remaining targets are pursued is represented as:

[0100] ;

[0101] Based on behavioral rewards and terminal state rewards, and combined with dynamic parameters, a chasing spacecraft is constructed. Game-theoretic escape spacecraft The reward function is:

[0102] ;

[0103] in, This signifies a reward for chasing spacecraft. Indicates the distance between the escaping spacecraft and the target. This indicates the distance between the escaping spacecraft and the pursuing spacecraft. Indicates chasing spacecraft Rewards for using behavior Represents logical values, It is a weighting coefficient, determined by the initial distance between the parties in the pursuit and capture; the reward for the escaping spacecraft is set as the negative of the reward for all corresponding pursuing spacecraft.

[0104] The reward for an escaping spacecraft is set to be the negative of the reward for all corresponding chasing spacecraft.

[0105] Unlike traditional reward function models, the above reward introduces... Item, and Shifting the focus of spacecraft maneuvers from primarily outcome-oriented to behavioral correctness can significantly reduce the complexity of chasing spacecraft and the training time required.

[0106] Specifically, step 4 includes:

[0107] When the speed of the defensive spacecraft exceeds the speed of the escape spacecraft And the relative distance continues to increase. A deceleration trend is used to suppress the continuous increase of relative distance, and the trend is determined by a deceleration pulse group (-M, -M, -M, -M, -M).

[0108] When the speed of the defensive spacecraft exceeds the speed of the escape spacecraft And the relative distance continues to increase. , The threshold constant is The distances between the pursuer and the escapee in the current round and the previous round, respectively, are used to suppress the continuous increase of the relative distance by using an acceleration trend. The trend is determined by the deceleration pulse group (M, M, M, M, M).

[0109] A semi-mandatory behavior switching mechanism is designed to ensure the adaptability of the chasing spacecraft at different game stages through flexible but infrequent behavior switching. Specifically:

[0110] The rules stipulate that the chasing spacecraft cannot switch behaviors twice within three consecutive rounds of the game, so that the spacecraft decides whether to change its behavior only after accumulating its advantages and disadvantages over a period of time.

[0111] Design a dynamic task allocation mechanism to maintain a uniform distribution of resources while effectively allocating chasing tasks. Specifically:

[0112] If the escaping spacecraft is not selected by any pursuing spacecraft, all pursuing spacecraft are free to choose their targets based solely on distance.

[0113] Once the escape spacecraft is selected, other pursuing spacecraft need to consider the current mission coalition when selecting that target. Has the quantity reached the expected value? ;

[0114] If the requirement is not met, you are free to choose. The specific process for choosing is as follows:

[0115] If chasing spacecraft distance Smaller than the current alliance The chasing spacecraft Maximum distance from the target ,but Can replace the league In The principle can be formalized as follows:

[0116] ;

[0117] This indicates the number of elements in the array. When the target missions are of similar difficulty or the spacecraft are isomorphic, the expected number of mission coalitions is based on the principle of equal distribution. When the target missions are of different difficulty, the expected number of mission coalitions can be designed based on the spacecraft's adversarial effectiveness. Spacecraft with strong escape spacecraft game-playing capabilities should be assigned more chasing spacecraft.

[0118] Under the assumption of spacecraft isomorphism, an equal-sharing mechanism can be adopted, meaning that the number of units in each mission coalition is... In cases where the value cannot be divided evenly, the value of each task alliance is determined manually or through some mechanism. .

[0119] Step 5: Train and update the proposed deep reinforcement learning algorithm. Based on the established framework, model, reward function, and mechanism, finally output a spacecraft capture and escape game strategy that meets the game expectations.

[0120] The present invention will be described in detail below with reference to embodiments:

[0121] like Figure 2 As shown in the figure, the multi-spacecraft encirclement and pursuit game decision-making method based on a group-independent learning strategy provided by this invention includes:

[0122] In the LVLH coordinate system of the orbital reference point, motion models of the pursuing and escaping spacecraft are established based on the equations of relative motion. The equation describes the relative position of the coordinate system at a given time relative to the origin of the LVLH coordinate system.

[0123] Set up a collection of chasing spacecraft and escape spacecraft collection To indicate the type of spacecraft;

[0124] Set up a chasing spacecraft cluster and escape spacecraft Payoff function and :

[0125] ;

[0126] in, and They are chasing spacecraft and escape spacecraft A set of speed impulse strategies, This indicates the distance between the escaping spacecraft and the pursuing spacecraft. Represents a specific mapping function;

[0127] According to zero-sum game theory, escape spacecraft The payoff function can be expressed as:

[0128] ;;

[0129] The Nash equilibrium solution of a zero-sum game can be expressed as: Further, the escape spacecraft and the corresponding collection of chasing spacecraft The equilibrium is defined as:

[0130]

[0131] and The Space Chasing Alliance Chasing the escape spacecraft At the time when each party is maximizing its own benefits, the Spacecraft Alliance is pursuing its goals. The strategy is The escape spacecraft strategy is , and This indicates a general strategy.

[0132] Set the maximum number of game rounds. and maximum speed pulse limit and pulse step size ;

[0133] set up To determine the minimum distance required to effectively capture an escaping spacecraft, and These represent the chasing spacecraft and the escape spacecraft respectively in the [missing information - likely a number]th ... Position of the game round;

[0134] The terminal objective that an escape spacecraft needs to achieve is defined as:

[0135] ;

[0136] When a chasing spacecraft fails to approach an escaping spacecraft, it needs to continuously optimize its own reward function.

[0137] Combining zero-sum game theory, the optimal mathematical model for chasing spacecraft is further obtained as follows:

[0138] ;

[0139] Then, the reinforcement learning method is improved based on the proximal policy optimization framework;

[0140] An observation information state space was constructed by extracting the positions and velocities of all chasing and escaping spacecraft. ;

[0141] Action space Set up with 3 indicators: behavioral index Target Index and pulse amplitude index or ;

[0142] Based on the above three indicators, and considering subsequent spacecraft behavior switching and mission allocation, the game strategy for chasing spacecraft is set as follows:

[0143] ;

[0144] After constructing the spacecraft's observation state space and action state space, the spacecraft's actor network inputs the initial state. or And output action strategy or And the probability of selection, where include , include ;

[0145] The action strategy is iterated through the spacecraft kinematics model to obtain the state space at the next moment. and and action rewards And store it in the experience pool;

[0146] Spacecraft critics network based on the current global state The output state value function predicts the long-term reward of the input action and calculates the time difference (TD) error to score the selected action.

[0147] After the reinforcement learning framework is built, the state reward function is further set.

[0148] Configure a reward function based on group-independent behavior, including direct pursuit of behavior rewards. Tracking behavior rewards Targeted behavior rewards Imitation behavior reward Encirclement behavior reward and predictive behavior rewards ;

[0149] Configure a terminal state-based reward function, including distance maintenance rewards. Distance Reward And mission completion rewards ;

[0150] chasing spacecraft Game-theoretic escape spacecraft The reward function is constructed as follows:

[0151] ;

[0152] in, This signifies a reward for chasing spacecraft. Indicates chasing spacecraft Rewards for using behavior Represents logical values, These are weighting coefficients, determined by the initial distances between the parties involved in the manhunt and pursuit, ensuring... This means that behavioral rewards account for 90% of the learning process, serving as the primary guide for spacecraft learning. The reward for an escaping spacecraft is set to the negative of the reward for all corresponding chasing spacecraft.

[0153] Then, a mechanism for determining forced acceleration and deceleration of spacecraft is constructed.

[0154] Then, a task allocation mechanism based on the hedonic game theory of alliances was designed; in this mechanism, when the escape spacecraft is not selected by any pursuing spacecraft, all pursuing spacecraft freely choose targets based solely on distance; once the escape spacecraft is selected, other pursuing spacecraft need to consider the current task alliance when selecting that target. Has the quantity reached the expected value? If not, you are free to choose; if distance Smaller than the current alliance The chasing spacecraft Maximum distance from the target ,if , Can replace the league In ;

[0155] Finally, the designed reinforcement learning intelligent learning method is fed with the corresponding observation information state. Output action state space and the state space of the next moment and and action rewards ;

[0156] Based on the optimization model of the pursuit and capture game, the state space at the next moment is... and Determine whether the task termination conditions have been met;

[0157] Based on the maximum number of game rounds Check if the current round number has exceeded the limit. If it has, the mission ends; otherwise, continue.

[0158] The gradient parameters of the actor-critic algorithm network are updated by extracting information from the experience pool according to conditions.

[0159] By continuously executing the previous step until the mission termination condition is met, a group-independent strategy for the spacecraft's pursuit and capture game is obtained.

[0160] The above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. A multi-spacecraft encirclement and pursuit game decision-making method based on a group-independent learning strategy, characterized in that, Includes the following steps: Step 1: In the scenario of a multi-spacecraft encirclement and pursuit game, establish motion models of the pursuing spacecraft, the monitoring spacecraft, and the target spacecraft in the LVLH coordinate system of the orbital reference point to describe the relative motion process of the spacecraft; Using velocity pulses as the basic strategies of both sides in the game and relative distance as the optimization objective, this paper considers the velocity pulse amplitude, time constraints, and terminal game constraints of spacecraft. Based on the zero-sum game principle, a multi-spacecraft encirclement and pursuit game model is constructed. The LVLH coordinate system represents the local vertical and local horizontal coordinate system. Step 2: Design a spacecraft intelligent learning algorithm based on the near-end strategy optimization framework, construct an observation state space including the position and velocity of both the pursuer and the pursuer, construct an action state space including velocity pulse size, behavior type and target assignment sequence, and design the training process for both the pursuer and the pursuer based on the principle of left and right mutual fighting. Step 3: Based on Step 2, design a reward function to decompose and transform the game strategies in spacecraft game, including forward flight, backward flight, rendezvous and formation flight, into six underlying behaviors that are conducive to the high dynamic game of spacecraft clusters. Combine these behaviors with terminal distance status rewards to form spacecraft training action incentives, with behavioral rewards accounting for more than 90%, ensuring that behavioral guidance is the core. Behavioral rewards include: 1) Directly pursue behavioral rewards Encourage the chasing spacecraft to move in the same direction as the escaping spacecraft, applying pressure and directly rewarding the chasing behavior. Depends on the speed of the chasing spacecraft and escape spacecraft speed The similarity is expressed as ,in The dot product of two unit vectors is represented; the chasing spacecraft in direct pursuit behavior is approximately in a formation flight state with the escaping spacecraft; 2) Tracking behavior rewards : Encourages the chasing spacecraft to move towards the previous position of the escaping spacecraft to achieve a tracking effect; rewards tracking behavior. Depends on the speed of the chasing spacecraft With relative distance vector The similarity in direction is represented as ; 3) Targeted behavioral rewards : Encourage the chasing spacecraft to adjust its position to align with the escape spacecraft's direction of motion, thereby enhancing the tracking effect in subsequent rounds; pointing behavior reward Depends on the direction of the relative distance vector relative to the velocity direction of the escape spacecraft The similarity is expressed as ; 4) Rewarding imitation behavior Encourage the imitation of the best-performing spacecraft in the group, and reward such behavior. Depends on the speed of the chasing spacecraft Similarity to the speed of the best-chasing spacecraft; 5) Rewarding the surrounding behavior The chasing spacecraft is encouraged to move in the direction in which the escaping spacecraft is most likely to escape. The velocity vectors of the chasing spacecraft and the escaping spacecraft are combined to obtain... ; Chasing spacecraft along The direction of movement reduces the angle and space for the escape spacecraft to evade; the encirclement behavior is rewarded. Depends on the speed of the chasing spacecraft direction and The similarity in direction is represented as ; 6) Predictive behavioral rewards : Encourages chasing spacecraft to move towards the location where the escaping spacecraft might escape in the next moment; predictive behavior rewards Depends on the speed of the chasing spacecraft Direction and the speed of the spacecraft at the next moment The similarity in direction is represented as ; Step 4: Design an acceleration / deceleration mechanism to selectively correct the spacecraft's actions, ensuring that the chasing spacecraft can quickly reverse its disadvantage as the relative distance continues to increase; design a semi-mandatory behavior switching mechanism to ensure the adaptability of the chasing spacecraft in different game stages through flexible but infrequent behavior switching; design a dynamic task allocation mechanism to maintain a uniform distribution of resources while effectively allocating chasing tasks. Step 5: Train and update the spacecraft pursuit and escape game decision algorithm constructed in Steps 1-4, and finally output a spacecraft encirclement and pursuit game strategy adapted to high dynamic game scenarios.

2. The multi-spacecraft encirclement and pursuit game decision-making method based on a group-independent learning strategy according to claim 1, characterized in that, In step 1, the relative motion model of the spacecraft is represented as follows: ; ; in, express The state of velocity and acceleration at any given moment. and It is the state transition matrix. These represent the real-time time and the initial time, respectively. These represent the sets of chasing spacecraft and the sets of escape spacecraft, respectively. The number of rounds in the game. It is the first Wheel speed increment, Indicates the current game round. Indicates the first Round number satellites The magnitude of the velocity pulse on the axis, with the superscript T indicating the transpose of the matrix; the equation describes the relative position of the spacecraft with respect to the origin of the LVLH coordinate system at a given time, without the need for control conditions; Using a round-based game theory approach, an optimization mathematical model is established for the spacecraft encirclement and pursuit game problem, expressed as follows: ; in, This indicates optimizing the objective function value for an escape spacecraft. The spacecraft encirclement alliance stated that it was , and Representing sets Benefits and escape spacecraft The benefits, Indicates the maximum pulse limit. This indicates the minimum distance at which a chasing spacecraft can effectively capture an escaping spacecraft. and These represent the chasing spacecraft in the [number]th [year]. The position of the round and the escape spacecraft in the The position of the round, and These represent the sets of chasing spacecraft and the sets of escaping spacecraft, respectively. This represents the maximum number of rounds in the game. This represents the 2-norm of a matrix.

3. The multi-spacecraft encirclement and pursuit game decision-making method based on a group-independent learning strategy according to claim 1, characterized in that, In step 2, a spacecraft intelligent learning algorithm following the zero-sum game principle is designed based on the near-end strategy optimization framework, and an observation state space including the position and velocity of both the pursuer and the pursuer is constructed. Construct velocity pulses including chasing and escaping spacecraft. Types of behavior in chasing spacecraft and target allocation sequence The action state space is used to solve the spacecraft's action strategy; Spacecraft action state space Perform the following coding design: Set the velocity pulse index Types of behavior in chasing spacecraft and target allocation sequence To follow the decimals in the interval [0,1], where The behavior type chosen by the chasing spacecraft is represented and mapped to an integer in the interval [1,6] after processing the cumulative probability density. This represents the selection of one of six game behaviors. The three codes conform to Gaussian, Beta, and Beta distributions, respectively. The number of the defensive escape spacecraft assigned to the chasing spacecraft is mapped to an integer in the interval [1, m] after processing with the accumulated probability density; the impulse index This represents a pulse combination containing pulses along three directional axes, with a total of [number] pulses per directional axis. A velocity pulse increment scheme, The pulse combination consists of a constant determined by the magnitude of the pulse gradient. kind, After processing the accumulated probability density, it is mapped to the interval [1, ...]. [A set of integers, each integer corresponding to a velocity pulse increment scheme;] Constructing the actor-critic neural network required for information transmission during the training process; the initial state of the spacecraft's actor network input. or Output action strategy or And the probability of selection, where include , include The action strategy is iterated through the spacecraft kinematics model to obtain the state space at the next moment. and and action rewards And store it in the experience pool; the spacecraft commentator network is based on the current global state. The output state value function predicts the long-term reward of the input action and calculates the time difference error, ultimately achieving a score for the selected action.

4. The multi-spacecraft encirclement and pursuit game decision-making method based on a group-independent learning strategy according to claim 1, characterized in that, In addition to behavioral rewards, terminal status rewards are added, specifically the following three settings: 1) Distance maintenance bonus To guide the pursuing spacecraft to maintain a distance from the escaping spacecraft and prevent the pursuing spacecraft from becoming too far away from the escaping spacecraft during the pursuit process, it is represented as: ; Where D0 and k0 are parameters. , ; 2) Closer Reward : Guiding the chasing spacecraft to approach the escaping spacecraft, inheriting the characteristics of traditional models to ensure training stability, is represented as: ; 3) Rewards for completing the mission Encouraging pursuing spacecraft to remain within the target's safe corridor until all remaining targets are pursued is represented as: ; Based on behavioral rewards and terminal state rewards, and combined with dynamic parameters, the spacecraft is chased. Game-theoretic escape spacecraft The reward function is constructed as follows: ; in, This signifies a reward for chasing spacecraft. Indicates the distance between the escaping spacecraft and the target. This indicates the distance between the escaping spacecraft and the pursuing spacecraft. Indicates chasing spacecraft Rewards for using behavior Represents logical values, It is a weighting coefficient, determined by the initial distance between the parties in the pursuit and capture; the reward for the escaping spacecraft is set as the negative of the reward for all corresponding pursuing spacecraft.

5. The multi-spacecraft encirclement and pursuit game decision-making method based on a group-independent learning strategy according to claim 1, characterized in that, The acceleration / deceleration mechanism in step 4 includes: 1) When the speed of the escaping spacecraft exceeds the speed of the pursuing spacecraft And the relative distance continues to increase. , The threshold constant is These represent the distances between the pursuer and the escapee in the current round and the previous round, respectively. A deceleration trend is used to suppress the continuous increase of the relative distance. The trend is determined by the deceleration pulse group (-M, -M, -M, -M, -M). 2) When the speed of the escaping spacecraft exceeds the speed of the pursuing spacecraft And the relative distance continues to increase. An acceleration trend is used to suppress the continuous increase in relative distance, and the trend is determined by a deceleration pulse group (M, M, M, M, M).

6. The multi-spacecraft encirclement and pursuit game decision-making method based on a group-independent learning strategy according to claim 1, characterized in that, The semi-mandatory behavior switching mechanism in step 4 is a semi-random behavior switching method that guides the spacecraft to make reasonable behavior switching. It stipulates that the chasing spacecraft cannot make two behavior switchings within three consecutive game rounds, so that the spacecraft decides whether to change its behavior only after the advantages and disadvantages of its behavior have accumulated over a period of time.

7. The multi-spacecraft encirclement and pursuit game decision-making method based on a group-independent learning strategy according to claim 1, characterized in that, The dynamic task allocation method in step 4 includes: If the escape spacecraft is not selected by any pursuing spacecraft, all pursuing spacecraft freely choose their targets based solely on distance. Once the escape spacecraft is selected, other pursuing spacecraft, when selecting that target, need to consider the current mission coalition. Has the quantity reached the expected value? If not achieved, then the choice is free, and the specific free choice process is as follows: If chasing a spacecraft distance Smaller than the current alliance The chasing spacecraft Maximum distance from the target ,but Replace the alliance In , is represented as: ; When the target missions are of similar difficulty or the spacecraft are isomorphic, the expected number of mission coalitions is determined by an equal distribution principle. When the target mission difficulties are inconsistent, the expected number of mission coalitions is designed based on the spacecraft's game-theoretic efficiency, allocating more chasing spacecraft to spacecraft with stronger escape game-theoretic capabilities. Under the assumption of isomorphic spacecraft, an equal distribution mechanism is adopted, meaning the number of coalitions per mission is... In cases where the value cannot be divided evenly, the value of each task alliance is determined manually or through some mechanism. .