Planning and command control method and system for spacecraft pursuit game

By constructing a spacecraft pursuit-escape game state model and multi-agent collaborative decision-making, and combining adaptive dynamic programming and perturbation compensation, the problems of low solution efficiency of high-dimensional differential game models and the influence of orbital perturbations are solved, and real-time and robust control of spacecraft pursuit-escape game is realized.

CN121553399APending Publication Date: 2026-02-24HARBIN INST OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511726074.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-24
Publication Date
2026-02-24

AI Technical Summary

Technical Problem

Existing technologies suffer from low efficiency in solving high-dimensional differential game models, making it difficult to generate control commands in real time. They also lack multi-spacecraft collaborative strategies and fail to consider the impact of orbital perturbations on control accuracy, thus limiting the application of spacecraft pursuit-escape game control methods in complex space environments.

Method used

Basic data is acquired through the track measurement and control system, a pursuit-escape game state model is constructed, multi-agent collaborative decision-making is used to generate the global situation field gradient, and adaptive dynamic programming and anti-perturbation compensation mechanism are combined to generate the optimal control command and achieve robust control.

Benefits of technology

It improves the efficiency of solving high-dimensional problems in pursuit and escape game, meets the requirements of millisecond-level real-time decision-making, increases the capture probability of multi-spacecraft collaborative encirclement and capture, and reduces the orbital control error to the meter level, significantly improving control accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121553399A_ABST
    Figure CN121553399A_ABST
Patent Text Reader

Abstract

The invention discloses a planning and command control method and system for a pursuit game of a spacecraft, and the method comprises the steps: obtaining the basic data of the pursuit game through an orbit measurement and control system; constructing a pursuit game state model, and updating a state vector of the pursuit game state model in real time; generating a global potential field gradient based on multi-agent collaborative decision; based on the state vector and the global potential field gradient, optimal control is solved through self-adaptive dynamic programming, and a final maneuvering instruction containing the thrust direction and amplitude is generated; and correcting the generated final maneuvering instruction through an anti-perturbation compensation mechanism to realize robust control. According to the planning and command control method for the pursuit game of the spacecrafts, by fusing adaptive dynamic planning, multi-agent collaborative decision-making and perturbation compensation mechanisms, a maneuvering instruction with optimal fuel can be efficiently generated, collaborative hunting of multiple spacecrafts in a complex perturbation environment is achieved, and the maneuvering efficiency of the spacecrafts is improved. And a core technical support is provided for a space security task.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of spacecraft autonomous control and orbital dynamics technology, and in particular to a planning and command control method and system for spacecraft pursuit and escape game, applicable to autonomous decision-making, path planning and real-time control of spacecraft in complex space environments. Background Technology

[0002] With the development of space warfare technology, the spacecraft pursuit-escape game problem has become a core issue in ensuring space security. While existing pursuit-escape models based on differential game theory can describe continuous dynamic adversarial processes, their solutions rely on high-dimensional two-point boundary value problems, resulting in high computational complexity and difficulty in meeting real-time requirements. Furthermore, existing methods mostly focus on one-to-one game scenarios, failing to fully consider the impact of multi-spacecraft cooperative strategies and actual orbital perturbations (such as Earth's oblateness J2 perturbation and atmospheric drag) on ​​control accuracy. For example, traditional formation control methods are prone to configuration instability in low Earth orbit due to perturbations, while single-spacecraft pursuit-escape models struggle to handle multi-target cooperative encirclement scenarios. Therefore, a pursuit-escape game control method that balances real-time performance, robustness, and cooperation is urgently needed. Summary of the Invention

[0003] The purpose of this invention is to provide a planning and command control method and system for spacecraft pursuit and escape game, in order to solve the problems of low solution efficiency of high-dimensional differential game models in the prior art, difficulty in generating control commands in real time; lack of multi-spacecraft cooperative strategies, making it impossible to achieve complex encirclement and capture missions; and failure to consider the impact of orbital perturbations on control accuracy, which limits practical application.

[0004] To achieve the above objectives, this invention provides a planning and command control method for spacecraft pursuit and escape game, comprising the following steps: S1. Obtain basic data for the pursuit and escape game through the track measurement and control system; S2. Based on the basic data of the pursuit and escape game collected in S1, construct the state model of the pursuit and escape game and update the state vector of the pursuit and escape game state model in real time. S3. Based on the pursuit-escape game state model, the global situation gradient is generated through multi-agent collaborative decision-making. S4. Based on the state vector of S2 and the global situation field gradient of S3, the optimal control is solved by adaptive dynamic programming, and the final maneuver command containing the thrust direction and amplitude is generated. S5. Through the anti-perturbation compensation mechanism, the final maneuver command generated by S4 is modified to achieve robust control.

[0005] Preferably, the basic data in S1 includes orbital dynamic data and environmental disturbance data, as shown below: Orbital dynamic data: including orbital state vectors and orbital observation vectors of the tracking spacecraft and the escape spacecraft; Environmental disturbance data: including but not limited to perturbation disturbance data such as J2 perturbation, atmospheric drag perturbation, solar radiation pressure perturbation, solar gravitational perturbation, and lunar gravitational perturbation.

[0006] Preferably, S2 is as follows: S21. Initialize the orbital state vectors of the tracking spacecraft and the escape spacecraft collected in S1, and establish a game state model of relative orbital dynamics based on the Clohessy-Wiltshire equation. S22. Define the initial state vector of the pursuit-escape game state model. As shown below: ; in, These represent the relative position coordinates of the tracking spacecraft and the escape spacecraft in the geocentric inertial coordinate system, respectively. These represent the derivatives of the relative position coordinates of the tracking spacecraft and the escape spacecraft in the geocentric inertial coordinate system with respect to time, respectively. S23. Decompose the J2 perturbation collected by S1 into J2 perturbation acceleration components, and use the J2 perturbation acceleration components as compensation terms; at the same time, extend the initial state vector to a 9-dimensional state vector, where the first 6 dimensions of the 9-dimensional state vector are the basic relative motion components, and the last 3 dimensions are the J2 perturbation acceleration components; update the last 3 dimensions of the J2 perturbation acceleration components in real time through the orbital perturbation differential equation; S24. Based on stochastic process theory, the unknown maneuvering behavior of the escape spacecraft is modeled, and Kalman filtering is used to estimate the instantaneous orbital parameters of the escape spacecraft in real time, updating the first 6 basic relative motion components in the 9-dimensional state vector.

[0007] Preferably, S3 is as follows: S31. Construct a distributed command and control architecture, treating multiple tracking spacecraft as collaborative intelligent agents, and achieve information sharing and task allocation through consensus algorithms; S32. A capture strategy is designed based on the potential field method. An exponential repulsive potential field is defined around the escaping spacecraft, and a linear attractive potential field is defined between the tracking spacecraft. The overall potential field gradient is generated by combining the potential field gradients. The exponential repulsive potential field is shown below: ; in, This represents the repulsive force of the exponential repulsive potential field at the current position; This represents the coefficient of exponential repulsive potential strength. This indicates the spatial distance between the tracking spacecraft and the escape spacecraft; Indicates the safe distance between tracking spacecraft and escape spacecraft; The linear attractive potential field is shown below: ; in, This represents the attractive force of the linear attractive potential field at the current location; This represents the linear attraction potential intensity coefficient.

[0008] Preferably, S4 is as follows: S41. Construct a two-layer optimization framework to transform the pursuit-escape game problem into a two-layer optimization problem: an upper-layer long-term task module and a lower-layer short-term task module; S42. Introduce time scale separation technology to decompose long-term tasks into short-term rolling optimization tasks to reduce computational complexity. S43. Design a reward function to balance fuel consumption and capture time; S44. Based on the output of the reward function in S43, update the network weights through policy iteration. S45. Based on the updated network weights, combined with the real-time state vector of S2 and the global state field gradient of S3, the final maneuver command containing the thrust direction and amplitude is generated after the escape spacecraft enters the capture radius through the upper-level Hamilton-Jacobi-Bellman equation iteration.

[0009] Preferably, in S41, the pursuit-escape game problem is transformed into a two-level optimization problem, as detailed below: (1) The upper-level long-term task module solves the game equilibrium point through the Hamilton-Jacobi-Bellman equation; (2) The lower-level short-term task module takes the real-time state vector of S2 and the global situation field gradient of S3 as input, and uses a neural network to approximate the optimal control law to generate the fuel-optimal maneuver command in real time.

[0010] Preferably, the neural network includes an evaluation network and an execution network. The evaluation network is used to evaluate the quality of the action and provide quantitative evaluation index data; the execution network is used to make decisions and select the action to be performed based on the current state.

[0011] Preferably, the reward function in S43 is used to balance fuel consumption and capture time, as shown below: ; in, This represents the reward value of the reward function; This represents the weighting coefficient for fuel consumption evaluation; The total speed pulse indicates the amount of fuel consumed; This represents the weighting coefficient for the captured event evaluation; Indicates the capture time.

[0012] Preferably, S5 is as follows: S51. Through the extended state observer, the J2 perturbation acceleration and atmospheric drag disturbance are estimated in real time; S52. Feed the estimated J2 perturbation acceleration forward to the control law; combine Lyapunov stability theory to modify the final maneuver command generated in S4, which includes the thrust direction and amplitude, to verify the robustness of the closed-loop system and ensure convergence under model uncertainty.

[0013] A planning and command control system for spacecraft pursuit and escape game, the system includes an orbit measurement and control system for measuring the orbital data of the escaping spacecraft and the tracking spacecraft, and realizing precise orbit determination at the system level; The state awareness module is used to identify the orbital maneuvering behavior or current orbital position of the escaping spacecraft target, and to classify and grade the orbital behavior patterns to form a unified space situational awareness result. The collaborative decision-making module is used to generate collaborative response strategies for tracking spacecraft and capturing escaped spacecraft based on the current space situation. The optimized control module is used to generate spacecraft-level optimized control commands based on the collaborative response strategy and send them to the tracking spacecraft. The execution compensation module, based on optimized control commands and combined with the state perception module, forms a closed-loop control command to achieve closed-loop compensation for perturbation interference force, perturbation interference torque, and actuator output error.

[0014] Therefore, the present invention employs the above-mentioned planning and command control method and system for spacecraft pursuit and escape game, and the beneficial effects are as follows: (1) This invention effectively improves the efficiency of solving high-dimensional problems of pursuit and escape game by using adaptive dynamic programming and a distributed collaborative framework. The above meets the requirements for millisecond-level real-time decision-making.

[0015] (2) Simulation experiments have verified that the multi-spacecraft cooperative capture strategy proposed in this invention increases the capture probability of escaping spacecraft to [percentage missing]. .

[0016] (3) The present invention uses a perturbation compensation module to reduce the track control error to the meter level, which significantly improves the control accuracy compared with the traditional track control method.

[0017] The technical solution of the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. Attached Figure Description

[0018] Figure 1 This is an architecture diagram of a planning and command control system for spacecraft pursuit and escape game according to the present invention. Figure 2This is a flowchart of the adaptive dynamic programming algorithm in an embodiment of the present invention; Figure 3 This is a block diagram of the perturbation compensation control module in an embodiment of the present invention. Detailed Implementation

[0019] The technical solution of the present invention will be further described below with reference to the accompanying drawings and embodiments.

[0020] This invention provides a planning and command control method for spacecraft pursuit and escape game, comprising the following steps: S1. Obtain basic data for the pursuit and escape game through the track measurement and control system; S2. Based on the basic data of the pursuit and escape game collected in S1, construct the state model of the pursuit and escape game and update the state vector of the pursuit and escape game state model in real time. S3. Based on the pursuit-escape game state model, the global situation gradient is generated through multi-agent collaborative decision-making. S4. Based on the state vector of S2 and the global situation field gradient of S3, the optimal control is solved by adaptive dynamic programming, and the final maneuver command containing the thrust direction and amplitude is generated. S5. Through the anti-perturbation compensation mechanism, the final maneuver command generated by S4 is modified to achieve robust control.

[0021] Example like Figure 1 As shown, the present invention provides a planning and command control system for spacecraft pursuit and escape game, including an orbit measurement and control system for measuring the orbital data of the escaping spacecraft and the tracking spacecraft, thereby achieving precise orbit determination at the system level.

[0022] The state awareness module is used to identify the orbital maneuvering behavior or current orbital position of the escaped spacecraft target, and to classify and grade the orbital behavior patterns to form a unified space situational awareness result.

[0023] The collaborative decision-making module is used to generate collaborative response strategies for tracking spacecraft and capturing escaped spacecraft based on the current space situation.

[0024] The optimized control module is used to generate spacecraft-level optimized control commands based on the content of the collaborative response strategy and send them to the tracking spacecraft.

[0025] The execution compensation module, based on optimized control commands and combined with the state perception module, forms a closed-loop control command to achieve closed-loop compensation for perturbation interference force, perturbation interference torque, and actuator output error.

[0026] Based on the above system, this invention provides a planning and command control method for spacecraft pursuit and escape game, comprising the following steps: S1. Obtain basic data for the pursuit and escape game through the orbit tracking and control system. Specifically, the basic data includes orbit dynamic data and environmental interference data, as shown below: Orbital dynamic data: including orbital state vectors and orbital observation vectors of tracking and escaping spacecraft.

[0027] Environmental disturbance data: including perturbation and disturbance force data such as J2 perturbation, atmospheric drag perturbation, solar radiation pressure perturbation, solar gravitational perturbation, and lunar gravitational perturbation.

[0028] S2. Based on the basic data of the pursuit and escape game collected in S1, construct the state model of the pursuit and escape game and update the state vector of the pursuit and escape game state model in real time.

[0029] S21. Initialize the orbital state vectors of the tracking spacecraft and the escape spacecraft collected in S1, and establish a game state model of relative orbital dynamics based on the Clohessy-Wiltshire equation.

[0030] S22. Define the initial state vector of the pursuit-escape game state model. As shown below: ; in, These represent the relative position coordinates of the tracking spacecraft and the escape spacecraft in the geocentric inertial coordinate system (ECI), respectively. These represent the derivatives of the relative position coordinates of the tracking spacecraft and the escape spacecraft in the geocentric inertial coordinate system with respect to time.

[0031] S23. Decompose the J2 perturbation collected by S1 into J2 perturbation acceleration components and use the J2 perturbation acceleration components as compensation terms; at the same time, extend the initial state vector to a 9-dimensional state vector, where the first 6 dimensions of the 9-dimensional state vector are the basic relative running components and the last 3 dimensions are the J2 perturbation acceleration components; update the last 3 dimensions of the J2 perturbation acceleration components in real time through the orbital perturbation differential equation.

[0032] S24. Based on stochastic process theory, the unknown maneuvering behavior of the escape spacecraft is modeled, and Kalman filtering is used to estimate the instantaneous orbital parameters of the escape spacecraft in real time, updating the first 6 basic relative motion components in the 9-dimensional state vector.

[0033] S3. Based on the pursuit-escape game state model, the global situation gradient is generated through multi-agent collaborative decision-making.

[0034] S31. Construct a distributed command and control architecture, treating multiple tracking spacecraft as collaborative intelligent agents, and achieve information sharing and task allocation through consensus algorithms.

[0035] In this embodiment, each tracking spacecraft broadcasts its own status via a wireless network and uses a consensus protocol to calculate the priority weight of the capture mission.

[0036] S32. A capture strategy is designed based on the potential field method. An exponential repulsive potential field is defined around the escaping spacecraft, and a linear attractive potential field is defined between the tracking spacecraft. The overall potential field gradient is generated by combining the potential field gradients. The exponential repulsive potential field is shown below: ; in, This represents the repulsive force of the exponential repulsive potential field at the current position; This represents the coefficient of exponential repulsive potential strength. This indicates the spatial distance between the tracking spacecraft and the escape spacecraft; This indicates the safe distance between tracking spacecraft and escape spacecraft.

[0037] The linear attractive potential field is shown below: ; in, This represents the attractive force of the linear attractive potential field at the current location; This represents the linear attraction potential intensity coefficient.

[0038] S4. Based on the state vector of S2 and the global situation field gradient of S3, the optimal control is solved through adaptive dynamic programming (ADP), and the final maneuver command containing the thrust direction and amplitude is generated, such as... Figure 2 As shown.

[0039] S41. Construct a two-layer optimization framework to transform the pursuit-escape game problem into a two-layer optimization problem: an upper-layer long-term task module and a lower-layer short-term task module, as detailed below: (1) The upper-level long-term task module solves the game equilibrium point through the Hamilton-Jacobi-Bellman (HJB) equation.

[0040] (2) The lower-level short-term task module takes the real-time state vector of S2 and the global situation field gradient of S3 as input, and uses a neural network to approximate the optimal control law to generate the fuel-optimal maneuver command in real time.

[0041] The neural network includes a Critic Network and an Actor Network. The Critic Network evaluates the quality of actions and provides quantifiable evaluation metrics. The Actor Network makes decisions and selects the action to be performed based on the current state.

[0042] S42. Introduce time scale separation technology to decompose long-term tasks into short-term rolling optimization tasks, thereby reducing computational complexity.

[0043] S43. Design a reward function to balance fuel consumption and capture time. The reward function is as follows: ; in, This represents the reward value of the reward function; This represents the weighting coefficient for fuel consumption evaluation; The total speed pulse indicates the amount of fuel consumed; This represents the weighting coefficient for the captured event evaluation; Indicates the capture time.

[0044] S44. Based on the output of the reward function in S43, the network weights are updated iteratively through policy.

[0045] S45. Based on the updated network weights, combined with the real-time state vector of S2 and the global state field gradient of S3, the final maneuver command containing the thrust direction and amplitude is generated after the escape spacecraft enters the capture radius through the upper-level Hamilton-Jacobi-Bellman equation iteration.

[0046] S5. Through an anti-perturbation compensation mechanism, the final maneuver command generated by S4 is modified to achieve robust control, such as... Figure 3 As shown.

[0047] S51. The J2 perturbation acceleration and atmospheric drag disturbance are estimated in real time through the Extended State Observer (ESO).

[0048] S52. Feed the estimated J2 perturbation acceleration forward to the control law; combine Lyapunov stability theory to modify the final maneuver command generated in S4, which includes the thrust direction and amplitude, to verify the robustness of the closed-loop system and ensure convergence under model uncertainty.

[0049] This invention constructs a near-Earth orbit (LEO) scenario in the STK / Matlab co-simulation platform, with the J2 perturbation coefficient being... The proposed algorithm was compared with traditional differential game algorithms through simulation tests. Experiments show that, compared to traditional algorithms, the target acquisition time of this invention is significantly reduced. Reduced fuel consumption After processing by the perturbation compensation strategy of this invention, the position tracking error is reduced. rice.

[0050] Therefore, the present invention employs the above-mentioned planning and command control method and system for spacecraft pursuit and escape game, and the beneficial effects are as follows: (1) This invention effectively improves the efficiency of solving high-dimensional problems of pursuit and escape game by using adaptive dynamic programming and a distributed collaborative framework. The above meets the requirements for millisecond-level real-time decision-making.

[0051] (2) Simulation experiments have verified that the multi-spacecraft cooperative capture strategy proposed in this invention increases the capture probability of escaping spacecraft to [percentage missing]. .

[0052] (3) The present invention uses a perturbation compensation module to reduce the track control error to the meter level, which significantly improves the control accuracy compared with the traditional track control method.

[0053] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can still be made to the technical solutions of the present invention, and these modifications or equivalent substitutions cannot cause the modified technical solutions to deviate from the spirit and scope of the technical solutions of the present invention.

Claims

1. A planning and command control method for spacecraft pursuit and escape game, characterized in that, Includes the following steps: S1. Obtain basic data for the pursuit and escape game through the track measurement and control system; S2. Based on the basic data of the pursuit and escape game collected in S1, construct the state model of the pursuit and escape game and update the state vector of the pursuit and escape game state model in real time. S3. Based on the pursuit-escape game state model, the global situation gradient is generated through multi-agent collaborative decision-making. S4. Based on the state vector of S2 and the global situation field gradient of S3, the optimal control is solved by adaptive dynamic programming, and the final maneuver command containing the thrust direction and amplitude is generated. S5. Through the anti-perturbation compensation mechanism, the final maneuver command generated by S4 is modified to achieve robust control.

2. The planning and command control method for spacecraft pursuit and escape game as described in claim 1, characterized in that, The basic data in S1 includes orbital dynamics data and environmental disturbance data, as shown below: Orbital dynamic data, including orbital state vectors and orbital observation vectors for tracking and escaping spacecraft; Environmental disturbance data, including but not limited to perturbation disturbance data such as J2 perturbation, atmospheric drag perturbation, solar radiation pressure perturbation, solar gravitational perturbation, and lunar gravitational perturbation.

3. The planning and command control method for spacecraft pursuit and escape game as described in claim 1, characterized in that, S2 specifically refers to: S21. Initialize the orbital state vectors of the tracking spacecraft and the escape spacecraft collected in S1, and establish a game state model of relative orbital dynamics based on the Clohessy-Wiltshire equation. S22. Define the initial state vector of the pursuit-escape game state model. As shown below: ; in, These represent the relative position coordinates of the tracking spacecraft and the escape spacecraft in the geocentric inertial coordinate system, respectively. These represent the derivatives of the relative position coordinates of the tracking spacecraft and the escape spacecraft in the geocentric inertial coordinate system with respect to time, respectively. S23. Decompose the J2 perturbation collected by S1 into J2 perturbation acceleration components, and use the J2 perturbation acceleration components as compensation terms; at the same time, extend the initial state vector to a 9-dimensional state vector, where the first 6 dimensions of the 9-dimensional state vector are the basic relative motion components, and the last 3 dimensions are the J2 perturbation acceleration components; update the last 3 dimensions of the J2 perturbation acceleration components in real time through the orbital perturbation differential equation; S24. Based on stochastic process theory, the unknown maneuvering behavior of the escape spacecraft is modeled, and Kalman filtering is used to estimate the instantaneous orbital parameters of the escape spacecraft in real time, updating the first 6 basic relative motion components in the 9-dimensional state vector.

4. The planning and command control method for spacecraft pursuit and escape game as described in claim 1, characterized in that, S3 specifically refers to: S31. Construct a distributed command and control architecture, treating multiple tracking spacecraft as collaborative intelligent agents, and achieve information sharing and task allocation through consensus algorithms; S32. A capture strategy is designed based on the potential field method. An exponential repulsive potential field is defined around the escaping spacecraft, and a linear attractive potential field is defined between the tracking spacecraft. The overall potential field gradient is generated by combining the potential field gradients. The exponential repulsive potential field is shown below: ; in, This represents the repulsive force of the exponential repulsive potential field at the current position; This represents the coefficient of exponential repulsive potential strength. This indicates the spatial distance between the tracking spacecraft and the escape spacecraft; Indicates the safe distance between tracking spacecraft and escape spacecraft; The linear attractive potential field is shown below: ; in, This represents the attractive force of the linear attractive potential field at the current location; This represents the linear attraction potential intensity coefficient.

5. The planning and command control method for spacecraft pursuit and escape game as described in claim 1, characterized in that, S4 specifically refers to: S41. Construct a two-layer optimization framework to transform the pursuit-escape game problem into a two-layer optimization problem: an upper-layer long-term task module and a lower-layer short-term task module; S42. Introduce time scale separation technology to decompose long-term tasks into short-term rolling optimization tasks to reduce computational complexity. S43. Design a reward function to balance fuel consumption and capture time; S44. Based on the output of the reward function in S43, update the network weights through policy iteration. S45. Based on the updated network weights, combined with the real-time state vector of S2 and the global state field gradient of S3, the final maneuver command containing the thrust direction and amplitude is generated after the escape spacecraft enters the capture radius through the upper-level Hamilton-Jacobi-Bellman equation iteration.

6. The planning and command control method for spacecraft pursuit and escape game as described in claim 5, characterized in that, In S41, the pursuit-escape game problem is transformed into a two-level optimization problem, as detailed below: (1) The upper-level long-term task module solves the game equilibrium point through the Hamilton-Jacobi-Bellman equation; (2) The lower-level short-term task module takes the real-time state vector of S2 and the global situation field gradient of S3 as input, and uses a neural network to approximate the optimal control law to generate the fuel-optimal maneuver command in real time.

7. A planning and command control method for spacecraft pursuit and escape game as described in claim 6, characterized in that, Neural networks consist of an evaluation network and an execution network. The evaluation network evaluates the quality of actions and provides quantifiable evaluation metrics. The execution network makes decisions and selects the action to be performed based on the current state.

8. The planning and command control method for spacecraft pursuit and escape game as described in claim 5, characterized in that, The reward function in S43 is used to balance fuel consumption and capture time, as shown below: ; in, This represents the reward value of the reward function; This represents the weighting coefficient for fuel consumption evaluation; The total speed pulse indicates the amount of fuel consumed; This represents the weighting coefficient for the captured event evaluation; Indicates the capture time.

9. A planning and command control method for spacecraft pursuit and escape game as described in claim 1, characterized in that, S5 specifically refers to: S51. Through the extended state observer, the J2 perturbation acceleration and atmospheric drag disturbance are estimated in real time; S52. Feed the estimated J2 perturbation acceleration forward to the control law; combine Lyapunov stability theory to modify the final maneuver command generated in S4, which includes the thrust direction and amplitude, to verify the robustness of the closed-loop system and ensure convergence under model uncertainty.

10. A planning and command control system for spacecraft pursuit and escape game, characterized in that, The system is used to execute the planning and command control method for spacecraft pursuit and escape game as described in any one of claims 1-9, the system including an orbit measurement and control system for measuring the orbital data of the escaping spacecraft and the tracking spacecraft, and realizing precise orbit determination at the system level; The state awareness module is used to identify the orbital maneuvering behavior or current orbital position of the escaping spacecraft target, and to classify and grade the orbital behavior patterns to form a unified space situational awareness result. The collaborative decision-making module is used to generate collaborative response strategies for tracking spacecraft and capturing escaped spacecraft based on the current space situation. The optimized control module is used to generate spacecraft-level optimized control commands based on the collaborative response strategy and send them to the tracking spacecraft. The execution compensation module, based on optimized control commands and combined with the state perception module, forms a closed-loop control command to achieve closed-loop compensation for perturbation interference force, perturbation interference torque, and actuator output error.