Vision-based distributed spacecraft on-orbit pursuit-evasion game control method and system

By combining multi-spacecraft visual perception and feature matching technology with orbital pursuit and escape game model and generative adversarial learning, the problem of high-precision target reconstruction and real-time pursuit and escape control based on pure visual information in complex orbital environments has been solved, achieving efficient autonomous identification and intelligent pursuit control.

CN122144180APending Publication Date: 2026-06-05HANGZHOU INTERNATIONAL INNOVATION INSTITUTE OF BEIHANG UNIVERSITY

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
HANGZHOU INTERNATIONAL INNOVATION INSTITUTE OF BEIHANG UNIVERSITY
Filing Date
2025-12-17
Publication Date
2026-06-05

AI Technical Summary

Technical Problem

Existing technologies struggle to achieve high-precision target reconstruction and real-time pursuit control based on purely visual information in complex orbital environments. Traditional methods suffer from high computational complexity in high-dimensional state games, making them difficult to apply in real time.

Method used

By employing multi-spacecraft visual perception and feature matching, combined with feature point matching and intrinsic matrix solving techniques, an orbital pursuit and escape game model is established. Furthermore, a strategy network is trained through adversarial imitation learning to form a perception-decision-control closed-loop system.

Benefits of technology

It has achieved high-precision autonomous identification and intelligent pursuit control in non-cooperative, highly dynamic, and complex lighting environments, improving the autonomy of spacecraft in orbit and the success rate of missions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122144180A_ABST
    Figure CN122144180A_ABST
Patent Text Reader

Abstract

The application discloses a kind of distributed spacecraft in-orbit pursuit game control method and system based on vision, which comprises the following steps: S1, the orbit target image sequence is collected using the forward camera of multiple pursuit spacecraft, and the effective target area is screened out;S2, the feature points are extracted using SuperPoint, and the high-confidence matching point pair is output by SuperGlue structure semantic matching;S3, the intrinsic matrix is calculated and the nine-dimensional state vector is generated by fusing the estimation results of multiple spacecraft;S4, the nine-dimensional state vector is used to establish an orbit pursuit game model, and the optimal control law is solved by combining the relative motion dynamics equation to generate the expert strategy trajectory;S5, the trajectory is used to train the control strategy network;S6, the trained control strategy network is deployed to the attitude and orbit control system, and the orbit control instruction is output to form a perception-decision-control closed loop;The application realizes the efficient cooperative control of distributed spacecraft in-orbit pursuit through spacecraft visual perception, feature matching, game modeling and generative adversarial imitation learning.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of autonomous navigation and intelligent control technology for spacecraft, and more specifically to a vision-based distributed on-orbit pursuit and escape game control method and system for spacecraft. Background Technology

[0002] With the rapid development of deep space exploration and on-orbit rendezvous missions, autonomous identification and control of non-cooperative targets has become a core capability for spacecraft missions. In complex orbital environments, traditional navigation methods relying on multi-source sensors such as radar and lidar often fail to meet the requirements of long-term autonomous missions due to sensor availability, communication latency, and resource constraints. Especially when targets exhibit escape behavior, control strategies based on predefined algorithms struggle to achieve real-time response and efficient acquisition.

[0003] In recent years, visual navigation technology has become crucial for achieving autonomous spacecraft control due to its high-precision attitude and position estimation capabilities. However, traditional visual methods are sensitive to illumination, background interference, and target deformation, and rely excessively on prior models or training data, resulting in limited generalization ability. Meanwhile, existing pursuit and escape control methods are mostly based on optimal control theory, facing real-time bottlenecks when dealing with high-dimensional state games. Therefore, integrating robust visual state estimation with intelligent game-theoretic decision control has become an urgent research direction.

[0004] At the visual perception level, deep learning technologies, exemplified by the "segment everything" model, have made it possible to segment any target without prior knowledge, effectively solving the uncertainty problems caused by non-cooperative targets due to their shape, coating, occlusion, etc. By combining feature point matching and intrinsic matrix solving techniques, spacecraft can accurately calculate the relative attitude of targets without the need for structured light and depth information, laying the foundation for subsequent orbit determination and control.

[0005] At the decision-making and control level, orbital pursuit is essentially a dynamic asymmetric information game problem. While traditional differential game methods can provide theoretically optimal solutions, their high computational complexity makes them difficult to apply in real time. In contrast, intelligent methods such as imitation learning and multi-agent reinforcement learning can train robust policy networks under incomplete observation conditions, enabling spacecraft to quickly respond to target maneuvers and optimize their own paths, significantly improving the system's intelligence level.

[0006] In summary, the core technological challenge currently faced by spacecraft in autonomous rendezvous and on-orbit combat missions lies in how to accurately reconstruct the relative state of a target based on purely visual information and generate adversarial maneuver commands in real time, thereby constructing a complete and efficient closed-loop system of "perception-estimation-decision-control". Summary of the Invention

[0007] The purpose of this invention is to provide a vision-based distributed spacecraft on-orbit pursuit and escape game control method and system. The method achieves target state estimation through multi-spacecraft visual perception and feature matching, and combines pursuit and escape game modeling and generative adversarial imitation learning to train the strategy network, ultimately forming a perception-decision-control closed-loop distributed spacecraft on-orbit pursuit and control method.

[0008] To achieve the above objectives, the present invention provides the following technical solution: A vision-based distributed spacecraft on-orbit pursuit and escape game control method includes the following steps: S1. Multiple orbital target image sequences are acquired using cameras mounted forward on multiple pursuing spacecraft. Each orbital target image sequence is preprocessed and input into an unsupervised semantic segmentation model to generate a target semantic mask. Candidate regions of target spacecraft in the target image sequence are selected based on the morphological features of the mask region. S2. Use the SuperPoint algorithm to extract feature points within the candidate region; and use the SuperGlue algorithm to match the feature points within the candidate region based on structural semantic matching constraints to obtain feature matching point pairs, and evaluate the matching confidence of each feature matching point pair. S3. Using high-confidence feature matching point pairs, the eigenvalue matrix between orbital target images is estimated using the eight-point method; the relative pose of the pursuing spacecraft relative to the target spacecraft is calculated by using singular value decomposition of the eigenvalue matrix; and the nine-dimensional state vector of the pursuing spacecraft is fused by a multi-source information fusion mechanism to output a joint nine-dimensional state vector. S4. Establish an orbital pursuit and escape game model using a joint nine-dimensional state vector; establish the system state equation based on the relative motion dynamics equation, solve the optimal control law with the relative state of the pursuing spacecraft as feedback, and generate an expert strategy trajectory. S5. Model the orbital pursuit game problem as a Markov decision process; use expert policy trajectories as training samples, and introduce a generative adversarial imitation learning algorithm to train the control policy network; S6. Deploy the trained strategy network to the spacecraft attitude and orbit control system, output orbit control commands based on the real-time estimated target state, control the spacecraft actuators to complete the operation, and restart image acquisition to enter the next control cycle, forming a perception-decision-control closed loop.

[0009] Furthermore, in S1, the preprocessing process includes: firstly, compressing the color images acquired by multiple pursuing spacecraft into single-channel data by grayscale conversion to reduce the amount of computation and retain light intensity information; then, using histogram equalization to improve image contrast and make the edges of the target image clearer; then, eliminating random noise in the image by Gaussian filtering and using an edge enhancement algorithm to strengthen the boundary information of the target image.

[0010] Furthermore, in S1, the masked area includes: the visible structures of the target spacecraft's main body, solar panels, and modules; The morphological features of the masked region include: the area of ​​the masked region, edge closure, and geometric proportions.

[0011] Furthermore, in S3, the relative pose includes: a relative rotation matrix and a translation vector; The nine-dimensional state vector includes the three-dimensional position, three-dimensional velocity, and attitude angle of the pursuing spacecraft.

[0012] Furthermore, the expert strategy trajectory includes a zero-sum game performance index function of pursuit distance, control cost, and escape incentive.

[0013] Furthermore, in S5, the state of the Markov decision process is defined as the relative position and velocity of the pursuing spacecraft, the action is the direction and magnitude of acceleration, and the state transition is given by the orbital dynamics equations.

[0014] Furthermore, in S4, the process of establishing the system state equation includes: in the LVLH coordinate system, based on the relative motion dynamics model between the pursuing spacecraft and the target spacecraft, constructing the system state equation; wherein, the state variables are defined as the relative position and relative velocity between the pursuing spacecraft and the target spacecraft; the control variables include the thrust direction angle and acceleration magnitude of the pursuing spacecraft, and the thrust direction angle and acceleration magnitude of the target spacecraft; Assume the pursuing side controls... The escapee is under control. The system status is Then the dynamic equations satisfy the following form:

[0015] in, The system matrix is ​​derived from the linearization of the CW equation. , These are the control input matrices for the pursuing and escaping sides, respectively; To quantify the adversarial objectives of the pursuer and the fleeing, a zero-sum game performance index function is established:

[0016] in, To determine the distance for pursuit, To control the cost for the pursuing side, The escaping party aims to control "negative gains"; the pursuing party's goal is to minimize them. The escapee's objective is to maximize This leads to a saddle point solution problem.

[0017] Furthermore, in S5, the control policy network is optimized through deterministic policy gradient based on the reward signal provided by the discriminator; The discriminator distinguishes between expert and policy trajectories by maximizing the classification log-likelihood loss function.

[0018] Furthermore, the maximum classification log-likelihood loss function is:

[0019] The training objective of the policy network is to maximize the reward provided by the discriminator.

[0020] In this system, the output of the discriminator is treated as a proxy reward signal, replacing the explicit reward function in reinforcement learning; the policy network adopts a deterministic policy gradient agent, and the discriminator and generator update their parameters in an alternating optimization manner until convergence.

[0021] This invention also provides a system for implementing a vision-based distributed spacecraft on-orbit pursuit and escape game control method, comprising: The image acquisition and preprocessing module is used to acquire multiple orbital target image sequences using cameras mounted forward on multiple pursuing spacecraft. After preprocessing each orbital target image sequence, it is input into an unsupervised semantic segmentation model to generate a target semantic mask. Based on the morphological features of the mask region, candidate regions of the target spacecraft in the target image sequence are selected. The image feature point extraction and matching module is used to extract feature points within candidate regions using the SuperPoint algorithm; and to match feature points within candidate regions using the SuperGlue algorithm based on structural semantic matching constraints, thereby obtaining feature matching point pairs and evaluating the matching confidence of each feature matching point pair. The pose estimation module is used to match point pairs using high-confidence features and estimate the intrinsic matrix between orbital target images using the eight-point method; it also calculates the relative pose of the pursuing spacecraft with respect to the target spacecraft by decomposing the intrinsic matrix using singular value decomposition; and it fuses the nine-dimensional state vector of the pursuing spacecraft through a multi-source information fusion mechanism to output a joint nine-dimensional state vector. The game modeling module is used to establish an orbital pursuit and escape game model using a joint nine-dimensional state vector; it establishes the system state equation based on the relative motion dynamics equation, solves the optimal control law with the relative state of the pursuing spacecraft as feedback, and generates an expert strategy trajectory. The strategy generation and attitude / orbit execution module is used to model the orbital pursuit and escape game problem as a Markov decision process. It uses expert policy trajectories as training samples and introduces a generative adversarial imitation learning algorithm to train the control policy network. The trained policy network is then deployed in the spacecraft attitude / orbit control system. Based on the real-time estimated target state, it outputs orbit control commands to control the spacecraft actuators to complete the operation and restarts image acquisition to enter the next control cycle, forming a perception-decision-control closed loop.

[0022] According to specific embodiments provided by the present invention, the present invention discloses the following technical effects: This invention deeply integrates cutting-edge technologies such as unsupervised visual perception, high-precision feature matching, multi-source information fusion, pursuit-escape game modeling, and generative adversarial learning to construct a complete intelligent closed-loop system from perception to decision-making to control. It achieves high-precision autonomous identification, robust state estimation, and intelligent pursuit control of target spacecraft in harsh space environments such as non-cooperative, highly dynamic, and complex lighting conditions. This significantly improves the autonomy, adaptability, and mission success rate of spacecraft on-orbit operations, providing strong technical support for future challenging tasks such as on-orbit space servicing, debris removal, and space attack and defense. Attached Figure Description

[0023] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.

[0024] The vision-based distributed spacecraft on-orbit pursuit and escape game control method and system of the present invention will be further described below with reference to the accompanying drawings. Figure 1 This is a schematic diagram of the vision-based distributed spacecraft on-orbit pursuit and escape game control method of the present invention; Figure 2 This is a schematic diagram of the target segmentation and registration process in Embodiment 1 of the present invention; Figure 3 This is a schematic diagram of the feature point extraction and inter-frame matching structure in Embodiment 1 of the present invention; Figure 4 This is a schematic diagram of the three-dimensional trajectory of the pursuit spacecraft in Embodiment 1 of the present invention; Figure 5 This is a schematic diagram of the GAIL strategy training structure in Embodiment 1 of the present invention. Detailed Implementation

[0025] The specific embodiments of the present invention will be described in further detail below with reference to the accompanying drawings and examples. The following examples are for illustrative purposes only and are not intended to limit the scope of the invention.

[0026] To better understand the purpose, structure, and function of this invention, the invention will be described in further detail below with reference to the accompanying drawings.

[0027] Example 1 like Figure 1 As shown, this invention provides a vision-based distributed spacecraft on-orbit pursuit and escape game control method, characterized by the following steps: S1. Multiple orbital target image sequences are acquired using cameras mounted forward on multiple pursuing spacecraft. After preprocessing, these multiple orbital target image sequences are input into an unsupervised semantic segmentation model to generate target semantic masks. Candidate regions of target spacecraft in the target image sequences are selected based on the morphological features of the mask regions. In S1, the masked region includes the visible structures of the spacecraft's main body, solar panels, and modules; the morphological characteristics of the masked region include the area, edge closure, and geometric proportions of the masked region.

[0028] In S1, the preprocessing process includes: first, compressing the color images acquired by each pursuing spacecraft into single-channel data by grayscale conversion to reduce the amount of computation and retain light intensity information; then, using histogram equalization to improve image contrast and make the target edges clearer; then, eliminating random noise in the image by Gaussian filtering and using an edge enhancement algorithm to strengthen the boundary information of the target structure.

[0029] This embodiment is specifically as follows: 1) Multiple pursuing spacecraft, each equipped with a camera, acquire deep-space image sequences at a fixed frequency, obtaining visible light images containing the target spacecraft from multiple perspectives. This multi-view, distributed image acquisition method improves the coverage and recognition stability of structural information of non-cooperative targets. To adapt to the spatial imaging characteristics of small spacecraft size, complex backgrounds, and uneven illumination in the images, the images undergo preprocessing steps such as grayscale conversion, histogram equalization, noise filtering, and edge enhancement to improve image quality and structural resolution, providing stable input for subsequent segmentation and estimation, such as... Figure 2 As shown; 2) Image segmentation and target candidate region extraction for subsequent feature extraction and state estimation. Each spacecraft inputs the preprocessed target image into an unsupervised semantic segmentation model. The unsupervised semantic segmentation model automatically identifies different semantic regions in the target image and generates multiple masks in the target image to locate the possible position and structural range of the spacecraft. These masks typically cover the visible structures of the spacecraft, such as the main body, solar panels, and modules. This method has zero-shot generalization capability and can automatically identify target spacecraft regions in the image without prior templates. It also filters regions based on geometric analysis, edge sharpness, and central symmetry, retaining candidate regions for tracking and matching, such as... Figure 3 As shown.

[0030] By analyzing the area, edge closure, and geometric proportions of the mask region, the mask with the largest area and smoothest boundaries is selected as the effective target region. To further improve robustness, mask extraction results can be shared among different frames for structural consistency verification. Subsequently, the system calculates the center point, boundary contour, and image content of the region, providing this information to subsequent modules for visual tracking. To enhance recognition stability, the system also tracks the centroid trajectory of the mask region across multiple frames and eliminates target regions that are discontinuous or abrupt in recognition.

[0031] S2. Use the SuperPoint algorithm to extract feature points within the candidate region; and use the SuperGlue algorithm to match the feature points within the candidate region based on structural semantic matching constraints to obtain feature matching point pairs, and evaluate the matching confidence of each feature matching point pair.

[0032] This embodiment specifically involves: extracting image feature points from candidate regions and performing inter-frame matching. Each spacecraft applies the SuperPoint algorithm to extract key points and descriptors from the feature points, ensuring feature stability under complex conditions such as illumination changes, viewpoint rotation, and image blurring. Using the structural semantic matching results as constraints, matching feature points and their confidence scores are extracted from different semantic matching regions. Finally, the feature points from different regions are integrated, and the confidence scores of each feature point matching result are evaluated. The final matching result is as follows: Figure 2 As shown.

[0033] S3. High-confidence feature matching is used to pair points, and the eigenvalue matrix is ​​estimated using the eight-point method. The relative pose of the pursuing spacecraft to the target spacecraft is calculated using singular value decomposition. The nine-dimensional state vector of the pursuing spacecraft is fused using a multi-source information fusion mechanism to output a joint nine-dimensional state vector. This nine-dimensional state vector includes: three-dimensional position, three-dimensional velocity, and attitude angles, such as... Figure 4 As shown.

[0034] This embodiment is specifically as follows: Based on region extraction, each spacecraft independently employs a feature matching algorithm to perform cross-frame matching of key points in candidate regions, constructing a correspondence between feature points in the images. This process, combining local descriptors and context graph structure modeling, significantly improves the robustness and accuracy of feature matching, making it suitable for orbital target tracking scenarios with large attitude changes and significant image distortion. Each spacecraft uploads the matched point pairs to the system center or shared node, fusing them into the core observations for state estimation. A relative attitude and position state estimation method based on eigenvalue decomposition is proposed. By calculating the eigenvalue matrix for feature point pairs and combining it with the RANSAC algorithm to remove outliers, the rotation matrix and translation direction between spacecraft are recovered in the two-dimensional image domain. Combining the intrinsic parameters of each camera, triangulation and singular value decomposition are used to achieve relative pose reconstruction in three-dimensional space. By fusing the estimation results from multiple spacecraft perspectives, a complete state vector including rotation angles (pitch, yaw, roll) and spatial positions (X, Y, Z) is obtained.

[0035] S4. Establish an orbital pursuit-escape game model using a joint nine-dimensional state vector; establish the system state equation based on the relative motion dynamics equation, solve the optimal control law with the relative state of the pursuing spacecraft as feedback, and generate an expert strategy trajectory; wherein the expert strategy trajectory includes: a zero-sum game performance index function of pursuit-escape distance, control cost, and escape incentive.

[0036] This embodiment specifically involves estimating the target's three-dimensional spatial pose state to provide accurate input state quantities for generating game strategies. Using the feature matching point pairs and their matching confidence levels obtained in step S2, the target's position and orientation from the perspective of each pursuing spacecraft are inferred through a camera geometric model. First, matching point pairs are divided according to their confidence levels, and a few high-confidence matching point pairs are extracted to calculate the intrinsic matrix between images. Furthermore, the RANSAC algorithm is used to eliminate outliers, improving computational efficiency while ensuring accuracy. The relative rotation matrix and translation vector are calculated using Singular Value Decomposition (SVD). To further enhance accuracy and fault tolerance, the local estimation results from multiple spacecraft can be fused through a fusion mechanism to form a joint nine-dimensional state vector, including three-dimensional position, three-dimensional velocity, and attitude angles, which serves as input to the game theory model.

[0037] S5. Model the orbital pursuit game problem as a Markov decision process; use expert policy trajectories as training samples, and introduce a generative adversarial imitation learning algorithm to train the control policy network, such as... Figure 5 As shown; This embodiment specifically involves establishing an orbital pursuit-escape game model and generating an expert strategy trajectory based on physical dynamic constraints. This invention models the multi-spacecraft pursuit-escape process as a continuous-time, zero-sum dynamic game problem, employing differential game theory to characterize the adversarial behavior between the pursuing and escaping spacecraft. To accurately describe the dynamic evolution of the orbit, 1) The state of a Markov decision process is defined as the relative position and velocity of the spacecraft, the action is the direction and magnitude of acceleration, and the state transition is given by the orbital dynamics equations.

[0038] In S5, the relative motion dynamics model in the LVLH coordinate system is selected as the system state equation. The state variables are the position and velocity difference between the pursuing and fleeing parties in three-dimensional space, and the control variables are the thrust direction angle and acceleration magnitude of each party. Assume the pursuing side controls... The escapee is under control. The system status is Then the dynamic equations satisfy the following form:

[0039] in, The system matrix is ​​derived from the linearization of the CW equation. , These are the control input matrices for the pursuing and escaping sides, respectively; To quantify the adversarial objectives of the pursuer and the fleeing party, a zero-sum game performance index function is established:

[0040] in, To determine the distance for pursuit, To control the cost for the pursuing side, The goal of the fleeing party is to control the "negative benefit" (i.e., incentivize them to flee); the goal of the pursuing party is to minimize the negative benefit. The escapee's objective is to maximize This leads to a saddle point solution problem.

[0041] In S5, the discriminator distinguishes between experts and policy trajectories by maximizing the classification log-likelihood loss function, and the policy network optimizes the policy gradient based on the reward signal provided by the discriminator.

[0042] The training objective of the discriminator is to maximize the classification log-likelihood loss function.

[0043] The training objective of the policy network is to maximize the reward provided by the discriminator.

[0044] In this system, the output of the discriminator is treated as a proxy reward signal, replacing the explicit reward function in reinforcement learning. To improve training stability, the policy network adopts a deterministic policy gradient (DDPG) agent, and the discriminator and generator update parameters in an alternating optimization manner until convergence.

[0045] 2) The process of using expert policy trajectories as training samples and introducing a generative adversarial imitation learning algorithm to train the control policy network is as follows: By imitation learning, the game strategy is approximated and trained to form a lightweight and efficient control strategy network. First, the orbital pursuit game problem is modeled as a Markov decision process, whose state... Indicates the relative position and velocity of the spacecraft, and its movements. The direction and magnitude of its acceleration are represented by the orbital dynamics equations, and the state transition is given by these equations. This is used to train the control policy network. A large amount of expert demonstration data was generated using differential game solutions. As a model to imitate, construct a discriminator network. With policy network Its training objectives are as follows: The training objective of the discriminator is to maximize the classification log-likelihood loss function.

[0046] The training objective of the policy network is to maximize the reward provided by the discriminator.

[0047] In this model, the discriminator's output is treated as a proxy reward signal, replacing the explicit reward function in reinforcement learning. To improve training stability, the policy network employs a deterministic policy gradient (DDPG) agent, with the discriminator and generator updating parameters alternately until convergence.

[0048] To reduce computational complexity, in a "one-on-one pursuit" scenario, both sides share the same policy network; in a "multiple pursuits, one escape" scenario, the pursuer uses a centralized network, while the escapee uses independent networks trained separately. The resulting policy network can be run online based on the input state. Real-time output control action To achieve closed-loop optimal decision control in orbital game theory. S6. The trained strategy network is deployed in the attitude and orbit control systems of each spacecraft, and a real-time visual perception closed loop is constructed to achieve distributed, purely vision-driven collaborative pursuit and escape control. After processing each frame of image, each pursuing spacecraft inputs the estimated local target state or fused state vector into the strategy network, and quickly outputs orbit control commands such as the current thrust direction, magnitude, and attitude adjustment angular rate. The control commands are converted and input to the spacecraft's AOCS system through an interface, and the actual operation is completed by devices such as reaction wheels and thruster arrays. After the operation is completed, the system immediately restarts image acquisition and enters the next control cycle. Multiple spacecraft can perform soft synchronization based on low-frequency communication or state prediction to ensure coordinated strategy execution and form a stable and efficient perception-decision-control closed loop process. Experiments show that this method still has excellent performance under conditions such as image occlusion, strong interference, and distributed detection, supporting the system's scalability and practical deployment capabilities.

[0049] In summary, this invention constructs an orbital pursuit-escape game control model based on differential games. First, the model models multiple pursuing spacecraft and the target spacecraft as a many-to-one adversarial system. Considering factors such as orbital relative dynamics, thrust constraints, communication delays, and target response strategies, a continuous-time zero-sum dynamic game system is established. By solving the Hamilton-Jacobi-Isaacs equations, an expert trajectory set describing the optimal strategy for multi-spacecraft cooperative pursuit is generated. This is to improve real-time performance and the flexibility of distributed system deployment.

[0050] This invention also proposes a policy network training method based on Generative Adversarial Imitation Learning (GAIL). Using expert trajectories generated by differential games as supervision signals, a deep policy network model including a generator and a discriminator is trained. Each spacecraft can independently deploy this network and perform decision-making inference based on local observations and collaborative information, achieving accuracy that highly approximates expert policies while reducing real-time computational burden and meeting the deployment requirements of onboard computing platforms. Finally, this invention designs an attitude and orbit control parsing module based on visual estimation of state and policy network output. This module, considering spacecraft thruster capabilities, attitude adjustment constraints, and mission requirements, parses the high-dimensional action vector output by the policy network into attitude adjustment angular rates and orbital propulsion commands, which are then input into the attitude and orbit control systems of each spacecraft for execution, and images are reacquired to form their respective control loops. The distributed structure enables the system to possess redundancy, fault tolerance, and tactical coordination capabilities, ensuring continuous optimization and dynamic adjustment of the pursuit and escape strategy. The method in this invention divides the pursuit mission into two main stages: the first is an image-based target recognition and relative state estimation module, and the second is a multi-agent pursuit game control strategy module based on differential games and imitation learning. Unlike traditional methods relying on a single spacecraft for observation and control, this invention introduces a multi-spacecraft distributed image information acquisition and collaborative processing mechanism in the visual perception stage. Multiple pursuit spacecraft simultaneously acquire image information from different perspectives, and the target's three-dimensional pose state is jointly estimated through image features, achieving robust enhancement of redundant observation and attitude recovery, providing more accurate input for subsequent control. Compared with existing technologies, it also has the following technical advantages: (1) A novel distributed spacecraft pursuit and escape control framework based on pure optical information is proposed to replace the traditional complex system that relies on multiple source sensors such as radar, laser ranging, and GNSS, achieving fully autonomous image navigation control under extreme conditions. This method supports multiple spacecraft to collaboratively acquire image information, construct space target structures from multiple perspectives, and achieve more robust target perception and relative pose reconstruction. It is particularly suitable for challenging mission scenarios such as non-cooperative targets, unmanned deep space approach, and complex lighting environments. The system architecture is modular, with low communication burden and flexible deployment, and has good scalability and engineering practicality.

[0051] (2) An orbital game modeling method based on differential game theory is adopted to model the spacecraft pursuit and escape problem as a continuous-time dynamic many-to-one zero-sum game process. It fully considers the target escape behavior, the cooperative constraints between pursuing spacecraft, the orbit control response characteristics, and the spatiotemporal evolution relationship, and can obtain expert strategy trajectories with global optimality. At the same time, the generative adversarial imitation learning method (GAIL) is introduced to train the policy network using the results of differential game theory, so as to realize the generation of high-precision and high-robust control strategies, and solve the key problems of high computational complexity and poor real-time performance of traditional orbital game theory in multi-spacecraft environments.

[0052] (3) A multi-spacecraft collaborative full-process control system consisting of five parts—visual observation, state estimation, game modeling, strategy control, and closed-loop execution—was constructed, forming a distributed autonomous closed loop of "image input-state recovery-strategy generation-orbit control output-re-observation update," ensuring that multiple spacecraft collaboratively execute game-based pursuit missions in complex dynamic environments. The system possesses high real-time performance, redundancy, and responsiveness, and can adapt to complex operating conditions such as drastic changes in target attitude, frequent obstruction interference, and nonlinear orbital dynamics.

[0053] (4) The overall method in this invention has good embedded deployment capabilities. All algorithm modules adopt lightweight visual models and neural network structures. The policy network supports distributed operation without large-scale online iteration, making it easy to integrate into computing platforms such as onboard SoCs, embedded GPUs, or FPGAs of multiple spacecraft. This method is applicable to various platforms with computing power and power consumption constraints, such as small spacecraft formations, autonomous satellite clusters, and space exploration collaborations, and has good prospects for engineering promotion.

[0054] (5) This invention provides an intelligent, lightweight, and scalable solution for pursuit and escape control tasks involving pure vision and multi-spacecraft collaboration. It breaks the dependence of existing systems on high-end sensors and high-bandwidth communication links, improves the survivability, autonomy, and multi-spacecraft collaborative mission completion rate of spacecraft in complex space environments, and has significant military strategic value and deep space exploration application potential. Example 2 The present invention also provides a system for implementing a vision-based distributed spacecraft on-orbit pursuit and escape game control method as described in Embodiment 1, comprising: The image acquisition and preprocessing module is used to acquire multiple orbital target image sequences using cameras mounted forward on multiple pursuing spacecraft, and input the preprocessed multiple orbital target image sequences into an unsupervised semantic segmentation model to generate target semantic masks; and filter out candidate regions of target spacecraft in the target image sequences based on the morphological features of the mask regions. The image feature point extraction and matching module is used to extract feature points within candidate regions using the SuperPoint algorithm; and to match feature points within candidate regions using the SuperGlue algorithm based on structural semantic matching constraints, thereby obtaining feature matching point pairs and evaluating the matching confidence of each feature matching point pair. The pose estimation module is used to match point pairs using high-confidence features and estimate the intrinsic matrix using the eight-point method; it also calculates the relative pose of the pursuing spacecraft relative to the target spacecraft through singular value decomposition; and it fuses the nine-dimensional state vector of the pursuing spacecraft through a multi-source information fusion mechanism to output a joint nine-dimensional state vector. The game modeling module is used to establish an orbital pursuit and escape game model using a joint nine-dimensional state vector; it establishes the system state equation based on the relative motion dynamics equation, solves the optimal control law with the relative state of the pursuing spacecraft as feedback, and generates an expert strategy trajectory. The strategy generation and attitude / orbit execution module is used to model the orbital pursuit and escape game problem as a Markov decision process. It uses expert policy trajectories as training samples and introduces a generative adversarial imitation learning algorithm to train the control policy network. The trained policy network is then deployed in the spacecraft attitude / orbit control system. Based on the real-time estimated target state, it outputs orbit control commands to control the spacecraft actuators to complete the operation and restarts image acquisition to enter the next control cycle, forming a perception-decision-control closed loop.

[0055] The above description of the disclosed embodiments enables those skilled in the art to make or use the present invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the invention. Therefore, the present invention is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A vision-based distributed spacecraft on-orbit pursuit and escape game control method, characterized in that, Includes the following steps: S1. Use cameras mounted forward on multiple pursuing spacecraft to acquire multiple orbital target image sequences, and input each orbital target image sequence into an unsupervised semantic segmentation model to generate target semantic masks; Candidate regions for target spacecraft in target image sequences are selected based on the morphological features of the mask region. S2. Use the SuperPoint algorithm to extract feature points within the candidate region; and use the SuperGlue algorithm to match the feature points within the candidate region based on structural semantic matching constraints to obtain feature matching point pairs, and evaluate the matching confidence of each feature matching point pair. S3. Using high-confidence feature matching point pairs, the eigenvalue matrix between orbital target images is estimated using the eight-point method; the relative pose of the pursuing spacecraft relative to the target spacecraft is calculated by using singular value decomposition of the eigenvalue matrix; and the nine-dimensional state vector of the pursuing spacecraft is fused by a multi-source information fusion mechanism to output a joint nine-dimensional state vector. S4. Establish an orbital pursuit and escape game model using a joint nine-dimensional state vector; establish the system state equation based on the relative motion dynamics equation, solve the optimal control law with the relative state of the pursuing spacecraft as feedback, and generate an expert strategy trajectory. S5. Model the orbital pursuit game problem as a Markov decision process; use expert policy trajectories as training samples, and introduce a generative adversarial imitation learning algorithm to train the control policy network; S6. Deploy the trained strategy network to the spacecraft attitude and orbit control system, output orbit control commands based on the real-time estimated target state, control the spacecraft actuators to complete the operation, and restart image acquisition to enter the next control cycle, forming a perception-decision-control closed loop.

2. The vision-based distributed spacecraft on-orbit pursuit and escape game control method according to claim 1, characterized in that, In S1, the preprocessing process includes: first, compressing the color images acquired by multiple pursuing spacecraft into single-channel data by grayscale conversion to reduce the amount of computation and retain light intensity information; then, using histogram equalization to improve image contrast and make the edges of the target image clearer; then, eliminating random noise in the image by Gaussian filtering and using an edge enhancement algorithm to strengthen the boundary information of the target image.

3. The vision-based distributed spacecraft on-orbit pursuit and escape game control method according to claim 1, characterized in that, In S1, the masked area includes the visible structures of the target spacecraft's main body, solar panels, and modules; The morphological features of the masked region include: the area of ​​the masked region, edge closure, and geometric proportions.

4. The vision-based distributed spacecraft on-orbit pursuit and escape game control method according to claim 1, characterized in that, In S3, the relative pose includes: a relative rotation matrix and a translation vector; The nine-dimensional state vector includes the three-dimensional position, three-dimensional velocity, and attitude angle of the pursuing spacecraft.

5. The vision-based distributed spacecraft on-orbit pursuit and escape game control method according to claim 1, characterized in that, The expert strategy trajectory includes a zero-sum game performance index function of pursuit distance, control cost, and escape incentive.

6. The vision-based distributed spacecraft on-orbit pursuit and escape game control method according to claim 1, characterized in that, In S5, the state of the Markov decision process is defined as the relative position and velocity of the pursuing spacecraft, the action is the direction and magnitude of acceleration, and the state transition is given by the orbital dynamics equations.

7. The vision-based distributed spacecraft on-orbit pursuit and escape game control method according to claim 1, characterized in that, In S4, the process of establishing the system state equation includes: in the LVLH coordinate system, based on the dynamic model of the relative motion between the pursuing spacecraft and the target spacecraft, the system state equation is constructed; wherein, the state variables are defined as the relative position and relative velocity between the pursuing spacecraft and the target spacecraft; the control variables include the thrust direction angle and acceleration magnitude of the pursuing spacecraft, and the thrust direction angle and acceleration magnitude of the target spacecraft; Assume the pursuing side controls... The escapee is under control. The system status is Then the dynamic equations satisfy the following form: in, The system matrix is ​​derived from the linearization of the CW equation. , These are the control input matrices for the pursuing and escaping sides, respectively; To quantify the adversarial objectives of the pursuer and the fleeing, a zero-sum game performance index function is established: in, To determine the distance for pursuit, To control the cost for the pursuing side, The escaping party aims to control "negative gains"; the pursuing party's objective is to minimize them. The escapee's objective is to maximize This leads to a saddle point solution problem.

8. The vision-based distributed spacecraft on-orbit pursuit and escape game control method according to claim 1, characterized in that, In S5, the control policy network is optimized through deterministic policy gradient based on the reward signal provided by the discriminator; The discriminator distinguishes between expert and policy trajectories by maximizing the classification log-likelihood loss function.

9. The vision-based distributed spacecraft on-orbit pursuit and escape game control method according to claim 8, characterized in that, The maximum classification log-likelihood loss function: The training objective of the policy network is to maximize the reward provided by the discriminator. In this system, the output of the discriminator is treated as a proxy reward signal, replacing the explicit reward function in reinforcement learning; the policy network adopts a deterministic policy gradient agent, and the discriminator and generator update their parameters in an alternating optimization manner until convergence.

10. A vision-based distributed spacecraft on-orbit pursuit and escape game control system, used to implement the vision-based distributed spacecraft on-orbit pursuit and escape game control method as described in any one of claims 1-9, characterized in that, include: The image acquisition and preprocessing module is used to acquire multiple orbital target image sequences using cameras mounted forward on multiple pursuing spacecraft, and input each orbital target image sequence into an unsupervised semantic segmentation model to generate a target semantic mask. Candidate regions for target spacecraft in target image sequences are selected based on the morphological features of the mask region. The image feature point extraction and matching module is used to extract feature points within candidate regions using the SuperPoint algorithm; and to match feature points within candidate regions using the SuperGlue algorithm based on structural semantic matching constraints, thereby obtaining feature matching point pairs and evaluating the matching confidence of each feature matching point pair. The pose estimation module is used to match point pairs using high-confidence features and estimate the intrinsic matrix between orbital target images using the eight-point method; it also calculates the relative pose of the pursuing spacecraft with respect to the target spacecraft by decomposing the intrinsic matrix using singular value decomposition; and it fuses the nine-dimensional state vector of the pursuing spacecraft through a multi-source information fusion mechanism to output a joint nine-dimensional state vector. The game modeling module is used to establish an orbital pursuit and escape game model using a joint nine-dimensional state vector; it establishes the system state equation based on the relative motion dynamics equation, solves the optimal control law with the relative state of the pursuing spacecraft as feedback, and generates an expert strategy trajectory. The strategy generation and attitude / orbit execution module is used to model the orbital pursuit and escape game problem as a Markov decision process. It uses expert policy trajectories as training samples and introduces a generative adversarial imitation learning algorithm to train the control policy network. The trained policy network is then deployed in the spacecraft attitude / orbit control system. Based on the real-time estimated target state, it outputs orbit control commands to control the spacecraft actuators to complete the operation and restarts image acquisition to enter the next control cycle, forming a perception-decision-control closed loop.