Optimal control method for cooperative tracking of non-cooperative target by cluster spacecraft
Patent Information
- Application Number
- CN202510382712.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-28
- Publication Date
- 2026-08-28
- Estimated Expiration
- 2045-03-28
AI Technical Summary
在实际应用中,许多任务中的目标难以准确建模,其轨迹随着时间发生不确定性变化,这使得精确建模变得愈加困难
(1)无需精确系统模型:克服了传统的最优控制方法需要精确的系统模型的缺陷;本申请通过数据驱动的方法,无需对系统完全了解,仅依赖状态测量数据(追方状态信息和相对状态信息)、初始控制策略、探索信号即可实现最优控制。
Smart Images

Figure CN120462660B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of spacecraft pursuit and escape game control technology, and specifically relates to an optimal control method for swarm spacecraft to collaboratively track non-cooperative targets. Background Technology
[0002] In space missions, coordinated control of swarm spacecraft has become a crucial means of achieving complex space tasks. Especially when tracking non-cooperative targets (such as spacecraft with unknown trajectories or behaviors, space debris or fragments, and other celestial bodies), swarm spacecraft need to coordinate efficiently without prior information. The dynamic characteristics of non-cooperative targets are typically highly uncertain. In some cases, targets may not have a definite trajectory, or their motion patterns may be difficult to predict. When performing missions, swarm spacecraft often need to track targets in real time without fully understanding their state, requiring control strategies to remain effective in highly uncertain environments. Therefore, how to achieve real-time swarm coordination without a precise system model has become a significant challenge in current space missions. Furthermore, during the coordinated tracking of non-cooperative targets, the trackers will eventually approach the target and reach the desired tracking position, thus necessitating a strategy that can avoid the risk of collision.
[0003] Existing traditional optimal control methods (such as the classic linear quadratic regulator LQR) typically rely on precise mathematical models of the system and the target to design globally optimal or near-optimal control strategies for the system. Furthermore, existing technologies apply Adaptive Dynamic Programming (ADP) to the game-theoretic control of spacecraft. ADP is a tool for solving optimal control problems in nonlinear systems. For example, patent document CN202310866825.0 discloses a method for solving spacecraft pursuit-escape game control based on adaptive dynamic programming. This method indirectly derives the control law by estimating the optimal value function, thus solving the problems of complex solution processes and long solution times in single-pair spacecraft game control strategies.
[0004] Existing methods have the following main drawbacks when addressing the issue of collaborative tracking of non-cooperative targets by swarm spacecraft: (1) Strong dependence on system model: Traditional optimal control methods usually require accurate understanding of the target's dynamic model. In practical applications, the target in many tasks is difficult to model accurately, and its trajectory changes uncertainly over time, which makes accurate modeling increasingly difficult. For example, although patent document CN202310866825.0 uses the ADP method in the field of pursuit and escape game control, it still needs to be based on a linearized model of near-circular orbit, and the generation of the control law still depends on the target's dynamic model.
[0005] (2) Low efficiency: Traditional control methods mostly rely on global information processing. This method faces the problem of excessive computational complexity when the cluster size is large, especially in tasks that require fast response, where real-time performance often cannot be guaranteed.
[0006] (3) High computational cost: For large-scale spacecraft clusters, the need for collaborative control between spacecraft within the cluster further increases the computational burden. As the cluster size increases, the computational load grows exponentially, leading to inefficiency.
[0007] (4) Failure to consider safety requirements: Traditional optimal control methods usually focus on achieving efficient target tracking. However, in the process of spacecraft swarm collaboration, safety is often not fully considered. Especially in pursuit-escape game and target tracking missions, as the swarm spacecraft approach the target, collisions between spacecraft or between spacecraft and the target may occur. Most existing technologies focus on the control of pursuit-escape game with limited input, without solving the dynamic obstacle avoidance problem in multi-spacecraft collaboration, and without designing effective obstacle avoidance functions. As a result, it is difficult to balance efficient tracking and safety assurance in dynamic and complex environments. Summary of the Invention
[0008] To address the above issues, this invention provides a method for collaborative tracking of non-cooperative targets by a cluster of spacecraft. This method combines a data-driven adaptive dynamic programming approach with a cluster-based collaborative obstacle avoidance mechanism. It eliminates the reliance on dynamic target models and can directly address the dynamic uncertainties and complex trajectory changes of non-cooperative targets. Furthermore, it solves the collision risk and mission conflict issues in multi-spacecraft collaborative scenarios, achieving the dual goals of efficient tracking and safety assurance, and providing a more adaptable and reliable solution for complex space missions.
[0009] The specific technical solution is as follows: The optimal control method for cooperative tracking of non-cooperative targets by swarm spacecraft includes the following steps: S1. Construct a relative dynamic model between non-cooperative target spacecraft and swarm spacecraft; S2. Establish a cluster spacecraft obstacle avoidance model based on the relative dynamics model, and track the relative positions of spacecraft and target spacecraft, and spacecraft and the cluster interior by constraining the spherical exclusion zone and cylindrical form field; S3. Define the objective function of the game between the pursuer and the pursuer, introduce the obstacle avoidance model into the objective function, and design optimization objectives for fuel consumption and relative distance for the cluster spacecraft and the target spacecraft respectively; S4. Based on the aforementioned relative dynamics model and game objective function, construct the Hamiltonian function and HJI equation to establish a dynamic game model; S5. The optimal solution of the HJI equation is approximated online using the data-driven ADP method to solve the optimal game control strategy between the cluster spacecraft and the target spacecraft: by combining the value function and the control law, both the value function and the control law are expressed as a weighted sum of basis functions, and the parameters are directly optimized online using real-time state data, thereby completely decoupling the dependence on the target model; S6. A global escape strategy for the target spacecraft is generated based on a null-space projection method. Optimal escape control under multi-task conflicts is achieved through priority sorting and multiple projections.
[0010] Further, in step S1, the relative dynamic model is: , x e For the position and velocity vector of the non-cooperative target spacecraft, ; x pi For the first in the cluster The position and velocity vector of a spacecraft ; For non-cooperative target spacecraft and the first in the cluster Relative state quantities of a spacecraft ; , These are the instantaneous orbital angular velocity and angular acceleration of the reference spacecraft, respectively. The gravitational constant of Earth; , and These are the reference orbit radius, the target orbit radius, and the first orbital radius in the cluster. The orbital radius of a spacecraft; , , For the three-axis thrust acceleration of the target spacecraft, , , For the first in the cluster The three-axis thrust acceleration of the spacecraft; and The state-space representation is as follows , , in and Nonlinear functions representing the intrinsic dynamics of the swarm spacecraft and target spacecraft system. This represents the control input matrix for the spacecraft cluster. This represents the target spacecraft control input matrix; The equations of relative motion are obtained , The extended equation is defined as follows: , in, , , , .
[0011] Furthermore, in step S2, the obstacle avoidance model is implemented using a potential function, establishing a spherical no-go zone for the tracking spacecraft and a cylindrical situational field for the target spacecraft; the potential function for: , in, This represents the relative position components between the tracking spacecraft and the target spacecraft. This represents a metric parameter indicating the relative distance between tracking spacecraft. This represents the geometric constraints between the target spacecraft and the tracking spacecraft.
[0012] Furthermore, in step S3, For the first in the cluster The spacecraft establishes the following objective function. : , in, For the relative state quantities of the spacecraft, and These are the thrust acceleration control vectors for both the pursuing and fleeing parties. It is a positive semi-definite symmetric matrix. It is a positive definite symmetric matrix. The weighting ratio factor is controlled by both sides in the pursuit and recovery process; Let be the potential function. Represents a time variable; Target spacecraft versus tracking spacecraft Establish the following objective function : .
[0013] Further, in step S4, the Hamiltonian function is: , The HJI equation is: .
[0014] in, For the relative state quantities of the spacecraft, and These are the thrust acceleration control vectors for both the pursuing and fleeing parties. Represents a value function. Let be the potential function. It is a positive semi-definite symmetric matrix. It is a positive definite symmetric matrix. This represents the gradient of the value function with respect to the state variables. This represents the optimal value function.
[0015] Furthermore, in step S5, the optimal game control strategy between the cluster spacecraft and the target spacecraft is: , , in, The weighting ratio factor is controlled by both sides in the pursuit and capture process. This represents the gradient of the optimal value function with respect to the state variables; When using the ADP method to approximate the optimal solution of the HJI equation online, a method with given basis functions is adopted as a means of function approximation. The value function and control law are approximated as follows: , , , By minimizing Obtain the coefficient , , ,in , in , , These are basis functions, used to approximate the value function, the cluster spacecraft control strategy, and the target spacecraft control strategy, respectively. , , Let them be the number of their corresponding basis functions. , , These are the weight coefficients corresponding to their respective basis functions. and Indicates adjacent times and The system state variables, The state weight matrix is... For the input weight matrix, and These represent the perturbation correction terms for the cluster spacecraft and the target spacecraft, respectively. The weighting ratio factor is controlled by both sides in the pursuit and capture process. This is the Bellman error.
[0016] Furthermore, in step S5, the specific process of using the data-driven ADP method to approximate the optimal solution of the HJI equation online and solve the optimal game control strategy between the cluster spacecraft and the target spacecraft includes: System initialization: Set the initial state information of the spacecraft cluster; the initial state information includes the position and velocity of the tracking spacecraft, the initial state of the target spacecraft, and algorithm parameters; the algorithm parameters include the approximate order of the basis functions, the initial feasible control strategy, the exploration signal, and the convergence threshold; Data collection and online learning: Real-time acquisition of trajectory data of target spacecraft, combined with cluster communication mechanisms to share dynamic data; Control strategy update: Based on dynamic data, iteratively optimize the value function parameters and control strategy parameters to generate coordinated control input; Convergence criterion: The convergence of the algorithm is determined by comparing the change in the iteration results with a threshold. Output the optimal value function and control strategy.
[0017] Furthermore, the method further includes the step of: S6. A global escape strategy for the target spacecraft is generated based on a null-space projection method. Optimal escape control under multi-task conflicts is achieved through priority sorting and multiple projections.
[0018] Further, step S6 specifically includes: Prioritize the target spacecraft according to its relative distance to the cluster of spacecraft, from smallest to largest. The global escape policy is generated through n-1 projections, and is represented as follows: , And there are ,in This represents the initial projection control value. Indicates intermediate variables. This represents the quantity in the i-th projection process. This represents the target spacecraft control quantity corresponding to the maximum relative distance.
[0019] Furthermore, the method further includes the step of: S7. Verify the quality of the control strategy through performance evaluation indicators, including tracking accuracy, fuel consumption, robustness, and real-time performance; if the performance evaluation indicators do not meet the requirements, return to step S5.
[0020] Compared with the prior art, one or more of the above technical solutions can achieve at least one of the following beneficial effects: (1) No need for precise system model: This application overcomes the shortcomings of traditional optimal control methods that require precise system models. Through a data-driven approach, optimal control can be achieved without fully understanding the system, relying only on state measurement data (tracking state information and relative state information), initial control strategy, and exploration signals.
[0021] (2) Improved computational efficiency: Through adaptive dynamic programming technology, the dependence on global information is reduced. Each spacecraft only needs to rely on local information and information shared by the cluster, thereby improving computational efficiency and adapting to real-time requirements.
[0022] (3) Enhanced adaptability and robustness: The ADP method is combined to adaptively adjust the control strategy, handle the uncertainty and changes of the target trajectory, and avoid the limitations of the fixed control strategy in the traditional method.
[0023] (4) Improved collaborative tracking performance: Through distributed control and dynamic optimization, this application can better coordinate the collaborative tracking behavior of multiple spacecraft in the cluster, avoiding the inefficiency problem of individual spacecraft optimization in traditional methods.
[0024] (5) Ensuring the safety of mission execution: In terms of avoiding collisions and ensuring the stability of mission execution, this invention fully considers the safety of spacecraft. By introducing an obstacle avoidance model and no-go zone strategy for clustered spacecraft, the risk of collisions between spacecraft can be effectively prevented. In addition, by managing mission priorities through the null-space projection method, multi-mission conflicts are avoided, ensuring the feasibility of the system in complex environments; this method can dynamically adjust mission execution strategies in real time, ensure the safe distance between spacecraft, and respond quickly in the event of emergencies, effectively improving the safety of the system. Attached Figure Description
[0025] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0026] Figure 1 This is a flowchart illustrating the optimal control method for collaborative tracking of non-cooperative targets by swarm spacecraft based on ADP in Example 1.
[0027] Figure 2 This is a schematic diagram of the collaborative tracking of a non-cooperative target by a cluster of spacecraft in Example 1. Detailed Implementation
[0028] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention. Example 1
[0029] This embodiment provides an optimal control method for swarm spacecraft to cooperatively track non-cooperative targets, including the following steps: S1. Construct a relative dynamic model of non-cooperative target spacecraft and swarm spacecraft.
[0030] Specifically, in the problem of swarm spacecraft cooperatively tracking a non-cooperative target spacecraft, the non-cooperative target spacecraft and the... The dynamic equations of a cluster of spacecraft based on a nonlinear relative motion model can be described as follows: , x e For the position and velocity vector of the non-cooperative target spacecraft, ; x pi For the first in the cluster The position and velocity vector of a spacecraft ; For non-cooperative target spacecraft and the first in the cluster Relative state quantities of a spacecraft ; , These are the instantaneous orbital angular velocity and angular acceleration of the reference spacecraft, respectively. The gravitational constant of Earth; , and These are the reference orbit radius, the target orbit radius, and the first orbital radius in the cluster. The orbital radius of a spacecraft; , , , , , These are the target spacecraft and the first in the cluster. The three-axis thrust acceleration of the spacecraft makes and The state-space representation is as follows , , in and Nonlinear functions representing the intrinsic dynamics of the swarm spacecraft and target spacecraft system. This represents the control input matrix for the spacecraft cluster. This represents the target spacecraft control input matrix.
[0031] Further, the equations of relative motion are obtained. , The extended equation is defined as follows: , in , , , .
[0032] S2. Establish a cluster spacecraft obstacle avoidance model based on the relative dynamics model, and track the relative positions of spacecraft and target spacecraft, and spacecraft and the cluster interior by using spherical restricted areas and cylindrical form field constraints.
[0033] In some implementations, for mission safety reasons, a spherical exclusion zone is established around the tracking spacecraft to prevent collisions within the formation, and a cylindrical potential field is established around the target spacecraft to enclose the tracking spacecraft body. The potential function can be expressed as: , in, This represents the relative position components between the tracking spacecraft and the target spacecraft. This represents a metric parameter indicating the relative distance between tracking spacecraft. This represents the geometric constraints between the target spacecraft and the tracking spacecraft.
[0034] S3. Define the objective function of the game between the pursuer and the pursuer, introduce the obstacle avoidance model into the objective function, and design optimization objectives for fuel consumption and relative distance for the cluster spacecraft and the target spacecraft respectively.
[0035] The swarm spacecraft aims to get as close as possible to the non-cooperative target spacecraft while minimizing fuel consumption during the capture process. The target spacecraft, on the other hand, aims to escape the swarm's encirclement with minimal fuel consumption. Both sides are simultaneously concerned with their relative distance and relative velocity, while also considering fuel optimization. In some implementations, the target spacecraft... The spacecraft establishes the following objective function. : , in, For the relative state quantities of the spacecraft, and These are the thrust acceleration control vectors for both the pursuing and fleeing parties. It is a positive semi-definite symmetric matrix. It is a positive definite symmetric matrix. The weighting ratio factor is controlled by both sides in the pursuit and recovery process; Let be the potential function. This represents the time variable, which is used to integrate the integral term in the objective function from the initial time 0 to infinity.
[0036] Target spacecraft versus tracking spacecraft Establish the following objective function :
[0037] S4. Based on the relative dynamics model and the game objective function, construct the Hamiltonian function and the HJI (Hamilton-Jacobi-Isaacs) equation to establish a dynamic game model.
[0038] Specifically, the mathematical models of the spacecraft on both sides consist of the system's nonlinear dynamics and the performance index functions of each spacecraft. Based on this, the Hamiltonian function is:
[0039] The following HJI equation can be obtained: , in Represents a value function. This represents the gradient of the value function with respect to the state variables. This represents the optimal value function.
[0040] S5. The data-driven ADP method is used to approximate the optimal solution of the HJI equation online and solve the optimal game control strategy between the cluster spacecraft and the target spacecraft. In this process, the joint value function and the control law are both expressed as a weighted sum of basis functions. The parameters are directly optimized online using real-time state data, thereby completely decoupling the dependence on the target model.
[0041] Specifically, the optimal game strategy for tracking and targeting spacecraft can be written as follows: , , in, The weighting ratio factor is controlled by both sides in the pursuit and capture process. This represents the gradient of the optimal value function with respect to the state variables.
[0042] To obtain the optimal game strategy, the HJI equation needs to be solved. For problems where differential equations are difficult to solve, a method using given basis functions is used to approximate the optimal solution of the equation. The value function and the control law can be approximated as follows: , , , By minimizing Obtain the coefficient , , ,in , in , , These are basis functions, used to approximate the value function, the cluster spacecraft control strategy, and the target spacecraft control strategy, respectively. , , Let them be the number of their corresponding basis functions. , , These are the weight coefficients corresponding to their respective basis functions. and Indicates adjacent times and The system state variables, The state weight matrix is... For the input weight matrix, and These represent the perturbation correction terms for the cluster spacecraft and the target spacecraft, respectively. The weighting ratio factor is controlled by both sides in the pursuit and capture process. This is the Bellman error.
[0043] S6. A global escape strategy for the target spacecraft is generated based on a null-space projection method. Optimal escape control under multi-task conflicts is achieved through priority sorting and multiple projections.
[0044] To obtain the global escape strategy of a target spacecraft facing a swarm of multiple spacecraft, a null-space projection method was adopted. This method projects low-priority subtasks onto the high-priority null space, allowing the control strategy to partially implement the low-priority subtasks without affecting the high-priority tasks. This strategy aims to resolve conflicts between multiple tasks, ensuring the system can effectively cope with swarm encirclement in complex environments while maintaining efficient execution of high-priority tasks. The null-space projection method optimizes the control strategy, improves the overall system performance, and makes it more adaptable to the application requirements of multi-task scenarios.
[0045] The target spacecraft's escape strategy for each spacecraft in the cluster is as follows: .
[0046] Based on the priority of the escape strategy, the spacecraft are arranged from high to low and then from small to large relative distance, i.e., the cluster of spacecraft with the largest relative distance is selected. Assign the lowest priority Then the projection from low-priority tasks to high-priority tasks is a total projection. n -1 times can be represented as , And there are ,in This represents the initial projection control value. Indicates intermediate variables. This represents the quantity in the i-th projection process. This represents the target spacecraft control quantity corresponding to the maximum relative distance. (Through...) n -1 projection is sufficient to obtain the global escape strategy of the escaper. .
[0047] In some implementations, after obtaining the globally optimal escape strategy, the following steps are also included: S7. Verify the quality of the control strategy through performance evaluation indicators; among which, the performance evaluation indicators include tracking accuracy, fuel consumption, robustness and real-time performance; when the performance estimation indicators do not meet the requirements, return to step S5 to solve again.
[0048] As a preferred embodiment, such as Figure 1 As shown, after establishing the dynamic game model, the specific implementation process is as follows: 1. System Initialization In the initial stage of this embodiment, the state information of the spacecraft cluster is first initialized, including the position and velocity of each tracking spacecraft. and the initial state of the target spacecraft. At the same time, set the initial parameters required by the algorithm, including: (1) Approximation order: the dimension of the basis functions used for approximation functions and control strategies; (2) Initial feasible control strategy (corresponding to) Figure 2 Initial control in the system: provides initial control input to the system; (3) Exploration signal: used to cover the state space to obtain more diverse dynamic data; (4) Threshold: used as a criterion for determining the convergence of the algorithm.
[0049] 2. Data Collection and Online Learning Each spacecraft collects real-time trajectory data about non-cooperative target spacecraft via sensors. This data includes the target spacecraft's relative position and velocity information. The swarm spacecraft will share this real-time measurement data, combined with the swarm's internal communication mechanisms. In the initial phase, the control input is... ,in Indicates initial feasible control. This indicates an exploration signal; dynamic data of the system is collected through operation for a sufficiently long period of time to construct the dynamic characteristics of the system.
[0050] 3. Control strategy update (1) Set the initial number of iterations .
[0051] (2) In each iteration, a matrix is generated based on the dynamic data of the system. sum vector These quantities form the basis of value functions and control law optimization problems.
[0052] (3) Obtain the value function parameters by solving the following minimization problem. and control strategy parameters , : .
[0053] Based on the learning results of the ADP algorithm, each spacecraft updates its control strategy in real time. These updates reflect changes in the relative state between the swarm spacecraft and the target spacecraft. Each spacecraft recalculates its control inputs according to the optimized strategy and shares the updated control information through the swarm communication mechanism. Swarm communication ensures that all spacecraft can work collaboratively when executing control strategies, avoiding inconsistencies caused by independent optimization.
[0054] 4. Convergence Criteria Calculate the change between the current iteration and the previous iteration, including: , If the change is less than the set threshold, the algorithm is considered to have converged, and the iteration ends. If the convergence condition is not met, the above iterative process is repeated until the convergence condition is met.
[0055] 5. Output Results After iteration, the optimal value function and optimal control strategy are obtained: , , .
[0056] 6. Obtain the global optimal escape strategy for the target spacecraft based on the null space method. Based on the priority of the escape strategy, the spacecraft are arranged from high to low and then from small to large relative distance, i.e., the cluster of spacecraft with the largest relative distance is selected. Assign the lowest priority Then the projection from low-priority tasks to high-priority tasks is a total of [number] projections. n -1 time. Passed. n -1 projection is sufficient to obtain the global escape strategy of the target spacecraft.
[0057] 7. Performance Evaluation As a preferred implementation, after completing the cooperative tracking task, a game-theoretic performance evaluation is performed on the results. This evaluation is a crucial step in assessing the superiority of control strategies between swarm spacecraft and non-cooperative target spacecraft. Key performance indicators include target tracking accuracy, spacecraft fuel consumption, mission execution time, robustness, and system stability. The specific evaluation in this embodiment mainly focuses on the following aspects: (1) Distance convergence: assess whether the target spacecraft can be effectively captured and whether the relative distance between the cluster spacecraft can converge stably within a safe range.
[0058] (2) Fuel consumption: The control strategy needs to optimize fuel use, reduce unnecessary thrust consumption, and ensure that the mission is completed with the lowest fuel consumption.
[0059] (3) Mission completion rate: measures whether the mission is successfully completed within the specified time, including the target spacecraft tracking time, accuracy and the trajectory error of the cluster spacecraft.
[0060] (4) Robustness: The evaluation system can still maintain good performance and task completion when facing uncertainties such as external disturbances and sensor errors.
[0061] (5) Computational complexity and real-time performance: The control strategy needs to have the ability to calculate and execute quickly to ensure effective operation under real-time constraints.
[0062] If the game performance evaluation meets the requirements, the process ends; if the game performance evaluation does not meet the performance requirements, the process returns to the steps of selecting initial control and exploration signals, and data collection and online learning are carried out again until the performance requirements are met.
[0063] like Figure 2 As shown, this invention is applicable to multi-spacecraft collaborative capture and escape control missions in near-Earth orbit and deep space exploration scenarios.
[0064] This invention, based on a data-driven method that does not require a precise system model, enables real-time optimization and adjustment of target tracking strategies through the collaborative work of multiple spacecraft in dynamic environments and complex mission requirements. This effectively addresses the dynamic changes of non-cooperative targets. This adaptive control strategy not only improves the accuracy and response speed of spacecraft swarms during space missions but also enhances the system's adaptability and robustness, thereby ensuring the success rate and stability of the entire mission. Furthermore, this invention introduces an obstacle avoidance model, incorporating it into the tracker's objective function. The introduction of the obstacle avoidance model effectively prevents collisions between trackers and targets, and also prevents collisions within the tracker swarm that could lead to spacecraft failure. This adaptive, safe, and collaborative tracking method ensures that spacecraft swarms can not only complete missions efficiently and stably but also effectively avoid potential risks and safety hazards, improving mission safety and reliability.
[0065] Obviously, the above embodiments are merely examples to clearly illustrate the technical solutions of the present invention, and are not intended to limit the implementation of the present invention. Those skilled in the art can make other variations or modifications based on the above description. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included within the protection scope of the claims of the present invention.
Claims
1. An optimal control method for collaborative tracking of non-cooperative targets by swarm spacecraft, characterized in that, The method includes the following steps: S1. Construct a relative dynamic model between non-cooperative target spacecraft and swarm spacecraft; S2. Establish a cluster spacecraft obstacle avoidance model based on the relative dynamics model, and track the relative positions of spacecraft and target spacecraft, and spacecraft and the cluster interior by constraining the spherical exclusion zone and cylindrical form field; S3. Define the objective function of the game between the pursuer and the pursuer, introduce the obstacle avoidance model into the objective function, and design optimization objectives for fuel consumption and relative distance for the cluster spacecraft and the target spacecraft respectively; S4. Based on the aforementioned relative dynamics model and game objective function, construct the Hamiltonian function and HJI equation to establish a dynamic game model; S5. The data-driven ADP method is used to approximate the optimal solution of the HJI equation online and solve the optimal game control strategy between the cluster spacecraft and the target spacecraft. By combining the value function and the control law, both the value function and the control law are expressed as a weighted sum of basis functions, and the parameters are directly optimized online using real-time state data.
2. The optimal control method for cooperative tracking of non-cooperative targets by swarm spacecraft according to claim 1, characterized in that, In step S1, the relative dynamic model is: , x e For the position and velocity vector of the non-cooperative target spacecraft, ; x pi For the first in the cluster The position and velocity vector of a spacecraft ; For non-cooperative target spacecraft and the first in the cluster Relative state quantities of a spacecraft ; , These are the instantaneous orbital angular velocity and angular acceleration of the reference spacecraft, respectively. The gravitational constant of Earth; , and These are the reference orbit radius, the target orbit radius, and the first orbital radius in the cluster. The orbital radius of a spacecraft; , , For the three-axis thrust acceleration of the target spacecraft, , , For the first in the cluster The three-axis thrust acceleration of a spacecraft; make and The state-space representation is as follows , , in and Nonlinear functions representing the intrinsic dynamics of the swarm spacecraft and target spacecraft system. This represents the control input matrix for the spacecraft cluster. This represents the target spacecraft control input matrix; The equations of relative motion are obtained , The extended equation is defined as follows: , in, , , , .
3. The optimal control method for cooperative tracking of non-cooperative targets by swarm spacecraft according to claim 1, characterized in that, In step S2, the obstacle avoidance model is implemented using a potential function, establishing a spherical no-go zone for the tracking spacecraft and a cylindrical situational field for the target spacecraft; the potential function for: , in, This represents the relative position components between the tracking spacecraft and the target spacecraft. This represents a metric parameter indicating the relative distance between tracking spacecraft. This represents the geometric constraints between the target spacecraft and the tracking spacecraft.
4. The optimal control method for cooperative tracking of non-cooperative targets by swarm spacecraft according to claim 1, characterized in that, In step S3 For the first in the cluster The following objective function is established for each spacecraft. ; , in, For the relative state quantities of the spacecraft, and These are the thrust acceleration control vectors for both the pursuing and fleeing parties. It is a positive semi-definite symmetric matrix. It is a positive definite symmetric matrix. The weighting ratio factor is controlled by both sides in the pursuit and capture process; Let be the potential function. Represents a time variable; Target spacecraft versus tracking spacecraft Establish the following objective function : 。 5. The optimal control method for cooperative tracking of non-cooperative targets by swarm spacecraft according to claim 2, characterized in that, In step S4, the Hamiltonian function is: , The HJI equation is: , in, For the relative state quantities of the spacecraft, and These are the thrust acceleration control vectors for both the pursuing and fleeing parties. Represents a value function. Let be the potential function. It is a positive semi-definite symmetric matrix. It is a positive definite symmetric matrix. This represents the gradient of the value function with respect to the state variables. This represents the optimal value function.
6. The optimal control method for cooperative tracking of non-cooperative targets by swarm spacecraft according to claim 2, characterized in that, In step S5, the optimal game control strategy between the cluster spacecraft and the target spacecraft is: , , in, The weighting ratio factor is controlled by both sides in the pursuit and capture process. This represents the gradient of the optimal value function with respect to the state variables; When using the ADP method to approximate the optimal solution of the HJI equation online, a method with given basis functions is adopted as a means of function approximation. The value function and control law are approximated as follows: , , , By minimizing Obtain the coefficient , , ,in , in , , These are basis functions, used to approximate the value function, the cluster spacecraft control strategy, and the target spacecraft control strategy, respectively. , , Let the number of their corresponding basis functions be . , , These are the weight coefficients corresponding to their respective basis functions. and Indicates adjacent times and The system state variables, The state weight matrix is... For the input weight matrix, and These represent the perturbation correction terms for the cluster spacecraft and the target spacecraft, respectively. The weighting ratio factor is controlled by both sides in the pursuit and capture process. This is the Bellman error.
7. The optimal control method for cooperative tracking of non-cooperative targets by swarm spacecraft according to claim 1, characterized in that, In step S5, the specific process of using the data-driven ADP method to approximate the optimal solution of the HJI equation online and solve the optimal game control strategy between the cluster spacecraft and the target spacecraft includes: System initialization: Set the initial state information of the spacecraft cluster; the initial state information includes the position and velocity of the tracking spacecraft, the initial state of the target spacecraft, and algorithm parameters; the algorithm parameters include the approximate order of the basis functions, the initial feasible control strategy, the exploration signal, and the convergence threshold; Data collection and online learning: Real-time acquisition of trajectory data of target spacecraft, combined with cluster communication mechanisms to share dynamic data; Control strategy update: Based on dynamic data, iteratively optimize the value function parameters and control strategy parameters to generate coordinated control input; Convergence criterion: The convergence of the algorithm is determined by comparing the change in the iteration results with a threshold. Output the optimal value function and control strategy.
8. The optimal control method for cooperative tracking of non-cooperative targets by swarm spacecraft according to claim 1, characterized in that, The method further includes the following steps: S6. A global escape strategy for the target spacecraft is generated based on a null-space projection method. Optimal escape control under multi-task conflicts is achieved through priority sorting and multiple projections.
9. The optimal control method for cooperative tracking of non-cooperative targets by swarm spacecraft according to claim 8, characterized in that, Step S6 specifically includes: Prioritize the target spacecraft according to its relative distance to the cluster of spacecraft, from smallest to largest. The global escape policy is generated through n-1 projections, and is represented as follows: , And there are ,in This represents the initial projection control value. Indicates intermediate variables. This represents the quantity in the i-th projection process. This represents the target spacecraft control quantity corresponding to the maximum relative distance.
10. The optimal control method for cooperative tracking of non-cooperative targets by swarm spacecraft according to any one of claims 1 to 9, characterized in that, The method further includes the following steps: S7. Verify the quality of the control strategy through performance evaluation indicators, including tracking accuracy, fuel consumption, robustness, and real-time performance; if the performance evaluation indicators do not meet the requirements, return to step S5.
Citation Information
Patent Citations
Spacecraft pursuit game control solving method based on adaptive dynamic programming
CN117034745A
Differential game-based aerial bomb distributed collaborative guidance method
CN118550321A
Deep reinforcement learning method for controlling orbital trajectories of spacecrafts in multi-spacecraft swarm
US20220363415A1