Satellite pursuit and capture decision-making method based on continuous-time hierarchical reinforcement learning
Patent Information
- Application Number
- CN202611299038.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-08-26
- Publication Date
- 2026-09-25
AI Technical Summary
[0005]为解决现有技术中多卫星追捕对抗方法难以同时刻画连续时间轨道动力学过程、未能有效考虑控制执行延迟以及在复杂对抗任务中缺乏分层决策机制的技术问题,本发明提供了一种基于连续时间分层强化学习的卫星追捕对抗决策方法,用于提高卫星在复杂对抗环境下的自主决策能力和策略稳定性
1、本发明公开的基于连续时间分层强化学习的卫星追捕对抗决策方法,通过在连续时间轨道动力学模型中引入包含动作调度机制的马尔可夫决策过程,并将单次决策周期划分为无控制输入的等待阶段和施加控制加速度的执行阶段,实现了对航天器真实“决策-指令传输-执行”延迟过程的数学建模;同时,本发明通过构建由管理层生成局部最优子目标、由控制层输出联合动作的连续时间分层强化学习框架,解决了长时序太空对抗任务中单一策略网络难以兼顾全局路径规划与高频底层姿态控制的工程问题。
Smart Images

Figure CN122808988A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of on-orbit intelligent decision-making and control technology for spacecraft, specifically a satellite pursuit and adversarial decision-making method based on continuous-time hierarchical reinforcement learning, as well as a computer terminal and computer-readable storage medium that apply this method. Background Technology
[0002] With the development of space technology, multi-satellite on-orbit missions are gradually evolving towards adversarial and cooperative approaches. Tasks such as satellite pursuit, avoidance, and protection have become crucial issues in space security and on-orbit servicing. How to achieve autonomous decision-making for satellites in complex and dynamic environments has become a key area of current research.
[0003] In existing technologies, differential game theory, optimal control, or optimization algorithms are often used to model and solve pursuit-adversarial problems. However, these methods typically rely on accurate models, have high computational complexity, and are difficult to apply to high-dimensional adversarial scenarios involving multiple agents. Meanwhile, reinforcement learning methods have been introduced into this field in recent years, but most are based on discrete-time models, usually assuming that control commands can be executed instantly. This makes it difficult to reflect the continuous-time characteristics of orbital dynamics and the control execution delays present in real-world systems.
[0004] Furthermore, in long-term complex tasks, traditional single-layer decision-making methods struggle to balance global planning and local control. Existing hierarchical reinforcement learning methods are mostly based on discrete-time frameworks and have not yet effectively combined continuous-time dynamics with multi-agent adversarial characteristics. Therefore, current technologies still lack a multi-satellite pursuit adversarial scheme that can simultaneously describe continuous-time evolution, execution latency, and hierarchical decision-making mechanisms. Summary of the Invention
[0005] To address the technical problems of existing multi-satellite pursuit and adversarial methods, such as difficulty in simultaneously characterizing continuous-time orbital dynamics, failure to effectively consider control execution delays, and lack of hierarchical decision-making mechanisms in complex adversarial missions, this invention provides a satellite pursuit and adversarial decision-making method based on continuous-time hierarchical reinforcement learning, which can improve the autonomous decision-making ability and strategy stability of satellites in complex adversarial environments.
[0006] To achieve the above objectives, the present invention provides the following technical solution: This invention discloses a satellite pursuit adversarial decision-making method based on continuous-time hierarchical reinforcement learning, comprising the following steps: S1. Construct a multi-satellite pursuit and adversarial environment, establish a continuous-time relative motion model of the pursuing satellite, target satellite and guardian satellite in the local orbital coordinate system, define the state space and set the termination conditions of the adversarial process; S2. Construct a Markov decision process model with an action scheduling mechanism. In this model, the control input for pursuing satellites is defined as a joint action including control acceleration and planned execution time. During the state evolution process, a single decision cycle is divided into a first evolution stage without control input and a second evolution stage with control acceleration applied, based on the planned execution time. The relative motion state of each satellite is continuously updated based on the continuous time relative motion model. S3. Construct a continuous-time hierarchical reinforcement learning framework, which includes a management layer and a control layer; the management layer generates a set of candidate sub-targets based on the current relative motion state of each satellite and evaluates and selects the optimal sub-target; the control layer receives the current relative motion state of each satellite and the optimal sub-target, and outputs the joint action for the current decision cycle; S4. Configure evasion strategies for target satellites in the multi-satellite pursuit and adversarial environment, and configure pre-trained interception strategies for guardian satellites; S5. In the multi-satellite pursuit adversarial environment, the Markov decision process model is used to iteratively train the continuous-time hierarchical reinforcement learning framework in combination with the preset reward function to obtain a converged satellite pursuit decision strategy.
[0007] As a further improvement to the above scheme, in step S1, a virtual reference satellite with the same orbital parameters as the tracking satellite is set as the origin of the local orbital coordinate system; the continuous-time relative motion model is described by the Clohessy-Wiltshire equations, expressed as: ; In the formula, For the first The relative positions of the satellites; For the first The relative speed of the satellites; For the first The relative acceleration of the satellites; For the first Control acceleration of a satellite; The orbital angular velocity of the virtual reference satellite.
[0008] As a further improvement to the above scheme, in step S1, the state space includes the system state, which includes the motion state of the pursuing satellite itself, the relative motion state between the target satellite and the pursuing satellite, and the relative motion state between the guardian satellite and the pursuing satellite, expressed by the following formula: ; In the formula, express The system state at any given moment; and They represent Constantly track the satellite's position and velocity; and They represent The relative position and relative velocity of the target satellite and the pursuing satellite at any given time; and These represent the relative positions and relative velocities of the guardian satellite and the pursuing satellite, respectively. In step S2, the continuous updating of the relative motion states of each satellite is represented by a piecewise state transition equation: ; ; In the formula, express The first derivative; Indicates the current decision-making moment. Indicates the next decision point; This represents the control acceleration generated at the current decision-making moment; Indicates the planned execution time; This represents the state transition function.
[0009] As a further improvement to the above scheme, in step S3, the specific process of the management layer generating the candidate sub-target set includes: Obtain the current distance between the pursuing satellite and the target satellite, and multiply the current distance by a preset scaling factor to obtain the search radius; The sub-target search area is constructed with the current position of the pursuing satellite as the center and the search radius as the specified area. Within the sub-target search area, a set of candidate sub-targets containing multiple spatial location points is generated; The optimal sub-objective is selected by evaluating it using the following formula: ; In the formula, The set of candidate sub-targets; Current system state Under the condition of selecting candidate sub-targets The expected cumulative reward that can be obtained; This is the optimal sub-objective.
[0010] As a further improvement to the above scheme, in step S3, the control layer outputs the joint action of the current decision cycle using the following formula: ; In the formula, For continuous action space; The output is the combined action of the current decision cycle; Given the current system state and optimal sub-objective Under the conditions, to carry out joint operations The expected cumulative reward that can be obtained.
[0011] As a further improvement to the above scheme, in step S4, the evasion strategy is as follows: when the distance between the pursuing satellite and the target satellite is greater than the preset threat distance, the target satellite performs a straight-line escape maneuver in the opposite direction of the relative position vector; when the distance is less than the preset threat distance, the target satellite performs a lateral evasion maneuver with periodic direction switching on the basis of the reverse escape.
[0012] As a further improvement to the above scheme, in step S4, the interception strategy is a continuous control strategy trained by a continuous-time reinforcement learning algorithm based on the Actor-Critic architecture. The satellite protection reward function used in the pre-training process consists of target protection reward, interception reward, occupancy reward and terminal reward. The target protection reward is used to constrain the guardian satellite to accompany the target satellite, and its value is negatively correlated with the deviation of the current distance between the guardian satellite and the target satellite from the expected protection distance. The interception reward is adjusted by a dynamic threat factor, the value of which is negatively correlated with the distance between the pursuing satellite and the target satellite. When the pursuing satellite enters the set threat perception range, the smaller the distance between the guardian satellite and the pursuing satellite, the greater the interception reward. Regarding the occupancy reward, a positive reward is given when the guardian satellite is located between the area connecting the pursuing satellite and the target satellite, and the position error is less than a preset threshold. The terminal rewards include positive rewards given when the guardian satellite successfully intercepts the pursuing satellite, and negative penalties given when the pursuing satellite successfully captures the target satellite, or when the guardian satellite collides with or engages with the target satellite and the confrontation exceeds the time limit.
[0013] As a further improvement to the above scheme, in step S5, the preset reward function includes management layer reward and control layer reward; the management layer reward consists of sub-target quality reward, target proximity reward and sub-target arrival reward; Specifically, for the sub-target quality reward, a positive reward is given when the generated sub-target is closer to the target satellite than the current position of the pursuing satellite; The value of the target proximity reward is positively correlated with the amount of reduction in distance between the pursuing satellite and the target satellite after the current sub-target is executed; For the sub-target arrival reward, a positive reward is given when the distance between the tracking satellite and the sub-target is less than a preset arrival threshold; The control layer reward consists of target approach reward, sub-target guidance reward, direction consistency reward, safety penalty, energy penalty, time penalty and terminal reward; The value of the target approach reward is positively correlated with the amount of distance reduction between the pursuing satellite and the target satellite at the current moment; The value of the sub-target guidance reward is positively correlated with the amount of distance reduction between the current tracking satellite and the current sub-target at the current moment; The directional consistency reward is determined by calculating the inner product of the direction vector of the pursuing satellite pointing to the target satellite and the control acceleration vector, and a positive reward is given when the control acceleration direction is consistent with the target direction. The safety penalty is triggered when the distance between the pursuing satellite and the guardian satellite is less than a safe distance threshold, in order to constrain the pursuing satellite to avoid the guardian satellite; The value of the energy penalty is positively correlated with the amplitude of the control acceleration in order to suppress excessive control input; The value of the time penalty is positively correlated with the size of the planned execution time, so as to constrain the control layer to output actions that approach the minimum execution time threshold; The terminal rewards include positive rewards for successfully capturing the target satellite and negative penalties for being intercepted by a protected satellite or for exceeding the timeout period.
[0014] The present invention also discloses a computer terminal, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the steps of the satellite pursuit adversarial decision-making method based on continuous-time hierarchical reinforcement learning as described above.
[0015] The present invention also discloses a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the satellite pursuit adversarial decision-making method based on continuous-time hierarchical reinforcement learning as described above.
[0016] Compared with the prior art, the beneficial effects of the present invention are: 1. The satellite pursuit adversarial decision-making method based on continuous-time hierarchical reinforcement learning disclosed in this invention introduces a Markov decision process with an action scheduling mechanism into the continuous-time orbital dynamics model, and divides a single decision cycle into a waiting phase without control input and an execution phase with applied control acceleration, thereby realizing mathematical modeling of the real "decision-command transmission-execution" delay process of spacecraft. At the same time, this invention solves the engineering problem that a single policy network in long-term space adversarial missions is difficult to balance global path planning and high-frequency low-level attitude control by constructing a continuous-time hierarchical reinforcement learning framework in which the management layer generates locally optimal sub-targets and the control layer outputs joint actions.
[0017] 2. This method describes the continuous-time relative motion in the local orbital coordinate system using the Clohessy-Wiltshire equations and updates the relative motion state of each satellite based on the piecewise state transition equations. This achieves a direct mapping between the algorithm's state space and the real orbital dynamics environment, enabling the output continuous control commands to be directly applied to the satellite's real orbital maneuvers in physical time.
[0018] 3. This method dynamically generates the search area by obtaining the target distance at the management level and evaluates the expected cumulative reward to select the optimal sub-target, thereby automatically decomposing long-distance cross-orbit pursuit tasks into multiple sequentially reachable spatial intermediate waypoints. At the same time, it combines the joint action strategy, which includes continuous acceleration and planned execution time, directly output by the control layer, thus solving the problems of ineffective fuel consumption and maneuver overshoot caused by fixed control step size in traditional discrete reinforcement learning.
[0019] 4. This method configures the guardian satellite with a combined reward function consisting of target protection, dynamic interception, occupancy blocking and terminal settlement, and introduces a dynamic threat factor adjustment mechanism related to the distance to the pursued satellite. This enables the guardian satellite to autonomously switch its space defense attitude and interception trajectory in different game stages such as escort protection, threat perception and active blocking.
[0020] 5. This method introduces sub-target quality and proximity rewards at the management level, and configures directional consistency rewards, collision avoidance safety penalties, and energy and time penalties for constrained acceleration and execution time at the control level. This enables the pursuit satellite to autonomously avoid the collision blind zone of the guardian satellite in a complex and disturbed environment, and to approach the final target with lower thrust consumption. Attached Figure Description
[0021] Figure 1 This is a flowchart of the satellite pursuit adversarial decision-making method based on continuous-time hierarchical reinforcement learning in Embodiment 1 of the present invention.
[0022] Figure 2 This is a schematic diagram of the motion of multiple satellites in the local orbital coordinate system in Embodiment 1 of the present invention.
[0023] Figure 3 This is a schematic diagram of the Markov decision process for action scheduling with execution delay in Embodiment 1 of the present invention.
[0024] Figure 4 This is a diagram of the continuous-time hierarchical reinforcement learning framework architecture in Embodiment 1 of the present invention.
[0025] Figure 5 This is a schematic diagram of the hierarchical strategy execution guidance process in Embodiment 1 of the present invention.
[0026] Figure 6This is a diagram showing the interception trajectory of the guardian satellite in Embodiment 1 of the present invention under different initial distributions.
[0027] Figure 7 This is a graph showing the reward curve and interception success rate curve during the training process of the guardian satellite in Embodiment 1 of the present invention; Figure 7 In the diagram, (a) is the reward curve and (b) is the loss curve.
[0028] Figure 8 This is a three-dimensional trajectory diagram of the pursuit satellite chasing and escaping under the interception of the guardian satellite in Embodiment 1 of the present invention.
[0029] Figure 9 This is a graph showing the training reward and capture success rate of the satellite in an adversarial environment in Embodiment 1 of the present invention. Figure 9 In the diagram, (a) represents the training reward, and (b) represents the capture success rate.
[0030] Figure 10 This is a comparison chart of the training normalized rewards of the CTHRL algorithm and the DDPG algorithm in Embodiment 1 of the present invention.
[0031] Figure 11 This is a comparison chart of the capture success rates of the CTHRL algorithm and the DDPG algorithm in Embodiment 1 of the present invention.
[0032] Figure 12 This is a comparison of the final pursuit-target relative distance during the training of the CTHRL algorithm and the DDPG algorithm in Embodiment 1 of the present invention.
[0033] Figure 13 This is a schematic diagram of the computer terminal structure in Embodiment 2 of the present invention. Detailed Implementation
[0034] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0035] Example 1 Please see Figure 1 This embodiment provides a satellite pursuit adversarial decision-making method based on continuous time hierarchical reinforcement learning, including the following steps, namely S1~S5.
[0036] S1. Construct a multi-satellite pursuit and adversarial environment. Establish a continuous-time relative motion model of the pursuing satellite, target satellite, and guardian satellite in a local orbital coordinate system, define the state space, and set the termination conditions for the adversarial process.
[0037] In step S1, the local orbital coordinate system is often used to describe the relative motion of spacecraft. It can simplify complex nonlinear absolute orbital motion into approximately linear relative motion equations, which facilitates the formulation of on-orbit autonomous control strategies.
[0038] Specifically, such as Figure 2 As shown, a virtual reference satellite with the same reference orbit parameters as the pursuing satellite is introduced as the origin of the local orbit coordinate system. The X-axis of the coordinate system points outward radially along the virtual satellite, the Y-axis points tangentially along the orbit, and the Z-axis is determined by the right-hand rule (i.e., the normal direction) of X and Y. Figure 2 middle, This forms the geocentric inertial coordinate system. This represents the absolute position vector of a specific satellite (such as a tracking, target, or guardian satellite) in the geocentric inertial coordinate system, i.e., from the geocentric origin. A vector that directly points to the actual location of the satellite. This represents the relative position vector of a specific satellite in the local orbital coordinate system. It originates from the centroid of the virtual reference satellite (the origin of the local orbital coordinate system). The vector pointing to the actual satellite position is used in subsequent relative motion models. This corresponds to the specific relative position coordinates of each satellite. ).
[0039] The continuous-time relative motion model is described by the Clohessy-Wiltshire equations, expressed as: ; In the formula, For the first The relative positions of the satellites; For the first The relative speed of the satellites; For the first The relative acceleration of the satellites; For the first Control acceleration of a satellite; The orbital angular velocity of the virtual reference satellite.
[0040] To ensure the equilibrium of the multi-satellite game, the pursuing satellite needs to have a maneuverability advantage, i.e., a maximum acceleration of the target satellite is set. Protecting the maximum acceleration of satellites .
[0041] The state space includes the system state, which includes the motion state of the pursuing satellite itself, the relative motion state between the target satellite and the pursuing satellite, and the relative motion state between the guardian satellite and the pursuing satellite, expressed by the following formula: ; In the formula, express The system state at any given moment; and They represent Constantly track the satellite's position and velocity; and They represent The relative position and relative velocity of the target satellite and the pursuing satellite at any given time; and These represent the relative positions and relative velocities of the guardian satellite and the pursuing satellite, respectively.
[0042] The termination condition of the confrontation process includes any of the following: successful pursuit (i.e., the distance between the pursuing satellite and the target satellite is...) The interception of the guardian satellite was successful (i.e., the distance between the pursuing satellite and the guardian satellite was...). ), and task timeout (current time) ).
[0043] S2. Construct a Markov decision process model with an action scheduling mechanism. In this model, the control input for pursuing satellites is defined as a joint action including control acceleration and planned execution time. During the state evolution process, a single decision cycle is divided into a first evolution stage without control input and a second evolution stage with the control acceleration applied, based on the planned execution time. The relative motion state of each satellite is continuously updated based on the continuous time relative motion model.
[0044] Markov Decision Processes (MDPs) are a classic mathematical framework in reinforcement learning used to simulate the interaction between an agent and its environment, typically assuming that action commands take effect instantaneously. However, in real-world on-orbit missions, due to policy computation, command transmission, and actuator response, control systems face inherent "decision-wait-actuated" delays. To address this, as... Figure 3 As shown, this embodiment constructs an action scheduling ordinary differential equation Markov decision process model to characterize the actual execution delay.
[0045] In this model, joint actions The continuous updating of the relative motion states of each satellite is represented by a piecewise state transition equation: ; ; In the formula, express The first derivative; Indicates the current decision-making moment; This represents the control acceleration generated at the current decision-making moment; Indicates the planned execution time; Indicates the next decision point; This represents the state transition function.
[0046] S3. Construct a continuous-time hierarchical reinforcement learning framework, which includes a management layer and a control layer; the management layer generates a set of candidate sub-targets based on the current relative motion state of each satellite and evaluates and selects the optimal sub-target; the control layer receives the current relative motion state of each satellite and the optimal sub-target, and outputs the joint action for the current decision cycle.
[0047] For long-range spatial pursuit and other long-term decision-making tasks, a single-layer architecture struggles to balance global planning and local control. Therefore, this embodiment employs a continuous-time hierarchical reinforcement learning architecture. Figure 4 As shown, it includes two control loops: a management layer and a control layer. (As illustrated...) Figure 5 As shown, the sub-targets generated by the management layer do not correspond to the final capture point, but rather serve as intermediate guiding nodes to guide the pursuing satellites to gradually approach the target.
[0048] The specific process by which the management layer generates the set of candidate sub-targets includes: Obtain the current distance between the pursuing satellite and the target satellite, and compare the current distance with a preset scaling factor. ( Multiply by , and you get the search radius; The sub-target search area is constructed with the current position of the pursuing satellite as the center and the search radius as the specified area. Within the sub-target search area, a set of candidate sub-targets containing multiple spatial location points is generated; The optimal sub-objective is selected by evaluating it using the following formula: ; In the formula, The set of candidate sub-targets; Current system state Under the condition of selecting candidate sub-targets The expected cumulative reward that can be obtained; This is the optimal sub-objective.
[0049] The control layer outputs the joint action for the current decision cycle using the following formula: ; In the formula, For continuous action space; The output is the combined action of the current decision cycle; Given the current system state and optimal sub-objective Under the conditions, to carry out joint operations The expected cumulative reward that can be obtained.
[0050] S4. Configure evasion strategies for target satellites in the multi-satellite pursuit and adversarial environment, and configure pre-trained interception strategies for guardian satellites.
[0051] The evasion strategy is as follows: the target satellite adopts a hierarchical escape rule based on relative distance. When the distance between the pursuing satellite and the target satellite is greater than the preset threat distance, the target satellite performs a straight-line escape maneuver in the opposite direction of the relative position vector. When the distance is less than the preset threat distance, the target satellite performs a lateral (zig-zag) evasion maneuver with periodic direction switching on the basis of the reverse escape.
[0052] The interception strategy is a continuous control strategy trained using a continuous-time reinforcement learning algorithm (CTDDPG, Continuous-Time Deep Deterministic Policy Gradient Algorithm) based on an Actor-Critic architecture. The guardian satellite reward function used in the pre-training process consists of target protection reward, interception reward, occupancy reward, and terminal reward. The total reward for the guardian satellite used in the pre-training process... Represented as: ; The definitions of each sub-reward are as follows: Target protection reward : ; Interception Rewards : ; Placement Rewards : ; Terminal rewards The value is constant when the satellite is successfully intercepted and pursued. The target satellite is a constant when it is captured. The constant value when a collision occurs. The timeout is a constant. .
[0053] In the above formula, to These are the weighting coefficients; to The amplitude truncation threshold; and These represent the distances between the guardian satellite and the target satellite at the previous and current moments, respectively. For the desired protection distance; To perceive the distance of threats; To track the distance between the satellite and the target satellite; and These represent the distances between the tracking satellite and the guardian satellite at the previous and current moments, respectively. To ensure effective interception distance; This refers to the position reward coefficient. As a threat factor, ; This is the position error threshold; This is the amplitude truncation function. This is the function for finding the maximum value.
[0054] S5. In the multi-satellite pursuit adversarial environment, the Markov decision process model is used to iteratively train the continuous-time hierarchical reinforcement learning framework in combination with the preset reward function to obtain a converged satellite pursuit decision strategy.
[0055] In step S5, the preset reward function includes management layer rewards and control layer rewards; wherein, management layer rewards... Represented as: ; The definitions of each sub-reward are as follows: Sub-target quality reward : ; Target proximity reward : ; Sub-goal achievement reward: When hour, ,otherwise .
[0056] In the above formula, These are the weighting coefficients; and These represent the distances between the tracking satellite and the target satellite at the current and next moments, respectively. The distance between the sub-target and the target satellite; Normalization factor; and These are the position vectors of the tracking satellite and the sub-target, respectively; Represents the L2 norm; The threshold for sub-targets to reach; To reach the reward constant.
[0057] Control layer rewards Represented as: ; The definitions of each sub-item of reward and punishment are as follows: Target Approach Reward : ; Directional Consistency Reward : ; Safety penalties :when hour, ,otherwise .
[0058] Sub-goal guided reward : ; Energy Punishment : ; Time penalty : ; Terminal rewards : A constant when the target satellite is successfully captured. When intercepted by the protected satellite, it is a constant. The timeout is constant. .
[0059] In the above formula, to These are the weighting coefficients; and This is the truncation threshold; Let be the position vector of the target satellite. To track the satellite's position vector; To track the satellite's control acceleration vector, Maximum acceleration constraints for tracking satellites; To determine the distance between the satellite being tracked and the satellite being protected, This is the safe distance threshold; and These represent the distances between the tracking satellite and the current sub-target at the current moment and the next moment, respectively. For the planned execution time, and These are the minimum and maximum allowed execution times, respectively.
[0060] To verify the effectiveness of the method proposed in this invention, this embodiment also provides a simulation experiment.
[0061] The simulation experiment in this embodiment is divided into two parts. The first part tests the protection capability of the guardian satellite. The second part introduces the guardian satellite trained in the first part as a jamming adversary to evaluate the pursuit satellite's pursuit capability when facing interception. All simulation experiments are implemented in a Python 3.12 environment, and model building and training use the PyTorch 2.4.1 framework. The hardware platform is equipped with an Intel i5 CPU and an NVIDIA GeForce GTX 1650 GPU, and is accelerated by CUDA 11.8 to improve training efficiency.
[0062] A. Satellite Simulation Experiment This experiment aims to train the CTDDPG agent protecting the satellite, enabling it to accurately intercept the pursuing satellite and protect the target satellite in a chase scenario. A rule-based policy agent is used as the training adversary. The target satellite executes a hierarchical escape strategy based on its distance from the pursuing satellite: lateral (zig-zag) maneuvers at close range and straight-line movement away at medium to long range. The pursuing satellite employs a straight-line pursuit strategy. This setup provides the protecting satellite with a stable training environment exhibiting clear behavioral patterns, thereby validating the interception and control performance of the CTDDPG algorithm in a continuous action space.
[0063] The target satellites are initially randomly distributed near the origin. The pursuing satellites maintain an initial relative distance of 300m to 500m from the target and are randomly distributed around the target. Guardian satellites are deployed along the line connecting the pursuing and target satellites, with a position weight between 0.3 and 0.7 for this line. Three sets of satellite trajectories from the later stages of training are selected for analysis, such as... Figure 6 As shown in the diagram. The results indicate that the guardian satellite (green curve) can rapidly adjust its acceleration direction based on the initial distribution and real-time status, actively entering the critical defense zone between the pursuing satellite (red curve) and the target satellite (blue curve), forming an interception barrier. It maintains a defensive posture that separates the pursuing satellite from the target satellite. When the pursuing satellite approaches the target in a straight line, the guardian satellite senses the relative motion in real time and gradually reduces its distance from the pursuing satellite by continuously fine-tuning its acceleration. In all three trajectories, the final distance between the guardian satellite and the pursuing satellite remains below 100m, while the distance to the target satellite remains strictly above 50m, thus achieving continuous suppression of the pursuing satellite while avoiding collision with the target.
[0064] Figure 7 The training curves of the guardian satellite are presented, including reward and loss curves for a total of 2000 training rounds. To verify the guardian satellite's ability to protect targets in different environments, comparative experiments were conducted. The interception success rate of the guardian satellite in two different target satellite maneuvering modes was recorded and plotted. Figure 7The success rate curves in the data are shown. The results indicate that the agent has effectively internalized the reward mechanisms for approaching and pursuing satellites, protecting target satellites, and avoiding collisions, thereby continuously learning and optimizing behavioral strategies that align with mission objectives.
[0065] B. Satellite tracking simulation experiment In this phase of the experiment, the guardian satellite trained in the first part was introduced as an environmental interference factor. The satellite pursuit not only needs to capture the target satellite but also must avoid being intercepted by the guardian satellite. The main scene parameters and algorithm parameters used in this experiment are summarized in Tables 1 and 2, respectively. The experimental results are as follows: Figure 8 As shown. For ease of demonstration, the generated sub-targets are labeled as... (in ).
[0066] Table 1: Simulation Scene Parameters
[0067] Table 2: Algorithm Parameters
[0068] Training curves for satellite tracking are as follows Figure 9 As shown, the curves include the reward curve and the capture success rate curve. For easier observation, the reward results were smoothed using a 50-round moving average. Experimental results show that even in the presence of interception interference from the guardian satellite, the pursuing satellite can still avoid the guardian satellite under the guidance of the sub-target, while gradually approaching and successfully capturing the target satellite, ultimately maintaining a stable capture success rate of over 95%.
[0069] To further evaluate the effectiveness of the proposed algorithm, a comparative experiment was conducted between the Continuous-Time Hierarchical Reinforcement Learning (CTHRL) algorithm and the Deep Deterministic Policy Gradient (DDPG) algorithm under the same training conditions. Normalized reward, capture success rate, and final pursuit satellite-target satellite distance were used as performance metrics. Here, capture success rate is defined as the percentage of successful captures in 20 independent evaluation rounds out of every 50 training rounds, and final pursuit satellite-target satellite distance is defined as the average Euclidean distance between the pursuit satellite and the target satellite at the end of the same evaluation round. The comparison results are presented below. Figures 10 to 12 middle.
[0070] Figure 10The normalized reward curves of the two algorithms during training were compared. Both algorithms continuously improve their rewards through interaction with the environment, indicating that they gradually learn effective pursuit strategies. In the early stages of training, DDPG achieved a slightly higher normalized reward than CTHRL. This is mainly because DDPG uses a single-layer architecture to directly learn continuous control policies, thus enabling faster initial policy improvement. In contrast, CTHRL requires simultaneous optimization of the management and control layer policies, necessitating additional exploration before generating effective sub-objectives. As training progresses, the management layer gradually learns to effectively guide high-quality sub-objectives to the control layer, leading to a rapid increase in normalized reward. Ultimately, CTHRL converges to a higher reward level after approximately 250 training epochs, demonstrating superior long-term policy performance.
[0071] The capture success rates of the two algorithms are as follows: Figure 11 As shown, similar to the reward curve, DDPG exhibits a slightly higher success rate in the initial training phase. This is attributed to its relatively simple policy structure, which facilitates faster learning of basic pursuit behaviors. With gradual optimization of the hierarchical policy, CTHRL's capture performance rapidly improves, maintaining a stable capture success rate above 95%. The results indicate that although the hierarchical framework introduces a longer exploration phase in the early stages of training, it significantly accelerates the policy convergence process once meaningful sub-objectives are established.
[0072] Figure 12 The final pursuit-target satellite distance during training is shown. This final distance is calculated as the average Euclidean distance between the pursuing and target satellites at the end of 20 independent evaluation rounds. It can be observed that in the early stages of training, DDPG reduces the final distance more rapidly, while CTHRL, due to the need to simultaneously optimize the management and control layer policies, decreases the distance relatively more slowly. However, after sufficient training, CTHRL quickly converges to a stable low-distance region, while DDPG requires significantly more training rounds to reach a comparable final distance metric. These results strongly demonstrate that the proposed continuous-time hierarchical reinforcement learning framework achieves a significantly faster convergence speed while maintaining superior pursuit performance.
[0073] Example 2 This embodiment provides a computer terminal, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the steps of the satellite pursuit adversarial decision-making method based on continuous-time hierarchical reinforcement learning as described in Embodiment 1.
[0074] like Figure 13As shown, the computer terminal provided in this embodiment includes: at least one processor 101, and a memory 102 connected to at least one processor 101. This embodiment does not limit the specific connection medium between the processor 101 and the memory 102. Figure 13 The example shown is the connection between processor 101 and memory 102 via bus 100. Bus 100 is... Figure 13 The connections between other components are shown in bold lines and are for illustrative purposes only, not as limiting information. Bus 100 can be divided into address bus, data bus, control bus, etc., for ease of representation. Figure 13 The bus is represented by a single thick line, but this does not indicate that there is only one bus or one type of bus. Alternatively, the processor 101 may also be called a controller; there is no restriction on the name.
[0075] In this embodiment, the memory 102 stores instructions that can be executed by at least one processor 101. The at least one processor 101 can execute the aforementioned method by executing the instructions stored in the memory 102.
[0076] The processor 101 is the control center of the device. It can connect to various parts of the control device through various interfaces and lines. By running or executing instructions stored in memory 102 and calling data stored in memory 102, the processor can perform various functions and process data, thereby monitoring the device as a whole.
[0077] In one possible design, processor 101 may include one or more processing units. Processor 101 may integrate an application processor and a modem processor, wherein the application processor mainly handles the operating system, user interface, and applications, and the modem processor mainly handles wireless communication. It is understood that the modem processor may also not be integrated into processor 101. In some embodiments, processor 101 and memory 102 may be implemented on the same chip; in some embodiments, they may also be implemented on separate chips.
[0078] Processor 101 can be a general-purpose processor, such as a central processing unit (CPU), digital signal processor, application-specific integrated circuit, field-programmable gate array or other programmable logic device, discrete gate or transistor logic device, or discrete hardware component, capable of implementing or executing the methods, steps, and logic block diagrams disclosed in the embodiments. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the satellite pursuit adversarial decision-making method based on continuous-time hierarchical reinforcement learning disclosed in Embodiment 1 can be directly implemented by the hardware processor, or implemented by a combination of hardware and software modules in processor 101.
[0079] Memory 102, as a non-volatile computer-readable storage medium, can be used to store non-volatile software programs, non-volatile computer-executable programs, and modules. Memory 102 may include at least one type of storage medium, such as flash memory, hard disk, multimedia card, card-type memory, random access memory (RAM), static random access memory (SRAM), programmable read-only memory (PROM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), magnetic storage, magnetic disk, optical disk, etc. Memory 102 can be any other medium capable of carrying or storing desired program code in the form of instructions or data structures that can be accessed by a computer, but is not limited thereto. In this embodiment, memory 102 can also be a circuit or any other device capable of implementing storage functions for storing program instructions and / or data.
[0080] By designing and programming the processor 101, the code corresponding to the satellite pursuit adversarial decision-making method based on continuous-time hierarchical reinforcement learning described in the foregoing embodiments can be embedded into the chip, thereby enabling the chip to execute the code during runtime. Figure 1 The steps of the satellite pursuit adversarial decision-making method based on continuous-time hierarchical reinforcement learning are shown. How to design and program the processor 101 is a technique well-known to those skilled in the art and will not be described further here.
[0081] Example 3 This embodiment provides a computer-readable storage medium storing a computer program thereon. When the program is executed by a processor, it implements the steps of the satellite pursuit adversarial decision-making method based on continuous-time hierarchical reinforcement learning as described in Embodiment 1.
[0082] The computer-readable storage medium may include flash memory, hard disk, multimedia card, card-type memory (e.g., SD or DX memory), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic memory, magnetic disk, optical disk, etc. In some embodiments, the storage medium may be an internal storage unit of a computer device, such as the hard disk or memory of the computer device. In other embodiments, the storage medium may also be an external storage device of the computer device, such as a plug-in hard disk, smart memory card, secure digital card, flash memory card, etc., provided on the computer device. Of course, the storage medium may include both internal storage units and external storage devices of the computer device. In this embodiment, the memory is typically used to store the operating system and various application software installed on the computer device. In addition, the memory can also be used to temporarily store various types of data that have been output or will be output.
[0083] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and inventive concept of the present invention, should be covered within the scope of protection of the present invention.
Claims
1. A satellite pursuit adversarial decision-making method based on continuous-time hierarchical reinforcement learning, characterized in that, Includes the following steps: S1. Construct a multi-satellite pursuit and adversarial environment, establish a continuous-time relative motion model of the pursuing satellite, target satellite and guardian satellite in the local orbital coordinate system, define the state space and set the termination conditions of the adversarial process; S2. Construct a Markov decision process model with an action scheduling mechanism. In this model, the control input for pursuing satellites is defined as a joint action including control acceleration and planned execution time. During the state evolution process, a single decision cycle is divided into a first evolution stage without control input and a second evolution stage with control acceleration applied, based on the planned execution time. The relative motion state of each satellite is continuously updated based on the continuous time relative motion model. S3. Construct a continuous-time hierarchical reinforcement learning framework, which includes a management layer and a control layer; the management layer generates a set of candidate sub-targets based on the current relative motion state of each satellite and evaluates and selects the optimal sub-target; the control layer receives the current relative motion state of each satellite and the optimal sub-target, and outputs the joint action for the current decision cycle; S4. Configure evasion strategies for target satellites in the multi-satellite pursuit and adversarial environment, and configure pre-trained interception strategies for guardian satellites; S5. In the multi-satellite pursuit adversarial environment, the Markov decision process model is used to iteratively train the continuous-time hierarchical reinforcement learning framework in combination with the preset reward function to obtain a converged satellite pursuit decision strategy.
2. The satellite pursuit adversarial decision-making method based on continuous-time hierarchical reinforcement learning according to claim 1, characterized in that, In step S1, a virtual reference satellite with the same orbital parameters as the tracking satellite is set as the origin of the local orbital coordinate system; the continuous-time relative motion model is described by the Clohessy-Wiltshire equations, expressed as: In the formula, For the first The relative positions of the satellites; For the first The relative speed of the satellites; For the first The relative acceleration of the satellites; For the first Control acceleration of a satellite; The orbital angular velocity of the virtual reference satellite.
3. The satellite pursuit adversarial decision-making method based on continuous-time hierarchical reinforcement learning according to claim 1, characterized in that, In step S1, the state space includes the system state, which includes the motion state of the pursuing satellite itself, the relative motion state between the target satellite and the pursuing satellite, and the relative motion state between the guardian satellite and the pursuing satellite, expressed by the following formula: In the formula, express The system state at any given moment; and They represent Constantly track the satellite's position and velocity; and They represent The relative position and relative velocity of the target satellite and the pursuing satellite at any given time; and These represent the relative positions and relative velocities of the guardian satellite and the pursuing satellite, respectively. In step S2, the continuous updating of the relative motion states of each satellite is represented by a piecewise state transition equation: In the formula, express The first derivative; Indicates the current decision-making moment. Indicates the next decision point; This represents the control acceleration generated at the current decision-making moment; Indicates the planned execution time; This represents the state transition function.
4. The satellite pursuit adversarial decision-making method based on continuous-time hierarchical reinforcement learning according to claim 3, characterized in that, In step S3, the specific process by which the management layer generates the candidate sub-target set includes: Obtain the current distance between the pursuing satellite and the target satellite, and multiply the current distance by a preset scaling factor to obtain the search radius; The sub-target search area is constructed with the current position of the pursuing satellite as the center and the search radius as the specified area. Within the sub-target search area, a set of candidate sub-targets containing multiple spatial location points is generated; The optimal sub-objective is selected by evaluating it using the following formula: In the formula, The set of candidate sub-targets; Current system state Under the condition of selecting candidate sub-targets The expected cumulative reward that can be obtained; This is the optimal sub-objective.
5. The satellite pursuit adversarial decision-making method based on continuous-time hierarchical reinforcement learning according to claim 4, characterized in that, In step S3, the control layer outputs the joint action for the current decision cycle using the following formula: In the formula, For continuous action space; The output is the combined action of the current decision cycle; Given the current system state and optimal sub-objective Under the conditions, to carry out joint operations The expected cumulative reward that can be obtained.
6. The satellite pursuit adversarial decision-making method based on continuous-time hierarchical reinforcement learning according to claim 1, characterized in that, In step S4, the evasion strategy is as follows: when the distance between the pursuing satellite and the target satellite is greater than the preset threat distance, the target satellite performs a straight-line escape maneuver in the opposite direction of the relative position vector; when the distance is less than the preset threat distance, the target satellite performs a lateral evasion maneuver with periodic direction switching on the basis of the reverse escape.
7. The satellite pursuit adversarial decision-making method based on continuous-time hierarchical reinforcement learning according to claim 1, characterized in that, In step S4, the interception strategy is a continuous control strategy trained by a continuous-time reinforcement learning algorithm based on the Actor-Critic architecture. The satellite protection reward function used in the pre-training process consists of target protection reward, interception reward, occupancy reward and terminal reward. The target protection reward is used to constrain the guardian satellite to accompany the target satellite, and its value is negatively correlated with the deviation of the current distance between the guardian satellite and the target satellite from the expected protection distance. The interception reward is adjusted by a dynamic threat factor, the value of which is negatively correlated with the distance between the pursuing satellite and the target satellite. When the pursuing satellite enters the set threat perception range, the smaller the distance between the guardian satellite and the pursuing satellite, the greater the interception reward. Regarding the occupancy reward, a positive reward is given when the guardian satellite is located between the area connecting the pursuing satellite and the target satellite, and the position error is less than a preset threshold. The terminal rewards include positive rewards given when the guardian satellite successfully intercepts the pursuing satellite, and negative penalties given when the pursuing satellite successfully captures the target satellite, or when the guardian satellite collides with or engages with the target satellite and the confrontation exceeds the time limit.
8. The satellite pursuit adversarial decision-making method based on continuous-time hierarchical reinforcement learning according to claim 1, characterized in that, In step S5, the preset reward function includes management layer reward and control layer reward; the management layer reward consists of sub-target quality reward, target proximity reward and sub-target arrival reward; Specifically, for the sub-target quality reward, a positive reward is given when the generated sub-target is closer to the target satellite than the current position of the pursuing satellite; The value of the target proximity reward is positively correlated with the amount of reduction in distance between the pursuing satellite and the target satellite after the current sub-target is executed; For the sub-target arrival reward, a positive reward is given when the distance between the tracking satellite and the sub-target is less than a preset arrival threshold; The control layer reward consists of target approach reward, sub-target guidance reward, direction consistency reward, safety penalty, energy penalty, time penalty and terminal reward; The value of the target approach reward is positively correlated with the amount of distance reduction between the pursuing satellite and the target satellite at the current moment; The value of the sub-target guidance reward is positively correlated with the amount of distance reduction between the current tracking satellite and the current sub-target at the current moment; The directional consistency reward is determined by calculating the inner product of the direction vector of the pursuing satellite pointing to the target satellite and the control acceleration vector, and a positive reward is given when the control acceleration direction is consistent with the target direction. The safety penalty is triggered when the distance between the pursuing satellite and the guardian satellite is less than a safe distance threshold, in order to constrain the pursuing satellite to avoid the guardian satellite; The value of the energy penalty is positively correlated with the amplitude of the control acceleration in order to suppress excessive control input; The value of the time penalty is positively correlated with the size of the planned execution time, so as to constrain the control layer to output actions that approach the minimum execution time threshold; The terminal rewards include positive rewards for successfully capturing the target satellite and negative penalties for being intercepted by a protected satellite or for exceeding the timeout period.
9. A computer terminal, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the satellite pursuit adversarial decision-making method based on continuous-time hierarchical reinforcement learning as described in any one of claims 1 to 8.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the steps of the satellite pursuit adversarial decision-making method based on continuous-time hierarchical reinforcement learning as described in any one of claims 1 to 8.