A Star Cluster Orbit Pursuit and Evasion Decision-Making Method Based on Multi-Proximal Reinforcement Learning
By introducing multi-near-end reinforcement learning methods in satellite orbit control, combining the 8th-order Longge-Kuta numerical integration and the near-end strategy optimization algorithm, the problem of traditional methods being difficult to cope with changes in complex space environments is solved, and more efficient and safe satellite cluster orbit control is achieved.
Patent Information
- Application Number
- CN202510439758.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-09
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2045-04-09
AI Technical Summary
Traditional satellite orbit control methods are difficult to cope with dynamic changes and sudden demands in complex space environments, especially when large-scale star clusters are operated dynamically.
The satellite state information is updated using the star cluster orbital pursuit decision-making method based on multi-near-end reinforcement learning, and the satellite state information is updated using the 8th-order Longge-Kuta numerical integral method, and the decision-making process is optimized by the near-end strategy optimization algorithm and the federal near-end algorithm to achieve coordination and policy fusion among satellites.
It improves the overall decision-making efficiency and response speed of the satellite cluster, ensuring an efficient and safe operation in a dynamic environment.
Smart Images

Figure CN119962403B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of aerospace orbit planning, and specifically, to a swarm orbit pursuit-evasion decision-making method based on multi-proximal reinforcement learning. Background Art
[0002] In the field of modern aerospace orbit planning, precisely adjusting and maintaining a satellite's predetermined orbit in a complex space environment is a key technical challenge. Traditional orbit control methods typically rely on detailed mathematical models and preset control strategies, which require in-depth prediction and understanding of the satellite's dynamic behavior during design. However, the dynamic nature and uncertainty of the actual space environment often make it difficult for traditional control strategies to meet sudden orbit adjustment requirements, especially when dealing with the complexity and real-time requirements of large-scale swarm dynamic operations.
[0003] With the rapid development of artificial intelligence technology, adaptive control technologies based on intelligent algorithms have been introduced into the field of satellite orbit control, demonstrating superior performance. Compared with traditional methods, these intelligent algorithms can learn and adapt to environmental changes in real time, effectively improving the accuracy and efficiency of satellite orbit control by dynamically adjusting control strategies. When managing and adjusting the orbits of a satellite swarm, intelligent algorithms can better handle the interactions between the satellites and uncertain environmental factors through continuous learning and optimization. Multi-agent reinforcement learning technology is particularly prominent in this field, as it allows multiple agents (i.e., satellites) to collaborate in learning and decision-making without central control, greatly enhancing the flexibility and scalability of the system.
[0004] Therefore, it is necessary to propose a swarm orbit pursuit-evasion decision-making method based on multi-proximal reinforcement learning to solve the problem of decision-making for the pursuit-evasion behavior of a satellite swarm in a dynamic and complex orbit environment. Summary of the Invention
[0005] The purpose of the present invention is to overcome the defects and deficiencies of the prior art, and to provide a swarm orbit pursuit-evasion decision-making method based on multi-proximal reinforcement learning, which uses the proximal policy optimization algorithm and combines the federated proximal algorithm to improve the overall decision-making efficiency and response speed of the satellite swarm.
[0006] To achieve the above purpose, the present invention provides the following technical solutions:
[0007] A swarm orbit pursuit-evasion decision-making method based on multi-proximal reinforcement learning, characterized in that it specifically includes the following steps:
[0008] S1. Use the 8th-order Runge-Kutta numerical integration method to update the state information of all satellites in the satellite swarm, including position information and velocity information;
[0009] S2. Based on the status information of all satellites obtained in step S1, each satellite in the satellite swarm collects the status information of all other satellites respectively, constructs a status vector, and then obtains a policy network and a value network. Then, the corresponding action is selected and executed through the policy network. After executing the action, the reward feedback by the environment is received, and the status vector of the next time step is updated. Then, the temporal difference error and the advantage function are calculated in turn. Then, the proximal policy optimization algorithm is used to update the policy network and the value network, and the federated proximal algorithm is used to fuse the policy networks of each satellite in the satellite swarm;
[0010] S3. Based on step S2, the comprehensive performance of each satellite in the satellite swarm is evaluated, and the update of the policy network parameters of each satellite is dynamically adjusted to continuously optimize the overall performance of each satellite in the satellite swarm.
[0011] Furthermore, the specific steps of step S1 are as follows:
[0012] S11. Using the 8th-order Runge-Kutta method, a time step is evenly divided into multiple calculation stages. In each calculation stage, the position and velocity of the satellite are estimated in the middle according to the calculation result of the previous calculation stage and the predefined coefficients, and a set of status information at intermediate times is obtained. Then, using this set of intermediate status information, 8 position slopes used to approximate the position change trend are gradually obtained and 8 velocity slopes used to approximate the velocity change trend ; Then, all 8 position slopes and 8 velocity slopes are added according to the weights, and the new position and new velocity of the satellite at the end of this time step can be obtained, approaching the ordinary differential equations of satellite dynamics, that is , where is the satellite dynamics equation, which determines the acceleration change of the satellite at a given position and velocity ;
[0013] The position slope is determined by the velocity vector of the satellite at the moment, and the formula is as follows:
[0014] (1)
[0015] The velocity slope is calculated at the th intermediate node. At this time, the satellite position and velocity are the intermediate state comprehensively estimated from the end state of the previous time step and several previous slopes , and this intermediate state can be regarded as at time The predicted values of the satellite position and velocity are calculated through the satellite dynamics equation to calculate the corresponding velocity slope , and the formula is as follows:
[0016] (2)
[0017] In equations (1) and (2), represents the time node coefficient related to the th order, represents the time step, and represent the coefficients related to the th order, and respectively represent the satellite position vector and velocity vector at the
[0018] S12. Calculate the 8th order position slope and 8th order velocity slope of all satellites in the satellite group according to step S11, and then sum them up with weights according to the weight coefficient to update the state information of all satellites in the satellite group. The formula is as follows:
[0019] (3)
[0020] (4)
[0021] In equations (3) and (4), P t+1 and v t+1 respectively represent the satellite position vector and velocity vector at time t + 1.
[0022] Furthermore, in the said step S11, , and The specific numerical values are determined according to the 8th order Runge - Kutta method adopted.
[0023] Furthermore, the said step S2 specifically includes the following steps:
[0024] S21. Satellite collects the state information of all other satellites in the satellite group, including the position vector and velocity vector, and constructs the state vector , that is:
[0025] (5)
[0026] In equation (5), represents the total number of satellites in the satellite group,
[0027] Denote the satellite at time the position vector at the moment,
[0028] Denote the satellite at time the velocity vector at the moment.
[0029] S22. Define the policy network of the satellite as , which receives the state vector as input through the policy network parameter and outputs the probability distribution of the action vector .
[0030] Among them, Denote the satellite the decision-making function, then there is:
[0031] (6)
[0032] In formula (6), Denote the policy neural network with parameter , the softmax function ensures that the output is a legal probability distribution;
[0033] And, define the value network of the satellite as , which receives the state vector as input through the value network parameter and outputs the expected cumulative return of this state;
[0034] Among them, Denote the value function of the satellite , then there is:
[0035] (7)
[0036] In formula (7), Denote the value neural network with parameter .
[0037] S23. At time t, the satellite selects and executes the action corresponding to the action vector according to its policy network. After executing the action, the satellite receives the reward fed back by the environment, and updates to obtain the next state through the satellite dynamics equation solved by the 8th-order Runge-Kutta method, that is, the state vector at time t+1.
[0038] S24. Define the advantage function of the satellite as , where represents the discount factor, represents the smoothing parameter of the Generalized Advantage Estimation, which is used to balance bias and variance, represents the temporal difference error;
[0039] S25. Use the Proximal Policy Optimization algorithm to update the policy network and value network of the satellite , and use the loss function of the policy network of the satellite and the loss function of the value network to represent, and the specific steps are as follows:
[0040] Define the loss function of the policy network of the satellite as:
[0041] (8)
[0042] In formula (8), , represents the probability ratio between the current policy and the old policy, is a hyperparameter used to control the update amplitude of the policy network, usually taking values between 0.1 and 0.3; represents the parameters of the policy network of the satellite , represents the input state of the policy network and outputs the probability distribution of taking each possible action in this state, represents the parameters of the old policy network, that is, the parameters of the policy network before the update; represents the estimated value of the advantage function, which is used to measure the quality of the selected action vector relative to the average level under the state vector ; represents the expected value of the time step .
[0043] And, define the loss function of the value network of the satellite as:
[0044] (9)
[0045] In formula (9), , represents the target value network, where represents the immediate reward feedback by the environment after the satellite executes the action at the moment, Denote the state vector of the value network at time t+1 for value prediction, where \(\gamma\) represents the discount factor, and \(V^{\pi}(s)\) represents the value network of the satellite .
[0046] S26. The policy networks of each satellite in the satellite cluster are fused using the federated proximal algorithm, that is, the policy network parameters of the satellite are synchronized with the policy network parameters of adjacent satellites, and the policy network parameters of the satellite are updated through the federated proximal algorithm. The specific steps are as follows:
[0047] After updating the policy network of the satellite according to step S25, the satellite sends its policy network parameters to adjacent satellites. The satellite performs a fusion update based on the policy network parameters of adjacent satellites, so there is:
[0048] (10)
[0049] In formula (10), \(\alpha\) represents the learning step size, \(N_i\) represents the set of adjacent satellites of the satellite i, and \(w_{ij}\) represents the weight calculated based on the communication distance.
[0050] Furthermore, in step S24, the calculation formula of \(r_{t}\) is: where, \(r_{t}\) represents the immediate reward obtained after the satellite i performs an action at time t, \(\gamma\) represents the discount factor, \(V^{\pi}(s_t)\) represents the value prediction of the state vector \(s_t\) at time t by the value network, and \(V^{\pi}(s_{t+1})\) represents the value prediction of the state vector \(s_{t+1}\) at time t+1 by the value network.
[0051] Furthermore, in step S26, the calculation formula of \(w_{ij}\) is: where, \(d_{ij}\) represents the communication distance between the satellite i and the adjacent satellite j, and \(d_{ik}\) represents the communication distance between the satellite Communication distance between
[0052] Furthermore, step S3 specifically includes the following steps:
[0053] S31. Conduct a comprehensive performance evaluation of the satellite , and the evaluation metrics specifically include:
[0054] Orbit maintenance : Used to measure the orbit maintenance performance of the satellite at moment,
[0055] Fuel efficiency : Used to measure the fuel usage efficiency of the satellite at moment,
[0056] Collision avoidance : Used to measure the collision avoidance performance of the satellite at moment;
[0057] Define the performance score of the satellite as , then there is:
[0058] (11)
[0059] In formula (11), respectively correspond to the weight coefficients used to weigh orbit maintenance , fuel efficiency and collision avoidance .
[0060] S33. Dynamically adjust the update of the policy network parameters according to the performance score of the satellite obtained in step S32, which specifically includes the following steps:
[0061] S331. Performance weighted learning step size: Use the performance score as the adjustment factor of the learning step size , then there is:
[0062] (12)
[0063] In formula (12), represents the benchmark learning step size, represents the highest performance score of all satellites in the satellite group at moment.
[0064] S332. Performance-weighted Policy Network Fusion: During the process of fusing the policy network using the Federated Proximal Algorithm in step S26, taking the performance score as the weight factor for policy network fusion, then equation (10) can be modified as:
[0065] (13)
[0066] In equation (13), represents the policy network parameters of satellite at time, represents the policy network parameters of satellite at time, represents the learning step size of satellite at time, represents the fusion weight used by satellite at time when obtaining parameters from adjacent satellite , represents the policy network parameters of adjacent satellite at time.
[0067] S34. Based on step S32, comprehensively evaluate the performance of all other satellites in the satellite cluster, and based on step S33, dynamically adjust the update of the policy network parameters of all other satellites in the satellite cluster, continuously optimizing the overall performance of each satellite in the satellite cluster.
[0068] Compared with the prior art, the beneficial effects of the present invention are:
[0069] 1. The present invention introduces the high-precision 8th-order Runge-Kutta (RK8) method, which can accurately update the state information of the satellite, including position information and velocity information, ensuring the accuracy and stability of orbit control.
[0070] 2. The present invention uses the Proximal Policy Optimization algorithm (PPO) and the Federated Proximal Algorithm to optimize the decision-making process of the satellite, realizing the coordination and policy fusion among the satellites in the satellite cluster.
[0071] 3. Through comprehensive performance evaluation and the federated learning mechanism, the present invention can continuously optimize the overall performance of the satellite cluster, ensuring that the satellite cluster maintains an efficient and safe operating state in a dynamic environment. Description of the Drawings
[0072] Figure 1 is a flowchart of an intelligent decision-making method for the orbital behavior of a satellite cluster based on multi-proximal multi-agent reinforcement learning according to the present invention. Detailed Embodiments
[0073] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0074] See Figure 1 , a star cluster orbit pursuit and evasion decision-making method based on multi-proximal reinforcement learning, specifically including the following steps:
[0075] S1. The 8th-order Runge-Kutta numerical integration method is used to update the state information of the satellite, including position information and velocity information, to ensure the accuracy and stability of orbit control.
[0076] Specifically, step S1 includes the following steps:
[0077] S11. Using the 8th-order Runge-Kutta method, a time step is evenly divided into multiple calculation stages. In each calculation stage, the position and velocity of the satellite are estimated in the middle according to the calculation results of the previous calculation stage and predefined coefficients, obtaining a set of state information at intermediate times. Then, using this set of intermediate state information, 8 position slopes used to approximate the position change trend and 8 velocity slopes used to approximate the velocity change trend are gradually obtained; then, all 8 position slopes and 8 velocity slopes are added according to weights, and the new position and new velocity of the satellite at the end of this time step can be obtained, approaching the ordinary differential equation system of satellite dynamics, that is , where is the satellite dynamics equation, which determines the acceleration change of the satellite at a given position and velocity ;
[0078] The position slope is determined by the velocity vector of the satellite at time, and the formula is as follows:
[0079] (1)
[0080] The velocity slope is calculated at the th intermediate node. At this time, the satellite position and velocity are the intermediate state comprehensively estimated from the end state of the previous time step and several previous slopes , this intermediate state can be regarded as at time The predicted values of the satellite position and velocity, through the satellite dynamics equation Calculate the corresponding velocity slope , the formula is as follows:
[0081] (2)
[0082] In formulas (1) and (2), Represents the time node coefficient related to the th order, Represents the time step, And Represents the coefficient related to the th order. The specific values of these coefficients are determined according to the adopted 8th-order Runge-Kutta method. And Respectively represent The satellite position vector and velocity vector at the
[0083] S12. Calculate the 8th-order position slope and 8th-order velocity slope of all satellites in the satellite group according to step S11, and then sum up all the 8th-order position slopes and 8th-order velocity slopes weighted by the weight coefficient to update and obtain the state information of all satellites in the satellite group. The formula is as follows:
[0084] (3)
[0085] (4)
[0086] In formulas (3) and (4), P t+1 and v t+1 Respectively represent the satellite position vector and velocity vector at time t + 1.
[0087] Through the above 8th-order Runge-Kutta method, the integral part is approximated with high precision, and the accurate update of the satellite position and velocity states after the time step is realized, thus ensuring the accuracy and stability of orbit control.
[0088] S2. Based on the status information of all satellites obtained in step S1, each satellite in the satellite swarm collects the status information of all other satellites, constructs a status vector, and then obtains a policy network and a value network. Then, the corresponding action is selected and executed through the policy network. After executing the action, the reward feedback from the environment is received, and the status vector of the next time step is updated. Then, the temporal difference error and the advantage function are calculated in turn. Then, the proximal policy optimization algorithm is used to update the policy network and the value network, and the federated proximal algorithm is used to fuse the policy networks of each satellite in the satellite swarm to achieve cooperation and policy fusion among satellites.
[0089] Specifically, step S2 includes the following steps:
[0090] S21. The satellite collects the status information of all other satellites in the satellite swarm, including the position vector and the velocity vector, and constructs a status vector , that is:
[0091] (5)
[0092] In formula (5), represents the total number of satellites in the satellite swarm,
[0093] represents the satellite at time the position vector,
[0094] represents the satellite at time the velocity vector.
[0095] S22. Define the policy network of satellite as , which receives the status vector as input through the policy network parameter , and outputs the probability distribution of the action vector .
[0096] Among them, represents the satellite decision function, then there is:
[0097] (6)
[0098] In formula (6), represents the policy neural network with parameter , and the softmax function ensures that the output is a legal probability distribution.
[0099] And, define the satellite The value network is defined as , which receives the state vector through the value network parameter as input and outputs the expected cumulative reward for that state;
[0100] Among them, represents the value function of the satellite , then there is:
[0101] (7)
[0102] In formula (7), represents the value neural network with parameter .
[0103] S23. At time step , the satellite selects and executes an action according to its policy network. After executing the action, the satellite receives the reward fed back by the environment, and updates to obtain the next state through the satellite dynamics equation solved by the 8th-order Runge-Kutta method .
[0104] S24. The advantage function of the satellite is defined as , where represents the discount factor, represents the smoothing parameter of the generalized advantage estimation, which is used to balance bias and variance, represents the temporal difference error, and its calculation formula is: .
[0105] Specifically, the calculation formula of is: where represents the immediate reward obtained after the satellite executes the action at time represents the discount factor, represents the value prediction of the value network for the state vector at time t, represents the value prediction of the value network for the state vector at time t+1.
[0106] S25. The proximal policy optimization algorithm is used to update the policy network and value network of the satellite , and the loss function of the policy network of the satellite and the loss function of the value network are respectively adopted It is shown that the specific steps are as follows:
[0107] For the satellite the loss function of the policy network is defined as:
[0108] (8)
[0109] In Equation (8), , which represents the probability ratio between the current policy and the old policy, is a hyperparameter used to control the magnitude of the policy network update, usually taking values between 0.1 and 0.3; represents the parameters of the policy network of the satellite , represents the input state of the policy network and outputs the probability distribution of taking each possible action in this state, represents the parameters of the old policy network, that is, the parameters of the policy network before the update; represents the estimated value of the advantage function, which is used to measure the superiority or inferiority of the selected action vector relative to the average level under the state vector ; represents the expected value for the time step .
[0110] And, for the satellite the loss function of the value network is defined as:
[0111] (9)
[0112] In Equation (9), , which represents the target value network, is used to assist in training the value network. It is a reference value used to make the output of the value network closer to the actual return, that is: at the state the immediate reward is obtained and the transition is made to the next state , and the target value of this state is the immediate reward plus the value estimate of the next state.
[0113] Among them, represents the immediate reward feedback by the environment after the satellite executes the action at time represents the value prediction of the value network for the state vector at time t + 1 , represents the discount factor, represents the value network of the satellite .
[0114] S26. The policy networks of each satellite in the satellite constellation are fused using the federated proximal algorithm, that is, the policy network parameters of the satellite are synchronized with the policy network parameters of adjacent satellites, and the policy network parameters of the satellite are updated through the federated proximal algorithm. The specific steps are as follows: are , and the specific steps are as follows:
[0115] After updating the policy network of the satellite according to step S25, the satellite sends its policy network parameters to adjacent satellites. The satellite performs fusion update based on the policy network parameters of adjacent satellites, so there is:
[0116] (10)
[0117] In formula (10), represents the learning step size, represents the set of adjacent satellites of the satellite , represents the weight calculated based on the communication distance, and its calculation formula is: , where represents the communication distance between the satellite and the adjacent satellite , represents the communication distance between the satellite and the adjacent satellite .
[0118] Through the above weight allocation mechanism, the satellite can effectively utilize the policy information of adjacent satellites, ensuring the consistency and coordination of group behavior.
[0119] S3. Based on step S2, the comprehensive performance of each satellite in the satellite constellation is evaluated, and the update of the policy network parameters of each satellite is dynamically adjusted to continuously optimize the overall performance of each satellite in the satellite constellation, so as to achieve an efficient and safe operating state of the satellite constellation in a dynamic environment.
[0120] Specifically, step S3 includes the following steps:
[0121] S31. Define the reward function as , represents the distance between the current orbital state and the target orbital state , represents the 2-norm of the action vector , represents the adjustment coefficient.
[0122] S32. Satellite Conduct a comprehensive performance evaluation, with the following evaluation indicators:
[0123] Track maintenance :Used to measure satellite exist Track maintenance performance at all times,
[0124] Fuel efficiency :Used to measure satellite exist Fuel efficiency at all times,
[0125] Collision Avoidance :Used to measure satellite exist Time-to-time collision avoidance performance;
[0126] Defining Satellite The performance score is , then:
[0127] (11)
[0128] In formula (11), They correspond to the values used to balance track maintenance , fuel efficiency and collision avoidance The weight coefficient of .
[0129] S33. Satellites obtained in step S32 Performance rating The update of the network parameters of the dynamic adjustment strategy specifically includes the following steps:
[0130] S331. Performance-weighted learning step: The performance score As the learning step The adjustment factor is:
[0131] (12)
[0132] In formula (12), represents the benchmark learning step size, Indicates that all satellites in the satellite group are The highest performance score at the moment.
[0133] S332. Performance-weighted policy network fusion: In the process of using the federated proximal algorithm to fuse the policy network in step S26, the performance score As the weight factor of the policy network fusion, equation (10) can be modified as follows:
[0134] (13)
[0135] In formula (13), represents the policy network parameters of the satellite at time, represents the policy network parameters of the satellite at time, represents the learning step size of the satellite at time, represents the fusion weight used by the satellite at time when obtaining parameters from the adjacent satellite and represents the policy network parameters of the adjacent satellite at time.
[0136] S34. Comprehensively evaluate the comprehensive performance of all other satellites in the satellite group according to step S32, and dynamically adjust the update of the policy network parameters of all other satellites in the satellite group according to step S33, continuously optimizing the overall performance of each satellite in the satellite group, so that each satellite can effectively utilize the information learned from other satellites and its own experience while maintaining its own operation autonomy, and jointly optimize the group behavior.
[0137] Through the above steps, the satellite group can achieve continuous training and performance optimization of reinforcement learning, ensure that the satellite group maintains an efficient and safe operating state in a dynamic environment, and complete the intelligent decision-making planning of multi-proximal multi-agent reinforcement learning.
[0138] Although this specification is described according to the implementation manners, not every implementation manner only contains an independent technical solution. This narrative manner of the specification is only for clarity. Those skilled in the art should regard the specification as a whole, and the technical solutions in each embodiment can also be appropriately combined to form other implementation manners that can be understood by those skilled in the art.
[0139] Therefore, the above description is only a preferred embodiment of the present application, and is not used to limit the scope of implementation of the present application; that is, all equivalent transformations made according to the scope of the claims of the present application are within the protection scope of the claims of the present application.
Claims
1. A satellite cluster orbit pursuit decision method based on multi-proximal reinforcement learning, characterized by: The specific steps include: S1. Use the 8th-order Runge-Kutta numerical integration method to update the status information of all satellites in the satellite constellation, including position information and velocity information; S2. Based on the state information of all satellites obtained in step S1, each satellite in the satellite cluster collects the state information of all other satellites and constructs a state vector, thereby obtaining a policy network and a value network, and then selects and executes the corresponding action through the policy network; after executing the action, it receives the reward from the environment feedback, updates the state vector of the next time step, and then calculates the temporal difference error and advantage function in turn, and then uses the proximal policy optimization algorithm to update the policy network and the value network, and uses the federated proximal algorithm to fuse the policy networks of each satellite in the satellite cluster; S3. Based on step S2, a comprehensive performance evaluation is performed on each satellite in the satellite cluster, and the update of the strategic network parameters of each satellite is dynamically adjusted to continuously optimize the overall performance of each satellite in the satellite cluster.
2. According to claim 1, a method for swarm orbit pursuit decision-making based on multi-proximal reinforcement learning is characterized by: The step S1 specifically includes the following steps: S11. Use the 8th-order Runge-Kutta method to convert a time step The calculation is divided into multiple stages. Each stage will make an intermediate estimate of the satellite position and speed based on the calculation results of the previous stage and the predefined coefficients to obtain a set of state information at the intermediate moment. Then, using this set of intermediate state information, 8 position slopes are gradually calculated to approximate the position change trend. and 8 velocity slopes to approximate velocity trends ; Then set the slopes of all 8 positions and 8 speed slopes By adding the weights, we can get this time step The new position and velocity of the final satellite approximate the ordinary differential equations of satellite dynamics, namely, ,in, is the satellite dynamics equation, which determines the satellite at a given position and speed Acceleration changes under Position slope By satellite Velocity vector at time Determine, the formula is as follows: (1) Speed slope In the The satellite position and velocity are calculated at the intermediate node. The satellite position and velocity at this time are the state at the end of the previous time step. and the intermediate state obtained by comprehensive estimation of several preceding slopes , this intermediate state can be regarded as The predicted values of satellite position and velocity are obtained through the satellite dynamics equation Calculate the corresponding velocity slope , the formula is as follows: (2) In formulas (1) and (2), Indicates The time node coefficients of order correlation, represents the time step, and Indicates The coefficient of the order correlation, and Respectively Satellite position vector and velocity vector at the moment; S12. Calculate the 8th order position slope and 8th order velocity slope of all satellites in the satellite constellation according to step S11, and then calculate all the 8th order position slopes and 8th order velocity slopes according to the weight coefficients Weighted summation is performed to update the status information of all satellites in the satellite cluster. The formula is as follows: (3) (4) In formulas (3) and (4), P t+1 and v t+1 They represent the satellite position vector and velocity vector at time t+1 respectively.
3. The method for swarm orbit pursuit decision-making based on multi-proximal reinforcement learning according to claim 2 is characterized in that: In the step S11, , and The specific value of is determined by the 8th-order Runge-Kutta method used.
4. The method for swarm orbit pursuit decision-making based on multi-proximal reinforcement learning according to claim 2 is characterized in that: The step S2 specifically includes the following steps: S21. Satellite Collect the status information of all other satellites in the satellite group, including position vector and velocity vector, and construct the state vector ,Right now: (5) In formula (5), Represents the total number of satellites in the satellite constellation, Indicates satellite In time The position vector at time, Indicates satellite In time The velocity vector at the moment; S22. Satellite The policy network is defined as , which is achieved through the policy network parameters Receive state vector As input, the output action vector The probability distribution of in, Indicates satellite The decision function is: (6) In formula (6), Indicates that the parameter is The policy neural network, the softmax function ensures that the output is a legal probability distribution; And, the satellite The value network is defined as , which is expressed by the value network parameter Receive state vector As input, output the expected cumulative reward of this state; in, Indicates satellite The value function of is: (7) In formula (7), Indicates that the parameter is The value of neural network; S23. At time t, the satellite Select and execute an action vector based on its policy network The corresponding action, after executing the action, the satellite Rewards for receiving feedback from the environment , and the satellite dynamics equations solved by the 8th-order Runge-Kutta method , update to get the next moment, that is, the state vector at moment t+1 ; S24. Satellite The advantage function is defined as ,in, represents the discount factor, represents the smoothing parameter of the generalized advantage estimator, which is used to balance the bias and variance, represents the timing difference error; S25. Use the proximal strategy optimization algorithm to The strategy network and value network are updated by using satellite The loss function of the policy network And the loss function of the value network The specific steps are as follows: Satellite The loss function of the policy network Defined as: (8) In formula (8), , represents the probability ratio between the current strategy and the old strategy, It is a hyperparameter used to control the magnitude of the policy network update, usually between 0.1 and 0.3; Indicates satellite The parameters of the policy network, Represents the policy network input state Then output the probability distribution of each possible action taken in this state, Represents the parameters of the old policy network, that is, the parameters of the policy network before updating; Represents the estimated value of the advantage function, which is used to measure the state vector The selected action vector The degree of superiority or inferiority relative to the average level; Represents the time step Expected value; And, the satellite The loss function of the value network Defined as: (9) In formula (9), , represents the target value network; among them, Indicates satellite implement The immediate reward from the environment after the action. Represents the state vector of the value network at time t+1 The value prediction of represents the discount factor, Indicates satellite value network; S26. Use the federated proximal algorithm to fuse the strategic networks of each satellite in the satellite cluster. The policy network parameters Synchronizes strategic network parameters with neighboring satellites and updates satellites through a federated proximal algorithm The policy network parameters , the specific steps are as follows: According to step S25, the satellite After the strategic network is updated, the satellite Its policy network parameters Send to adjacent satellites, satellite Based on the strategic network parameters of adjacent satellites, the fusion update is: (10) In formula (10), represents the learning step length, Indicates satellite The set of neighboring satellites of Represents the weight calculated based on the communication distance.
5. The method for swarm orbit pursuit decision-making based on multi-proximal reinforcement learning according to claim 4 is characterized in that: In the step S24, The calculation formula is: ,in, Indicates satellite exist The instant reward after performing the action at any time, represents the discount factor, Represents the state vector of the value network at time t The value prediction of Represents the state vector of the value network at time t+1 value prediction.
6. The method for swarm orbit pursuit decision-making based on multi-proximal reinforcement learning according to claim 4, characterized in that: In the step S26, The calculation formula is: ,in, Indicates satellite With adjacent satellites The communication distance between Indicates satellite With adjacent satellites The communication distance between them.
7. The method for swarm orbit pursuit decision-making based on multi-proximal reinforcement learning according to claim 4, characterized in that: The step S3 specifically includes the following steps: S31. Satellite Conduct a comprehensive performance evaluation, with the following evaluation indicators: Track maintenance :Used to measure satellite exist Track maintenance performance at all times, Fuel efficiency :Used to measure satellite exist Fuel efficiency at all times, Collision Avoidance :Used to measure satellite exist Time-to-time collision avoidance performance; Defining Satellite The performance score is , then: (11) In formula (11), They correspond to the values used to balance track maintenance , fuel efficiency and collision avoidance The weight coefficient of S33. Satellites obtained in step S32 Performance rating The update of the network parameters of the dynamic adjustment strategy specifically includes the following steps: S331. Performance-weighted learning step: The performance score As the learning step The adjustment factor is: (12) In formula (12), represents the benchmark learning step size, Indicates that all satellites in the satellite cluster are The highest performance score at the moment; S332. Performance-weighted policy network fusion: In the process of using the federated proximal algorithm to fuse the policy network in step S26, the performance score As the weight factor of the policy network fusion, equation (10) can be modified as follows: (13) In formula (13), Indicates satellite exist The policy network parameters at time t, Indicates satellite exist The policy network parameters at time t, Indicates satellite exist The learning step length at each moment, Indicates satellite exist Time from adjacent satellites The fusion weights used when obtaining parameters, Indicates adjacent satellites exist Policy network parameters at the moment; S34. Perform a comprehensive performance evaluation on all other satellites in the satellite cluster according to step S32, and dynamically adjust the update of the strategic network parameters of all other satellites in the satellite cluster according to step S33, so as to continuously optimize the overall performance of each satellite in the satellite cluster.
Citation Information
Patent Citations
Multi-spacecraft chasing game orbit control method
CN116449714A
TD3 soft reinforcement learning spacecraft attitude control method and computer readable medium
CN116788524A