An Observation Constellation Optimization Method Based on Deep Reinforcement Learning

Through the improved TD3 algorithm and dynamic strategy noise mechanism, the problem of observation constellation instability caused by traditional TD3 technology under large noise changes is solved, the training efficiency and stability of constellation optimization design is improved, and the constellation observation performance is enhanced.

CN119004972BActive Publication Date: 2025-07-11SICHUAN UNIV
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202411072310.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-08-06
Publication Date
2025-07-11
Estimated Expiration
2044-08-06

AI Technical Summary

Technical Problem

Traditional TD3 technology causes unstable observation constellations, low learning and training efficiency, and unstable model optimization when the noise is changed before and after.

Method used

The observation constellation optimization method based on deep reinforcement learning is adopted, and the state space model, action space model and reward function are constructed to optimize the constellation strategy through the improved TD3 algorithm combined with dynamic strategy noise mechanism and priority experience playback.

Benefits of technology

It improves the training efficiency and stability of the zodiac optimization design decision model, enhances the zodiac observation performance to the ground, reduces energy consumption and extends service life.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119004972B_ABST
    Figure CN119004972B_ABST
Patent Text Reader

Abstract

The present invention discloses an observation constellation optimization method based on deep reinforcement learning, which relates to the field of satellite intelligence technology and includes the following steps: determining the constellation configuration of the constellation to be observed and obtaining its satellite orbit parameters; constructing a constellation optimization strategy by using a reinforcement learning algorithm based on the satellite orbit parameters; establishing a state space model, an action space model and designing a reward function; processing the state space model, the action space model and the reward function by using the improved TD3 algorithm based on the optimal policy function to obtain a final optimization strategy; and completing the optimization of the observation constellation by the constellation to be observed executing the final optimization strategy. The present invention improves the disadvantages of the traditional TD3, such as low learning and training efficiency and instability of the optimization model, improves the training efficiency and stability of the constellation optimization design decision model by introducing prioritized experience replay and dynamic policy noise, and improves the constellation's ground observation performance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of satellite intelligence technology, and particularly to an observation constellation optimization method based on deep reinforcement learning. Background Art

[0002] An observation constellation is a type of satellite constellation for observing targets on the earth, sea, and air. The optimization of an observation constellation refers to a series of technical activities aimed at improving the observation efficiency, reducing costs, expanding application fields, etc. through comprehensive and integrated analysis and optimization of the satellite constellation system. In constellation design, the application of optimization techniques can greatly improve the observation performance of the constellation and is one of the important technical means to keep the satellite constellation with high observation performance and reduce operating costs. Therefore, it is of great significance to study efficient observation constellation optimization methods.

[0003] One of the challenges in observation constellation optimization is that to meet the requirements of variable observation tasks and complex environmental constraints, it is necessary to quickly adjust the observation path. As an intelligent control method, reinforcement learning learns the optimal strategy by interacting with the environment, improving the satellite's autonomous decision-making ability. In case of emergencies, the satellite can quickly adjust the plan, prioritize the observation of key areas, and provide instant information for decision-making. By optimizing the constellation, not only its adaptability is enhanced, but also the energy consumption is reduced and the service life is extended. Studying the application of reinforcement learning in observation constellation optimization is the key to promoting the intelligence of space technology and improving the observation efficiency, and has important practical significance for national security, environmental protection, etc.

[0004] Currently, in the traditional TD3 technology, the experience replay is that the agent model learns through uniform distribution sampling. However, its high-value samples are diluted by ordinary samples, making the model update ineffective and inefficient. Moreover, the noise added during action output is Gaussian noise with a stable standard deviation, which will cause the agent model to be unstable in the case of large changes in noise before and after. Summary of the Invention

[0005] Aiming at the above deficiencies in the prior art, an observation constellation optimization method based on deep reinforcement learning provided by the present invention solves the problem that the traditional TD3 technology will cause the observation constellation to be unstable in the case of large changes in noise before and after.

[0006] To achieve the above invention purpose, the technical solution adopted by the present invention is as follows:

[0007] An observation constellation optimization method based on deep reinforcement learning is provided, which includes the following steps:

[0008] S1. Determine the constellation configuration of the constellation to be observed and obtain its satellite orbital parameters;

[0009] S2. Based on the satellite orbital parameters, use the reinforcement learning algorithm to construct a constellation optimization strategy;

[0010] S3. Establish a state space model, an action space model, and design a reward function;

[0011] S4. Based on the optimal policy function, use the improved TD3 algorithm to process the state space model, the action space model, and the reward function to obtain the final optimized policy;

[0012] S5. Execute the final optimized policy through the constellation to be observed to complete the optimization of the observation constellation.

[0013] Furthermore, step S1 includes the following steps:

[0014] S1-1. Determine the constellation configuration of the constellation to be observed; where the constellation configuration includes a general constellation configuration and a Walker constellation configuration;

[0015] S1-2. Judge whether the constellation configuration of the constellation to be observed is a general constellation configuration; if so, enter step S1-3; otherwise, determine that the constellation configuration of the constellation to be observed is a Walker constellation configuration and enter step S1-4;

[0016] S1-3. Take the 6 first orbital parameters corresponding to each satellite in the general constellation configuration as the satellite orbital parameters and enter step S2; where the 6 first orbital parameters are respectively the semi-major axis of the satellite orbit, the orbital eccentricity, the orbital inclination, the argument of perigee, the right ascension of the ascending node, and the true anomaly;

[0017] S1-4. Determine the total number of satellites in the Walker constellation configuration and the orbital elements, phase factor, number of orbital planes, semi-major axis of the orbit, and orbital inclination of each satellite;

[0018] S1-5. Take the first satellite in the Walker constellation configuration as the seed satellite and obtain the corresponding right ascension of the ascending node and the true anomaly at the initial moment;

[0019] S1-6. According to the formula:

[0020]

[0021] Obtain the right ascension of the ascending node of the remaining satellites and the true anomaly ; where, represents the right ascension of the ascending node of the seed satellite, represents the number of orbital planes, represents pi, represents that the satellite is located in the th orbital plane, represents the th satellite on the orbital plane, Denote the true anomaly of the seed satellite at the initial moment;

[0022] S1-7. The total number of satellites in the Walker constellation , the phase factor, the number of orbital planes, the semi-major axis of the orbit, the inclination of the orbital plane, and the right ascension of the ascending node of all satellites, and the true anomaly at the initial moment are used as satellite orbit parameters, and then proceed to step S2.

[0023] Furthermore, step S2 includes the following steps:

[0024] S2-1. Take the constellation to be observed as an agent and determine the optimization task of the constellation to be observed. Build the corresponding MDP model based on the Markov decision process. The corresponding formula is:

[0025]

[0026] where denotes the state set, including satellite orbit parameters, denotes the action set, denotes the state transition function, denotes the immediate reward, denotes the discount factor;

[0027] S2-2. Calculate the cumulative reward corresponding to the execution of all decisions of the MDP model;

[0028] S2-3. Calculate the optimal state value function and the optimal state-action value function ; where denotes the action in the action set ;

[0029] S2-4. Based on the optimal state value function and the optimal state-action value function , determine the constellation optimization strategy.

[0030] Furthermore, the calculation formula for the cumulative reward is:

[0031]

[0032] where , respectively denote the cumulative rewards corresponding to time , time , denotes the summation function, , , , respectively denote time , time , time At a moment corresponding immediate rewards and punishments;

[0033] Optimal state value function The calculation formula is:

[0034]

[0035] wherein represents the maximum value function, and represents the state, represents the state corresponding optimal state value function, represents that when in the state the agent executes the action and transfers to the state with a probability of, represents that when in the state the agent executes the action the immediate rewards and punishments;

[0036] The calculation formula of the optimal state-action value function is:

[0037]

[0038] wherein represents that when in the state the agent executes the action the optimal state-action value function.

[0039] Furthermore, step S3 includes the following steps:

[0040] S3-1. Determine whether the constellation configuration of the constellation to be observed is a general constellation configuration; if so, go to step S3-2; otherwise, incorporate it into step S3-4;

[0041] S3-2. Set the semi-major axis of the satellite orbit of the general constellation configuration to 8000 km, set the satellite orbit to a circular orbit, set the eccentricity to 0, and set the argument of perigee to 0. Take the orbital inclination, right ascension of the ascending node, true anomaly, and position parameters as the state, that is, the state space model, and the corresponding expression is:

[0042]

[0043] wherein and respectively represent the orbital inclination of the first satellite and the second satellite of the constellation to be observed, , , respectively represent the right ascension of the ascending node of the first satellite and the second satellite of the constellation to be observed , ; , , , , respectively represent the right ascension of the ascending node of the first satellite and the second satellite of the constellation to be observed , , th satellite ; , , , , , respectively represent the longitude, latitude, and altitude values of the constellation to be observed in the GPS84 coordinate system;

[0044] S3 - 3. Establish an action space model for the general constellation configuration, that is, a state space model, and enter step S3 - 6; among them, the state space model of the general constellation configuration corresponding expression is:

[0045]

[0046] where , respectively represent the actions of the first satellite and the second satellite of the constellation to be observed at the orbital inclination , ; , ; , respectively represent the actions of the first satellite and the second satellite of the constellation to be observed at the right ascension of the ascending node , ; , ; , , respectively represent the actions of the first satellite, the second satellite, and the th satellite of the observed constellation at the true anomaly ; ; , , ;

[0047] S3-4. Set the semi-major axis of the satellite orbit in the Walker constellation configuration to a fixed value, and use the number of orbital planes, phase factor, orbital inclination, right ascension of the ascending node, true anomaly, and position parameters as states, i.e., the state space model. The corresponding expression is:

[0048]

[0049] Among them, 、 、 、 、 represent the states of the number of orbital planes, phase factor, orbital inclination, right ascension of the ascending node, and true anomaly in the Walker constellation orbit, respectively;

[0050] S3-5. Establish the action space model of the Walker constellation configuration, i.e., the state space model, and enter step S3-6; among them, the state space model of the Walker constellation configuration The corresponding expression is:

[0051]

[0052] Among them, 、 、 、 、 represent the actions of the number of orbital planes, phase factor, orbital inclination, right ascension of the ascending node, and true anomaly in the Walker constellation orbit, respectively;

[0053] S3-6. According to the formula:

[0054]

[0055] Construct the reward function ; among them, 、 represent the target observation performance reward and the optimization times reward, respectively, 、 represent the coefficients, respectively, and .

[0056] Furthermore, the improved TD3 algorithm adopts a dynamic policy noise mechanism and includes an Actor policy network and a Critic value network; among them, the Actor policy network includes a target policy network and an online policy network; the Critic value network includes a target value network and an online value network; the target value network includes a target1 value network and a target2 value network; the online value network includes an online1 value network and an online2 value network.

[0057] Further, the training process of the improved TD3 algorithm in step S4 includes the following steps:

[0058] S4-1. Set the total number of training episodes and the total number of steps per episode;

[0059] S4-2. Initialize the environment, that is, randomly initialize the training satellite orbit parameters within a preset value range and obtain the current state ;

[0060] S4-3. Based on the current state , select an action through the online policy network;

[0061] S4-4. Execute the action , and synchronize the simulation scenario based on STK to obtain the corresponding reward and the next state , and update the constellation optimal policy;

[0062] S4-5. Store the current state , the next state , the action and the reward as sample data into the experience pool ;

[0063] S4-6. Randomly sample N sample data from the experience pool , and select an action through the target policy network, and calculate the loss function ;

[0064] S4-6. According to the loss function , update the parameters of the target value network and the online value network;

[0065] S4-7. Calculate the policy gradient and update the parameters of the online policy network;

[0066] S4-8. Update the target policy network and the target value network in step S4-6 in a soft update manner;

[0067] S4-9. Determine whether the current number of steps has reached the total number of steps per episode; if so, the current episode ends and proceeds to step S4-10; otherwise, the current number of steps is incremented by 1 and returns to step S4-3; where the initial value of the number of steps is 0;

[0068] S4-10. Repeat steps S4-1 to S4-9 until the total number of training episodes is reached.

[0069] Furthermore, the formula in step S4-3 is as follows:

[0070]

[0071] where represents the action output by the online policy network when the current network parameters are for the current state ; represents the policy noise of the current step, 、 represent the value range of the action ; represents a Gaussian distribution with a mean of 0 and a standard deviation of ; represents the online policy network when the current network parameters are and in the current state ; ;

[0072] The loss function of S4-6 is calculated as follows:

[0073]

[0074] where represents the minimum value function, represents the network parameters of the target value network, represents the network parameters of the online value network, 、 respectively represent the th sample data and the state of the th sample data, represents the online value network, represents the action of the th sample data, represents the serial number of the online value network;

[0075] The policy gradient of step S4-7 is calculated as follows:

[0076]

[0077] where represents the value gradient output by the online value network when the current network parameters are for the current state , action ; represents the online policy network When the current network parameters are and according to the current state the output action gradient;

[0078] Furthermore, the experience pool adopts the prioritized experience replay mechanism.

[0079] The beneficial effects of the present invention are as follows: This method improves the disadvantages of the traditional TD3, such as low learning and training efficiency and instability of the optimized model. By introducing prioritized experience replay and dynamic policy noise, the training efficiency and stability of the constellation optimization design decision model are improved, and the constellation earth observation performance is enhanced. Brief Description of the Drawings

[0080] Figure 1 is the flowchart of the method;

[0081] Figure 2 is the flowchart of the improved TD3 algorithm;

[0082] Figure 3 is the reward curve graph of the Walker constellation configuration training process;

[0083] Figure 4 is the reward curve graph of the general constellation configuration training process;

[0084] Figure 5 is the comparison graph of the optimization results of different constellation configurations;

[0085] Figure 6 is the result graph of 20 particle swarm parameters;

[0086] Figure 7 is the optimization result graph of 100 particle swarm parameters;

[0087] Figure 8 is the constellation optimization observation efficiency curve graph;

[0088] Figure 9 is the constellation optimization observation efficiency curve graph based on the particle swarm algorithm; among them, (a) is the constellation optimization observation efficiency curve graph when the particle swarm size is 20, and (b) is the constellation optimization observation efficiency curve graph when the particle swarm size is 100;

[0089] Figure 10 is the comparison histogram based on the particle swarm algorithm and this method; among them, (a) is the comparison histogram of the observation performance based on the particle swarm algorithm and this method, and (b) is the comparison histogram of the time consumption based on the particle swarm algorithm and this method. Detailed Implementation Manner

[0090] The following describes the specific embodiments of the present invention to facilitate those skilled in the art of this technology to understand the present invention. However, it should be clear that the present invention is not limited to the scope of the specific embodiments. For those of ordinary skill in the art of this technology, as long as various changes are within the spirit and scope of the present invention defined and determined by the appended claims, these changes are obvious, and all inventions and creations using the concept of the present invention are within the scope of protection.

[0091] As Figure 1 shown, an observation constellation optimization method based on deep reinforcement learning includes the following steps:

[0092] S1. Determine the constellation configuration of the constellation to be observed and obtain its satellite orbital parameters;

[0093] Step S1 includes the following steps:

[0094] S1-1. Determine the constellation configuration of the constellation to be observed; among them, the constellation configuration includes a general constellation configuration and a Walker constellation configuration;

[0095] S1-2. Judge whether the constellation configuration of the constellation to be observed is a general constellation configuration; if so, proceed to step S1-3; otherwise, determine that the constellation configuration of the constellation to be observed is a Walker constellation configuration and proceed to step S1-4;

[0096] S1-3. Take the 6 first orbital parameters corresponding to each satellite of the general constellation configuration as the satellite orbital parameters, and proceed to step S2; among them, the 6 first orbital parameters are respectively the semi-major axis of the satellite orbit, the orbital eccentricity, the orbital inclination, the argument of perigee, the right ascension of the ascending node, and the true anomaly;

[0097] S1-4. Determine the total number of satellites in the Walker constellation configuration and the orbital elements, phase factor, number of orbital planes, semi-major axis of the orbit, and orbital inclination of each satellite;

[0098] S1-5. Take the first satellite of the Walker constellation configuration as the seed satellite, and obtain the corresponding right ascension of the ascending node and the true anomaly at the initial moment;

[0099] S1-6. According to the formula:

[0100]

[0101] Obtain the right ascension of the ascending node of the remaining satellites and the true anomaly ; among them, represents the right ascension of the ascending node of the seed satellite, represents the number of orbital planes, represents pi, represents that this satellite is located in the an orbital plane, indicating the th satellite on the orbital plane, indicating the true anomaly of the seed satellite at the initial moment;

[0102] S1-7. Take the total number of satellites in the Walker constellation configuration , the phase factor, the number of orbital planes, the semi-major axis of the orbit, the inclination of the orbital plane, and the right ascension of the ascending node of all satellites, and the true anomaly at the initial moment as satellite orbital parameters, and proceed to step S2.

[0103] Among them, the inclination of the orbital plane, the right ascension of the ascending node, and the true anomaly are all periodic.

[0104] S2. Based on the satellite orbital parameters, construct a constellation optimization strategy using the reinforcement learning algorithm;

[0105] Step S2 includes the following steps:

[0106] S2-1. Take the constellation to be observed as an agent and determine the optimization task of the constellation to be observed. Based on the Markov decision process, construct the corresponding MDP model, and the corresponding formula is:

[0107]

[0108] Among them, represents the state set, including satellite orbital parameters, and any element represents a possible state in the environment; represents the action set, and any element represents an action that the agent can choose; represents the state transition function, represents the immediate reward and punishment, is the discount factor, that is, a weight, indicating the tendency towards future persistent rewards and current immediate rewards;

[0109] S2-2. Calculate the cumulative reward corresponding to all decisions executed by the MDP model;

[0110] The formula for calculating the cumulative reward is:

[0111]

[0112] Among them, , respectively represent the cumulative rewards corresponding to time , time , represents the summation function, , , , respectively represent moments , moment , moment , moment , and the corresponding immediate rewards and punishments;

[0113] S2-3. Calculate the optimal state value function and the optimal state-action value function ; where represents an action in the action set ;

[0114] The optimal state value function and the optimal state-action value function are derived from the state value function and the state-action value function and the Bellman optimality principle.

[0115] The state value function and the state-action value function are used to estimate the cumulative reward and evaluate the quality of each state and state-action combination. The corresponding formulas are:[[]]

[0116]

[0117] where represents the policy, represents the state, represents the expected cumulative reward of the agent when executing the policy in the state , represents the expected cumulative reward of the agent when executing the policy in the state and performing the action ;

[0118] According to the Bellman optimality principle, the state value function and the state-action value function need to be rewritten using the Bellman equation, that is, we get:[[]]

[0119]

[0120]

[0121] Thus, the calculation formula for the optimal state value function is:[[]]

[0122]

[0123] where represents the maximum value function, , Indicates the state, Indicates the state The corresponding optimal state value function, Indicates that at state when the agent executes the action the probability of transitioning to state is Indicates that at state when the agent executes the action the immediate reward and punishment;

[0124] The calculation formula of the optimal state-action value function is:

[0125]

[0126] where, Indicates that at state when the agent executes the action the optimal state-action value function.

[0127] S2-4. Determine the constellation optimization strategy based on the optimal state value function and the optimal state-action value function .

[0128] Among them, the constellation optimization strategy must meet the following conditions:

[0129]

[0130] where, Indicates the probability that the agent takes the action according to the policy at state , Indicates that when the state is the probability that the agent executes the action .

[0131] S3. Establish a state space model, an action space model and design a reward function;

[0132] Step S3 includes the following steps:

[0133] S3-1. Determine whether the constellation configuration of the constellation to be observed is a general constellation configuration; if so, go to step S3-2; otherwise, go to step S3-4;

[0134] S3-2. Set the semi-major axis of the satellite orbit of the general constellation configuration to 8000 km, set the satellite orbit to a circular orbit, the eccentricity to 0, and the argument of perigee to 0. Take the orbital inclination, right ascension of the ascending node, true anomaly and position parameters as the state, that is, the state space model. The corresponding expression is:

[0135]

[0136] Among them, and respectively represent the orbital inclination of the first satellite and the second satellite of the constellation to be observed. and , and respectively represent the right ascension of the ascending node of the first satellite and the second satellite of the constellation to be observed. and , and and respectively represent the true anomaly of the first satellite , the second satellite , and the th satellite of the constellation to be observed. and and , and and respectively represent the longitude, latitude, and altitude values of the constellation to be observed in the GPS84 coordinate system;

[0137] S3-3. Establish an action space model for the general constellation configuration, i.e., the state space model, and enter step S3-6; among them, the state space model of the general constellation configuration corresponding expression is:

[0138]

[0139] Among them, and respectively represent the actions of the first satellite and the second satellite of the constellation to be observed in the orbital inclination and . and respectively represent the actions of the first satellite and the second satellite of the constellation to be observed in the right ascension of the ascending node and . and and respectively represent the first satellite and the second satellite and the a satellite at the true anomaly 、 、 actions;

[0140] S3-4. Set the semi-major axis of the satellite orbit in the Walker constellation configuration as a fixed value, and use the number of orbital planes, phase factor, orbital inclination, right ascension of the ascending node, true anomaly, and position parameters as states, that is, the state space model. The corresponding expression is:

[0141]

[0142] Among them, 、 、 、 、 respectively represent the states of the number of orbital planes, phase factor, orbital inclination, right ascension of the ascending node, and true anomaly in the Walker constellation orbit;

[0143] S3-5. Establish the action space model of the Walker constellation configuration, that is, the state space model, and enter step S3-6; among them, the state space model of the Walker constellation configuration The corresponding expression is:

[0144]

[0145] Among them, 、 、 、 、 respectively represent the actions of the number of orbital planes, phase factor, orbital inclination, right ascension of the ascending node, and true anomaly in the Walker constellation orbit;

[0146] Among them, whether it is the Walker constellation configuration or the general constellation configuration, the corresponding state values need to be normalized and mapped to the range of 0 to 1; the corresponding action output values need to be normalized and mapped to the range of -1 to 1. For example, the actual value range of the inclination angle is 0 to 180°, and the relationship between the initial state value and the state value after state transition and the action value satisfy the following formula:

[0147]

[0148] S3-6. According to the formula:

[0149]

[0150] Construct the reward function ; among which, , respectively represent the target observation performance reward and the optimization times reward, , respectively represent coefficients, and .

[0151] Among which, the relationship between the constellation's target observation performance reward and the constellation's target observation performance score obtained by the constellation's target observation performance evaluation method based on multi-level fuzzy comprehensive evaluation is as follows:

[0152]

[0153] The optimization times reward and the change amount of all parameters to be optimized before and after the current round of optimization is .

[0154] Among which, the process of the constellation's target observation performance evaluation method based on multi-level fuzzy comprehensive evaluation is as follows:

[0155] Firstly, construct the evaluation index system. The total target layer is "the constellation's observation performance on the target", the intermediate levels are "observation accuracy", "observation time resolution", "observation stability", etc., and the index layer is specific observation data or parameters. Then, determine the factor set and the comment set. The factor set is a general set composed of various factors affecting the evaluation target, represented by , among which represents the observation accuracy, represents the observation time resolution, etc. The comment set is a set composed of various results that the evaluator may give to the evaluation target, represented by , among which represents "excellent", represents "good", etc. Secondly, specify the weights of each factor, construct the fuzzy evaluation matrix, and construct the fuzzy evaluation matrix according to the evaluation grades in the comment set. Each element in the matrix represents the membership degree of a certain factor to a certain evaluation grade. Finally, conduct multi-level fuzzy comprehensive evaluation and analyze the results. Firstly, start from the bottom layer, conduct single-factor fuzzy evaluation for each sub-factor to obtain their respective fuzzy evaluation matrices. Then, use their respective weights and the fuzzy evaluation matrices for synthesis operations to obtain the fuzzy evaluation vector of the upper level. This process is carried out layer by layer upwards until the total target layer is reached to obtain the final fuzzy evaluation vector. Conduct quantitative analysis on the final fuzzy evaluation vector, which can be converted into a comprehensive score or determine the optimal evaluation grade to determine the final evaluation grade of the evaluation object.

[0156] S4. Based on the optimal policy function, use the improved TD3 algorithm to process the state space model, action space model, and reward function to obtain the final optimized policy;

[0157] The improved TD3 algorithm adopts a dynamic policy noise mechanism and includes an Actor policy network and a Critic value network; among them, the Actor policy network includes a target policy network and an online policy network; the Critic value network includes a target value network and an online value network; the target value network includes a target1 value network and a target2 value network; the online value network includes an online1 value network and an online2 value network.

[0158] Because traditional TD3 adds Gaussian noise with a stable standard deviation when outputting actions, when encountering a large change in noise before and after, it will cause the instability of the agent model. Therefore, a dynamic policy noise mechanism is adopted, and its formula is:

[0159]

[0160] Among them, represents the noise standard deviation of the output action, 、 respectively represent the maximum noise standard deviation and the minimum noise standard deviation of the output action, represents the total number of training episodes, represents the current number of training episodes, 、 respectively represent the number of training episodes corresponding to the boundaries between the early stage and the middle stage and between the middle stage and the late stage of the agent model training phase, represents the current number of training episodes. Its purpose is to give the agent the maximum action noise in the early stage of training to enrich the state-action sample experience pool for training; in the middle stage of training, as the number of training times increases, the action noise gradually decreases; in the late stage of training, give the agent the minimum action noise to make the agent model gradually stable.

[0161] As Figure 2 shown, the training process of the improved TD3 algorithm in step S4 includes the following steps:

[0162] S4-1. Set the total number of training episodes and the total number of steps per episode;

[0163] S4-2. Initialize the environment, that is, randomly initialize the training satellite orbit parameters within the preset value range and obtain the current state ;

[0164] S4-3. Based on the current state , select an action through the online policy network ;

[0165] The formula for step S4-3 is:

[0166]

[0167] Where represents represents the policy noise of the current step, , represents the action 's value range, represents a Gaussian distribution with a mean of 0 and a standard deviation of ; represents that the current network parameter is and the online policy network under the current state ;

[0168] S4-4, execute this action , and synchronize the simulation scenario based on STK to obtain the corresponding reward and the next state , and update the constellation optimal policy;

[0169] S4-5, take the current state , the next state , the action and the reward as sample data and store them in the experience pool ;

[0170] The experience pool adopts the prioritized experience replay mechanism, that is:

[0171]

[0172] Take the absolute value of the TD-error value of the data as the sampling weight of this data, give a larger sampling weight to the samples with higher learning value, and achieve higher training efficiency. When the TD-error is larger, it means that the current value prediction is more inaccurate, and this data sample is more effective for the learning of the agent's value model. Where represents the output data of the online value network when the state is and the action is executed, represents the output data of the online value network when the state is and the action is executed.

[0173] S4-6, from the experience pool Randomly sample N sample data and select actions through the target policy network and calculate the loss function ;

[0174] The loss function of S4-6 The calculation formula is:

[0175]

[0176] Among them, represents the minimum value function, represents the network parameters of the target value network, represents the network parameters of the online value network, , respectively represent the th sample data, the state of the th sample data, represents the online value network, represents the th sample data's action, represents the serial number of the online value network;

[0177] S4-6. According to the loss function , update the parameters of the target value network and the online value network;

[0178] S4-7. Calculate the policy gradient and update the parameters of the online policy network;

[0179] The policy gradient of step S4-7 The calculation formula is:

[0180]

[0181] Among them, represents the value gradient output according to the current state when the current network parameters of the online value network are , action , represents the action gradient output according to the current state when the current network parameters of the online policy network are ; Output;

[0182] S4-8. Update the target policy network and the target value network in step S4-6 in a soft update manner. The corresponding formula is:

[0183]

[0184] Among them, , , respectively represent the network parameters of the target1 value network, the target2 value network, and the target policy network. represents the soft update coefficient. , , respectively represent the network parameters of the online1 value network, the online2 value network, and the online policy network. S4-9. Determine whether the current step reaches the total number of steps in a round; if so, the current round ends and proceeds to step S4-10; otherwise, the current step is incremented by 1 and returns to step S4-3; among them, the initial value of the step is 0.

[0185] S4-10. Repeat steps S4-1 to S4-9 until the total number of training rounds is reached.

[0186] S5. Execute the final optimization strategy through the constellation to be observed to complete the optimization of the observed constellation.

[0187] In one embodiment of the present invention, simulation training is performed based on the constellation optimization design algorithm improved by TD3. The weight set of the third-level indicators is obtained by calculating through the analytic hierarchy process, and the analytic hierarchy process judgment matrix of the weight set of the third-level indicators is constructed. The obtained values of the weight set of the third-level evaluation indicators are shown in Table 1.

[0188] Table 1

[0189]

[0190] The main hyperparameter settings for the training of each policy network and value network based on the TD3 algorithm are shown in Table 2.

[0191] Table 2

[0192]

[0193] Such as Figure 3 and Figure 4As shown in the figure, the reward curves of the general constellation configuration and the Walker constellation configuration are analyzed. Compared with the simplified general constellation configuration, for the same satellite scale, only in the design case with a relatively small constellation scale, the state and action dimensions are basically the same. As the constellation scale expands, the optimization parameters required for the simplified general constellation configuration increase in multiples, that is, the state and action dimensions of the simplified general constellation configuration increase in multiples, which increases the difficulty of network learning. The direct impact is that the training process converges more slowly, while the state and action dimensions of the Walker constellation configuration basically do not change. The simplified general constellation configuration requires a longer convergence time. In addition, the training results for both the simplified general constellation and the Walker constellation configurations can achieve convergence, further proving that the improved TD3 algorithm for constellation optimization in this method is effective.

[0194] Due to the inconsistency of the state and action dimensions of the two constellation configurations and the different value ranges of the optimization variables, the optimization times reward in the reward function is not comparable, and the rewards obtained for the environmental design are inconsistent. Therefore, it is meaningless to compare the magnitudes of the reward values here. The constellation can be optimized through the decision model obtained by training, and the optimization results are as Figure 5 shown.

[0195] The general constellation configuration is simulated, and the simulation results of 12,000 - time training are compared with those of the particle swarm algorithm. Among them, the particle swarm is set to two cases of 20 and 100 respectively for comparison.

[0196] As Figure 6 、 Figure 7 、 Figure 8 、 Figure 9 、 Figure 10As shown, the observation efficiency curve of this method is smoother than the two observation efficiency curves of the particle swarm algorithm. Moreover, the observation performance value of this method is 71.04, which is 7.81 higher than the observation performance value of 20 for the particle swarm algorithm and 2.1 higher than the observation performance value of 100 for the particle swarm algorithm. The processing time of this method is only 69.39 s, while the processing times of the particle swarm algorithm with settings of 20 and 100 are 384.2 s and 1890.64 s respectively, both far exceeding the processing time of this method. This further shows that for the constellation optimization method based on the particle swarm algorithm, as the scale of the particle swarm increases, the observation performance of the optimized constellation improves, but the required optimization time increases significantly. Compared with the constellation optimization method based on the particle swarm algorithm, the constellation optimized by the constellation optimization algorithm proposed in this chapter has basically the same observation performance and shows an obvious speed advantage. This is because, compared with the particle swarm algorithm, the constellation optimization method based on the improved TD3 algorithm obtains a decision mapping through training, while the constellation optimization method based on the particle swarm algorithm is one-time and requires iterative search from scratch each time it is used, which proves the superiority of the algorithm in this chapter in terms of efficiency.

[0197] In summary, the present invention improves the disadvantages of the traditional TD3, such as low learning and training efficiency and instability of the optimization model. By introducing prioritized experience replay and dynamic policy noise, the training efficiency and stability of the constellation optimization design decision model are improved, and the constellation's ground observation performance is enhanced.

Claims

1. An observation constellation optimization method based on deep reinforcement learning, characterized in that: It includes the following steps: S1. Determine the constellation configuration of the constellation to be observed and obtain its satellite orbit parameters; S2. Based on the satellite orbit parameters, construct a constellation optimization strategy using the reinforcement learning algorithm; S3. Establish a state space model, an action space model and design a reward function; S4. Based on the optimal policy function, use the improved TD3 algorithm to process the state space model, the action space model and the reward function to obtain the final optimization strategy; S5. Execute the final optimization strategy through the constellation to be observed to complete the optimization of the observation constellation; The step S1 includes the following steps: S1-1. Determine the constellation configuration of the constellation to be observed; wherein, the constellation configuration includes a general constellation configuration and a Walker constellation configuration; S1-2. Judge whether the constellation configuration of the constellation to be observed is a general constellation configuration; if so, enter step S1-3; otherwise, determine that the constellation configuration of the constellation to be observed is a Walker constellation configuration and enter step S1-4; S1-3. Take the 6 first orbit parameters corresponding to each satellite in the general constellation configuration as the satellite orbit parameters and enter step S2; wherein, the 6 first orbit parameters are respectively the semi-major axis of the satellite orbit, the orbital eccentricity, the orbital inclination, the argument of perigee, the right ascension of the ascending node, and the true anomaly; S1-4. Determine the total number of satellites in the Walker constellation configuration and the orbital elements, phase factors, number of orbital planes, semi-major axis of the orbit, and inclination of the orbital plane of each satellite; S1-5. Take the first satellite in the Walker constellation configuration as the seed satellite and obtain the corresponding right ascension of the ascending node and the true anomaly at the initial moment; S1-6. According to the formula: Obtain the right ascension of the ascending node of the remaining satellites and the true anomaly ; where represents the right ascension of the ascending node of the seed satellite, represents the number of orbital planes, represents pi, indicates that the satellite is located in the th orbital plane, represents the th satellite on the orbital plane, represents the true anomaly of the seed satellite at the initial moment; S1-7. Take the total number of satellites in the Walker constellation configuration , the phase factor, the number of orbital planes, the semi-major axis of the orbit, the inclination of the orbital plane, and the right ascension of the ascending node and the true anomaly at the initial moment of all satellites as satellite orbit parameters, and proceed to step S2; Among them, the experience pool used in the training process of the improved TD3 algorithm adopts a prioritized experience replay mechanism; the improved TD3 algorithm adopts a dynamic policy noise mechanism, and the expression of the dynamic policy noise mechanism is: Among them, represents the noise standard deviation of the output action, , respectively represent the maximum noise standard deviation and the minimum noise standard deviation of the output action, represents the total number of training rounds, represents the current number of training rounds, , respectively represent the number of training rounds corresponding to the boundaries between the early stage and the middle stage and between the middle stage and the late stage of the agent model training phase, represents the current number of training rounds; the agent model is the constellation model to be observed.

2. The method for optimizing an observation constellation based on deep reinforcement learning according to claim 1, wherein: The step S2 includes the following steps: S2-1. Take the constellation to be observed as an agent and determine the optimization task of the constellation to be observed, and construct a corresponding MDP model based on the Markov decision process, and the corresponding formula is: Among them, represents the state set, including satellite orbit parameters, represents the action set, represents the state transition function, represents the immediate reward and punishment, represents the discount factor; S2-2. Calculate the cumulative reward corresponding to all decisions executed by the MDP model; S2-3. Calculate the optimal state value function and the optimal state-action value function ; where represents an action in the action set ; S2-4. Determine the constellation optimization strategy based on the optimal state value function and the optimal state-action value function .

3. The method for optimizing an observation constellation based on deep reinforcement learning according to claim 2, characterized in that: The calculation formula of the cumulative reward is: Among them, , respectively represent the cumulative rewards corresponding to the time , time ; represents the summation function, , , , respectively represent the immediate rewards and punishments corresponding to the time , time , time , time ; The optimal state value function has the following calculation formula: Among them, represents the maximum value function, and represent the state, represents the state corresponding optimal state value function, represents that when the agent is in state and executes action the probability of transferring to state ; represents the immediate reward and punishment when the agent executes action in state ; The calculation formula of the optimal state-action value function is: Among them, represents the best state-action value function when the agent executes the action in the state .

4. The method for optimizing an observation constellation based on deep reinforcement learning according to claim 2, characterized in that: The step S3 includes the following steps: S3-1. Judge whether the constellation configuration of the constellation to be observed is a general constellation configuration; if so, enter step S3-2; otherwise, enter step S3-4; S3-2. Set the semi-major axis of the satellite orbit in the general constellation configuration to 8000 km, set the satellite orbit as a circular orbit, the eccentricity as 0, and the argument of perigee as 0. Take the orbital inclination, right ascension of the ascending node, true anomaly, and position parameters as the state, that is, the state space model. The corresponding expression is as follows: Among them, and respectively represent the orbital plane inclinations of the first satellite and the second satellite of the constellation to be observed; and , and respectively represent the right ascensions of the ascending nodes of the first satellite and the second satellite of the constellation to be observed; and , and and respectively represent the true anomaly angles of the first satellite , the second satellite , and the th satellite of the constellation to be observed; and and , and and respectively represent the longitude, latitude and altitude values of the constellation to be observed in the GPS84 coordinate system; S3-3. Establish the action space model of the general constellation configuration, i.e., the state space model, and proceed to step S3-6; among them, the state space model of the general constellation configuration The corresponding expression is: Among them, and respectively represent the actions of the first satellite and the second satellite in the orbital inclination and . and respectively represent the actions of the first satellite and the second satellite in the right ascension of the ascending node and . , , respectively represent the actions of the first satellite , the second satellite , and the th satellite in the true anomaly , , . S3-4. Set the semi-major axis of the satellite orbit in the Walker constellation configuration to a fixed value, and take the number of orbital planes, the phase factor, the orbital inclination, the right ascension of the ascending node, the true anomaly and the position parameter as the state, that is, the state space model, and the corresponding expression is: Among them, , , , , respectively represent the number of Walker constellation orbital planes, phase factor, orbital inclination, right ascension of the ascending node, and the state at the true anomaly; S3-5. Establish the action space model, i.e., the state space model, of the Walker constellation configuration, and proceed to step S3-6; among them, the state space model of the Walker constellation configuration The corresponding expression is: Among them, , , , , respectively represent the number of Walker constellation orbital planes, phase factors, orbital plane inclination angles, right ascensions of ascending nodes, and actions at true anomalies; S3-6. According to the formula: Construct the reward function ; where , respectively represent the target observation performance reward and the optimization times reward, , respectively represent the coefficients, and .

5. The method for optimizing an observation constellation based on deep reinforcement learning according to claim 4, wherein: The improved TD3 algorithm adopts a dynamic policy noise mechanism and includes an Actor policy network and a Critic value network. Among them, the Actor policy network includes a target policy network and an online policy network; the Critic value network includes a target value network and an online value network; the target value network includes a target1 value network and a target2 value network; the online value network includes an online1 value network and an online2 value network.

6. The method for optimizing an observation constellation based on deep reinforcement learning according to claim 5, wherein: The training process of the improved TD3 algorithm in step S4 includes the following steps: S4-1. Set the total number of training episodes and the total number of steps per episode. S4-2. Initialize the environment, that is, randomly initialize the training satellite orbit parameters within a preset value range and obtain the current state ; S4-3. Based on the current state , select an action through the online policy network ; S4-4. Execute this action , and synchronize the STK-based simulation scenario to obtain the corresponding reward and the next state , and update the constellation optimal strategy; S4-5. Store the current state , the next state , the action and the reward as sample data into the experience pool ; S4-6. Randomly sample N sample data from the experience pool and select actions through the target policy network , calculate the loss function ; S4-7. Update the parameters of the target value network and the online value network according to the loss function ; S4-8. Calculate the policy gradient and update the parameters of the online policy network. S4-9. Update the target policy network and the target value network in step S4-7 in a soft update manner. S4-10. Determine whether the current number of steps has reached the total number of steps per episode. If so, the current episode ends and step S4-11 is entered; otherwise, the current number of steps is incremented by 1 and step S4-3 is returned. Among them, the initial value of the number of steps is 0. S4-11. Repeat steps S4-1 to S4-10 until the total number of training episodes is reached.

7. The method for optimizing an observation constellation based on deep reinforcement learning according to claim 6, wherein: The formula for step S4-3 is: Among them, represents the online policy network At the current network parameter when according to the current state the output action, represents the policy noise of the current step, and represents the action value range, represents a Gaussian distribution with a mean of 0 and a standard deviation of ; represents the current network parameter and at the current state the online policy network ; The loss function of step S4-7 The calculation formula is as follows: Among them, represents the minimum value function, represents the network parameters of the target value network, represents the network parameters of the online value network, 、 respectively represent the th sample data and the state of the th sample data, represents the online value network, represents the action of the th sample data, represents the serial number of the online value network; The policy gradient of step S4-8 has the following calculation formula: Among them, represents the value gradient output according to the current state when the current network parameters are in the online value network and the action ; represents the action gradient output according to the current state when the current network parameters are in the online policy network .

Citation Information

Patent Citations

  • Star group collaborative task planning method based on mixed expert experience playback

    CN117068393A