An Interplanetary Orbit Transfer Method Based on Hidden State and Reinforcement Learning
Through the hidden state and reinforcement learning methods, the sequence hidden variable model and reinforcement learning controller predict the next state is solved, and the problem of agents not understanding uncertainty and reward sparseness in the prior art is solved, and a more efficient interplanetary orbit transfer design is achieved.
Patent Information
- Application Number
- CN202510340687.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-21
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2045-03-21
AI Technical Summary
The interplanetary orbit transfer method based on reinforcement learning in the prior art relies on continuous interactive learning with the environment, and the agent fails to understand the uncertainty and the reward structure is sparse, resulting in learning difficulties.
Using a method based on hidden state and reinforcement learning, the uncertain environment is represented by the sequence hidden variable model, and a reinforcement learning controller is constructed in combination with the critic network and the actor network, predict the next state and incorporate it into the reward structure to improve learning efficiency.
The training of agents in uncertain environments is accelerated, the processing ability of uncertainty is improved, the robustness and learning efficiency of orbital transfer design is enhanced, fuel consumption is reduced, and short-sighted behavior is avoided.
Smart Images

Figure CN119861572B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of spacecraft control, and particularly to an interplanetary orbit transfer method based on hidden states and reinforcement learning. Background Art
[0002] In recent years, small and micro spacecraft equipped with low-thrust electric thrusters have gradually become a research hotspot for deep space exploration due to their advantages such as low fuel consumption, high mission flexibility, and strong scalability. The design of low-thrust interplanetary orbit transfer is one of the basic guarantees for the success of small and micro spacecraft in deep space exploration missions. The design of low-thrust interplanetary orbit transfer aims to use a low-thrust propulsion system for long-term continuous propulsion to achieve efficient orbit transfer.
[0003] In practical applications, due to the limited on-board resources of small and micro spacecraft, there is a high requirement for lightweight computing. At the same time, the spacecraft will be affected by various external interferences and internal uncertainties in interplanetary space, and these factors will affect the accuracy and reliability of orbit transfer.
[0004] For the space environment with multi-source uncertainties, a robust design method for low-thrust interplanetary trajectories based on meta-reinforcement learning has been proposed. By continuously interacting between the agent and the environment, the adaptability and robustness to complex environments and uncertainties are achieved, and good real-time performance is realized by virtue of the characteristics of pre-training with artificial intelligence methods. However, relying only on the characteristics of continuous interaction between reinforcement learning and the environment for learning, the agent does not understand uncertainties. In addition, the key performance indicators of traditional methods include terminal fuel consumption and the errors of terminal position and velocity, which makes the effect evaluation only carried out at the end of the task, resulting in an extremely sparse reward structure.
[0005] Therefore, the existing technology still needs to be improved and developed. Summary of the Invention
[0006] The technical problem to be solved by the present invention is to provide an interplanetary orbit transfer method based on hidden states and reinforcement learning in view of the above-mentioned defects of the existing technology, aiming to solve the problems in the existing robust design method for interplanetary trajectories based on reinforcement learning, where the agent does not understand uncertainties by relying only on the characteristics of continuous interaction between reinforcement learning and the environment, and the reward structure is sparse, resulting in difficult learning for the agent.
[0007] The technical solution adopted by the present invention to solve the problem is as follows:
[0008] An embodiment of the present invention provides an interplanetary orbit transfer method based on hidden states and reinforcement learning, and the method includes:
[0009] Obtain the observation data of the spacecraft; wherein, the observation data is determined based on state data and observation uncertainty; the state data is calculated based on a state transition matrix constructed for a preset orbit transfer mission;
[0010] Input the observation data into a preset sequential latent variable model to obtain latent variables;
[0011] Input the latent variables and the historical trajectory record of the spacecraft into a reinforcement learning controller to obtain the execution actions of the spacecraft;
[0012] Predict the next state data of the spacecraft according to the observation data and the execution actions, calculate a reward value according to the state data and the next state data, and enable the reinforcement learning controller to update the orbit transfer strategy based on the reward value;
[0013] Wherein, the reinforcement learning controller includes:
[0014] A critic network, which is used to estimate the state value according to the input latent variables and its own value estimation function, and update the orbit transfer strategy according to the state value;
[0015] An actor network, which is used to calculate the execution actions of the spacecraft according to the historical trajectory record and the orbit transfer strategy.
[0016] Advantages of the present invention: In the embodiments of the present invention, by establishing a sequential latent variable model, the uncertain environment can be explicitly represented and learned, and the hidden information hidden under the observations of the uncertain environment can be extracted. And a reinforcement learning controller is composed of an actor network and a critic network. A reinforcement learning algorithm framework is constructed by using the sequential latent variable model and the reinforcement learning controller, thereby accelerating the training of the agent in the uncertain environment and improving the agent's ability to handle uncertainty. In addition, the embodiments of the present invention also predict the next state based on the current observation and the expected operation, and incorporate the quality of the predicted next state into the reward structure, so that the immediate reward can capture the effectiveness of the current and subsequent strategies at the same time, thereby improving the learning efficiency of the algorithm. Description of the Drawings
[0017] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments recorded in the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0018] Figure 1 It is a schematic flowchart of an interplanetary orbit transfer method based on hidden state and reinforcement learning provided by an embodiment of the present invention.
[0019] Figure 2 It is the structure diagram of the sequence latent variable model provided by the embodiment of the present invention.
[0020] Figure 3 It is the schematic diagram of the framework of stochastic latent state proximal policy optimization provided by the embodiment of the present invention.
[0021] Figure 4 It is the comparison diagram of the reward convergence effect before and after the automatic adjustment of weights provided by the embodiment of the present invention.
[0022] Figure 5 It is the size distribution diagram of the velocity increment applied by the pulse thrust provided by the embodiment of the present invention.
[0023] Figure 6 It is the orbit transfer result diagram generated by Monte Carlo simulation provided by the embodiment of the present invention. Detailed implementation manners
[0024] The present invention discloses an interplanetary orbit transfer method based on latent state and reinforcement learning. To make the objectives, technical solutions and effects of the present invention clearer and more definite, the following further describes the present invention in detail with reference to the accompanying drawings and by way of examples. It should be understood that the specific examples described herein are only used to explain the present invention and are not used to limit the present invention.
[0025] Those skilled in the art of the present technology can understand that unless specifically stated otherwise, the singular forms "a", "an", "the" and "said" used herein may also include the plural forms. It should be further understood that the term "including" used in the specification of the present invention means the presence of the described features, integers, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or their groups. It should be understood that when we say that an element is "connected" or "coupled" to another element, it can be directly connected or coupled to other elements, or there may also be intermediate elements. In addition, the "connection" or "coupling" used herein may include wireless connection or wireless coupling. The phrase "and / or" used herein includes all or any unit and all combinations of one or more related listed items.
[0026] Those skilled in the art of the present technology can understand that unless otherwise defined, all terms (including technical terms and scientific terms) used herein have the same meaning as the general understanding of those of ordinary skill in the art to which the present invention belongs. It should also be understood that terms such as those defined in a general dictionary should be understood to have a meaning consistent with the meaning in the context of the prior art, and will not be interpreted with an idealized or overly formal meaning unless specifically defined as here.
[0027] In view of the above defects of the prior art, the present invention provides an interplanetary orbit transfer method based on hidden states and reinforcement learning, as Figure 1 shown. The method specifically includes the following steps:
[0028] Step S100: Obtain the observation data of the spacecraft; wherein, the observation data is determined based on state data and observation uncertainty; the state data is calculated based on a state transition matrix constructed for a preset orbit transfer task;
[0029] Step S200: Input the observation data into a preset sequential latent variable model to obtain latent variables;
[0030] Step S300: Input the latent variables and the historical trajectory record of the spacecraft into a reinforcement learning controller to obtain the execution actions of the spacecraft;
[0031] Step S400: Predict the next state data of the spacecraft according to the observation data and the execution actions, calculate a reward value according to the state data and the next state data, and enable the reinforcement learning controller to update the orbit transfer strategy based on the reward value;
[0032] Wherein, the reinforcement learning controller includes:
[0033] A critic network, which is used to estimate the state value according to the input latent variables and its own value estimation function, and update the orbit transfer strategy according to the state value;
[0034] An actor network, which is used to calculate the execution actions of the spacecraft according to the historical trajectory record and the orbit transfer strategy.
[0035] Generally speaking, considering that the environment has various uncertainties, the system cannot directly obtain the state and observation errors caused by these uncertainties. In this embodiment, the uncertain environment is regarded as a partially observable Markov decision process. By integrating a sequential latent variable model with a proximal policy optimization algorithm, a reinforcement learning framework that can handle uncertainties and partial observability is proposed (which can be called a stochastic hidden state proximal policy optimization algorithm framework). The algorithm process of this embodiment is mainly divided into three steps. The first step is problem modeling, that is, constructing a state transition matrix based on a preset orbit transfer task and constructing an expression for observation data. The second step is to construct a reinforcement learning algorithm framework that can handle uncertainties and partial observability, that is, first using a sequential latent variable model (such as Figure 2Perform representation learning on the orbit transfer model containing uncertainties, autonomously extract the hidden state information that can characterize the system state, and predict the future state of the system to improve the ability to handle uncertainties. Secondly, combine the sequential latent variable model and the proximal policy optimization technique to establish a stochastic hidden state proximal policy optimization reinforcement learning framework. Use the learned latent variable information to replace the uncertain state information for the design of the orbit transfer policy, and reduce the impact of system uncertainties on the overall control. The third step is reward design. The sequential latent variable model in this embodiment can not only facilitate the extraction of hidden information in the uncertain environment, but also predict the next state based on the current observation and expected operation. Incorporate the data related to predicting the future state into the calculation of the reward value for each step, which can alleviate the sparsity of the reward structure, and further improve the future expected effect, data efficiency, and model learning speed of the algorithm.
[0036] Specifically, the reinforcement learning algorithm framework is as Figure 3 shown, and it includes two parts: a sequential latent variable model and a reinforcement learning controller composed of an actor network and a critic network. That is, the intelligent agent is jointly composed of the sequential latent variable model and the reinforcement learning controller. Explicitly perform representation learning on the uncertain environment through the sequential latent variable model, and extract the hidden information hidden under the observations of the uncertain environment, so as to accelerate the training of the intelligent agent in the uncertain environment and improve the ability of the intelligent agent to handle uncertainties. In this embodiment, the latent variable learned by the sequential latent variable model is used as the input of the critic network, and the current and historical environment trajectories are used as the input of the actor network, so as to solve the problem of insufficient decision-making information that may be caused by uncertainties. In addition, different from the traditional reward design that is only designed based on the current state of the environment and the actions taken, which only reflects the quality of the current action. In the training process of this embodiment, the sequential latent variable model and the current policy are comprehensively considered and used to predict the next state, and the quality of the predicted state is quantified and incorporated into the reward structure, so that the immediate reward can capture the effectiveness of the current and subsequent policies at the same time, and further improve the learning efficiency of the algorithm.
[0037] In one implementation, the state transition matrix is:
[0038] , formula (1);
[0039] where, and respectively represent the positions of the spacecraft at the k-th time step and the (k + 1)-th time step; and respectively represent the velocities of the spacecraft at the k-th time step and the (k + 1)-th time step; and respectively represent the masses of the spacecraft at the k-th time step and the (k + 1)-th time step; and denotes the Lagrangian coefficient at the k-th time step; and denotes the derivative of the Lagrangian coefficient at the k-th time step; the vector is the output of the reinforcement learning controller, representing the velocity change obtained by impulse inference ; denotes the equivalent exhaust velocity of the thruster; denotes the position deviation caused by state uncertainty at the (k + 1)-th time step; denotes the velocity deviation caused by state uncertainty at the (k + 1)-th time step;
[0040] The observed data is in vector form, and the calculation formula for the observation vector is:
[0041] , formula (2);
[0042] where denotes the observation vector; denotes time, , denotes the total number of time steps; denotes the transpose of a matrix or vector; denotes the position deviation caused by observation uncertainty at the k-th time step; denotes the velocity deviation caused by observation uncertainty at the k-th time step.
[0043] Specifically, the derivation process of the above calculation formulas for the state transition matrix and the observation vector is as follows:
[0044] First, perform problem modeling: In a low-thrust interplanetary rendezvous mission, a spacecraft equipped with a low-thrust electric propulsion device departs from a given initial position and, through low-thrust orbital maneuvers, transfers to the target orbit within a specified time and reaches the same velocity as the target planet. Assuming that the spacecraft is only affected by the solar gravity, the spacecraft model can be approximated as a two-body model. Considering the electric thrust and space perturbations acting on the spacecraft, its initial kinematic and dynamic models can be described as:
[0045] , formula (3);
[0046] , formula (4);
[0047] , formula (5);
[0048] where and respectively denote the position vector and velocity vector pointing from the center of the celestial body's mass to the spacecraft; denotes the distance from the central celestial body to the spacecraft; denotes the standard gravitational parameter, and denote the mass of the central celestial body and the mass of the spacecraft respectively, is the gravitational constant; is the thrust amplitude of the electric propulsion system; is the direction vector of the thrust; is the perturbation acceleration suffered by the spacecraft; is the gravitational acceleration at sea level; is the specific impulse of the electric propulsion system.
[0049] Discretize the low-thrust orbit and approximate the low-thrust orbit as trajectories connected by impulses Due to the maximum thrust limit of the low-thrust thruster, the impulse at the time step has a maximum amplitude limit.
[0050] The maximum impulse value and the impulse value at the time step are defined as follows:
[0051] , Equation (6);
[0052] , Equation (7);
[0053] In the formula, denotes the maximum thrust amplitude that the spacecraft engine can provide; denotes the total transfer time, denotes the mass of the spacecraft at time .
[0054] The initial kinematic and dynamic models can be further rewritten in the following discrete Markov decision process form:
[0055] , Equation (8);
[0056] , Equation (9);
[0057] Among them, is a stochastic discrete dynamic model describing the environmental state transition; is a function representing the policy mapping; and respectively denote and the system states at the step, and these states include the position of the spacecraft velocity and mass ; vector is the output of the reinforcement learning controller, representing the velocity change obtained by the impulse thrust ; represents the observation vector of the spacecraft, which contains all the information in and the current time , .
[0058] Based on the initial kinematic and dynamic models, establish the initial state transition matrix of the Markov decision process:
[0059] , formula (10);
[0060] The action is defined as:
[0061] , formula (11);
[0062] where and represent the Lagrange coefficients at the th time step; represents the mass change formula, given by the Tsiolkovsky equation; is the equivalent exhaust velocity of the thruster.
[0063] In addition, to accurately describe the state transition process, the present invention considers two different types of uncertainties:
[0064] 1. State uncertainty: Uncertainty caused by external disturbances and unmodeled dynamics;
[0065] 2. Observation uncertainty: Uncertainty caused by sensor errors and communication failures.
[0066] Time (k = 0, 1, 2,..., ) The state uncertainty and observation uncertainty at time are represented as additive Gaussian noise in position and velocity. Specifically:
[0067] , formula (12);
[0068] , formula (13);
[0069] where represents the covariance matrix of the state uncertainty , represents the position deviation caused by the state uncertainty at the kth time step, denotes the velocity deviation caused by state uncertainty at the k-th time step; denotes the observation uncertainty of the covariance matrix, denotes the position deviation caused by observation uncertainty at the k-th time step, denotes the velocity deviation caused by observation uncertainty at the k-th time step; , , , denotes the standard deviations of position and velocity in state uncertainty and observation uncertainty; is the identity matrix of dimension , while is the zero matrix of dimension .
[0070] Therefore, the previously established initial state transition matrix, i.e., formula (10) can be rewritten as formula (1). And the observation vector can be expanded as formula (2).
[0071] In one implementation, the sequential latent variable model consists of an action model, latent variables, an observation model, and latent variable dynamics, and the latent variable dynamics is used to establish the connection between the observation model, the action model, and the latent variables; the objective function of the sequential latent variable model is established based on the equivalence principle of maximizing the observation data probability and maximizing the evidence lower bound.
[0072] Specifically, as Figure 2 shown, the structure of the sequential latent variable model in this embodiment mainly includes four parts: action, latent variables, an observation model, and latent variable dynamics. Figure 2 in is the latent variable at the k-th time step, which summarizes the current and historical information, including position, velocity, mass, time, and the state transition probability to the next state. denotes the observation model, which includes the sum of the state information and the deviation caused by uncertainty. To highlight the ability of the sequential latent variable model to represent an environment with uncertainty, Figure 2 the deviation caused by uncertainty is separately extracted in is the time step Actions taken at that time. Latent variable dynamics establish the connection between the observation model, the action model, and the latent variables. To ensure that the established latent variable model accurately represents the environment, the probability of each observed data from the training set must be maximized throughout the generation process. Since it is difficult to directly calculate this probability due to the marginalization of the latent variables, in this embodiment, the evidence lower bound of this probability is derived in a manner similar to amortized variational inference. Maximizing the probability of the observed data is equivalent to maximizing the evidence lower bound, thereby obtaining the objective function of the sequential latent variable model.
[0073] In one implementation, the latent variable dynamics include a generative model and an inference model;
[0074] The generative model is used to sequentially generate the next latent variable and the prior distribution of each latent variable according to the previous latent variable information;
[0075] The inference model is used to calculate the posterior distribution of the current latent variable based on the subsequent latent variable information and the observed data.
[0076] Specifically, as Figure 2 shown, where the solid arrows represent the generative model and the dashed arrows represent the inference model. The generative model sequentially generates the next latent variable according to the previous information, including the prior distribution of each latent variable. Here, it includes the initial prior distribution and , the prior distribution and ( ), and the decoder . In the generative model, the initial distribution is defined as a multivariate standard normal distribution, that is, , while the other distributions are parameterized by a neural network with parameters . On the contrary, the inference model uses the subsequent latent variable information and the observed information to infer the probability of the current latent variable, that is, the posterior distribution of each latent variable. The inference model includes the initial posterior distribution and , the posterior distribution and . To simplify the subsequent steps in the inference model, it is assumed that 's posterior distribution is the same as its prior distribution. 's posterior distribution is represented by and is parameterized by a neural network with parameters .
[0077] Furthermore, the probability of each observed data from the training set is maximized throughout the generation process, . The evidence lower bound of this probability is derived:
[0078] , formula (14);
[0079] wherein, is the sequence length of the single-sequence latent variable model; represents the relative entropy of two distributions.
[0080] Furthermore, maximizing the probability of the observed data is equivalent to maximizing the evidence lower bound, so the objective function of the sequence latent variable model is:
[0081] , formula (15);
[0082] wherein, represents the objective function of the sequence latent variable model; is the sequence length of the single-sequence latent variable model; represents the relative entropy of two distributions; represents the probability of each observed data; represents the evidence lower bound; represents the observation vector, is the latent variable at the k-th time step, and respectively represent the first-layer latent variable and the second-layer latent variable at the k-th time step; represents the execution action at the k-th time step; represents the function with respect to the trajectory sampled from the probability distribution q
[0083] In one implementation, the overall objective function of the reinforcement learning controller consists of the objective formula of the actor network, the mean squared value function error term calculated based on the value function, and the entropy term calculated based on the historical trajectory record.
[0084] Specifically, as Figure 3 shown, the latent variable obtained from the sequence latent variable model is used as the input of the evaluation network to estimate the value function of the state, denoted as . At the same time, the historical observation and decision sequence are directly used as the input of the actor network to output the action, denoted as , wherein is the hyperparameter that controls the length of the sequence latent variable model. Corresponding modifications are made to the original formula of the proximal policy optimization algorithm to obtain the objective formula of the actor network. Furthermore, the overall objective function of the reinforcement learning controller (i.e., the proximal policy optimization algorithm part) is composed of the mean squared value function error term and the entropy term.
[0085] In one implementation, the overall objective function of the reinforcement learning controller is:
[0086] , formula (16);
[0087] wherein, represents the overall objective function of the reinforcement learning controller; represents the objective formula of the actor network; represents the mean square value function error term; represents the entropy term; represents the trainable parameters of the neural network corresponding to the reinforcement learning controller; , represents the hyperparameters that control the relative importance of the three objectives;
[0088] The objective formula of the actor network is:
[0089] , formula (17);
[0090] wherein, represents the objective formula of the actor network; represents along the policy the expectation of the trajectory sampled; represents the total number of time steps; represents the importance weight, represents the old policy used when sampling data, represents the execution of an action, represents the system state; represents the value range limited to function, represents the importance weight at the represents the hyperparameter that controls the pruning range; represents the generalized advantage estimation;
[0091] The calculation formula of the mean square value function error term is:
[0092] , formula (18);
[0093] wherein, represents the mean square value function error term; represents the latent variable; represents the self-value estimation function established by the neural network with parameters ; represents the power of the reward discount factor; represents the environmental reward at the
[0094] The calculation formula of the entropy term is:
[0095] , formula (19);
[0096] wherein, represents the entropy term; represents the th time step and the trajectory of the previous steps, that is, .
[0097] In one implementation, the generalized advantage estimate (generalized advantage estimate, GAE) is used as an unbiased estimate of the advantage function , and the calculation formula is as follows:
[0098] , formula (20);
[0099] wherein, is an adjustable hyperparameter in GAE; is the reward discount factor.
[0100] In one implementation, calculating the reward value according to the state data and the next state data includes:
[0101] Obtain the reference trajectory of the spacecraft, wherein the reference trajectory is the optimal trajectory obtained based on a deterministic environment;
[0102] Calculate a reward term according to the next state data and the reference trajectory;
[0103] Obtain the target state data corresponding to the target planet, calculate the final position error and the final velocity error of the spacecraft according to the state data and the target state data, and calculate a penalty term according to the final position error and the final velocity error respectively;
[0104] Calculate the single-step fuel consumption of the spacecraft according to the state data, and calculate a penalty term according to the single-step fuel consumption;
[0105] Calculate a penalty term according to the state data and a preset maximum thrust threshold;
[0106] Calculate a penalty term according to the state data and a preset no-entry area;
[0107] Calculate the position deviation and the velocity deviation according to the state data and the reference trajectory, and calculate a penalty term according to the position deviation and the velocity deviation respectively;
[0108] Calculate the reward value according to the reward term and all the penalty terms.
[0109] Specifically, in this embodiment, a reward function is used to calculate the reward value. The reward value, as the environmental feedback, can be used to update the orbit transfer strategy, specifically, to update the network parameters of the critic network. The reward function is jointly composed of multiple data items, and the types of data items are divided into reward items and penalty items. The reward function specifically includes the following data items:
[0110] Reward item for the predicted next state: Through the sequential latent variable model, the next state is predicted based on the current observation and expected operation. A reward item is generated from the predicted next state and the reference trajectory to alleviate the sparsity of the reward structure. The value of the reference trajectory is determined based on the trajectory generated by the agent trained in the deterministic environment.
[0111] Penalty items for the final position error and the final velocity error, where the final position error and the final velocity error represent the differences between the final position and velocity of the spacecraft and the final position and velocity of the target planet. Incorporating the terminal position and velocity errors as penalties into the calculation process of the reward value can effectively reduce the final position error and the final velocity error.
[0112] Penalty item for single-step fuel consumption: Small microspacecraft used for deep space exploration usually lack the ability to carry a large amount of fuel or refuel midway through the mission. Therefore, in the design of the orbit transfer mission, it is crucial to minimize fuel consumption while ensuring that the terminal error remains as small as possible. To this end, this embodiment introduces a penalty item related to fuel consumption, which penalizes the fuel consumption at each time step to make the spacecraft conserve fuel as much as possible.
[0113] Penalty item for the maximum thrust threshold: This item reflects the constraint related to the maximum thrust limit. Spacecraft equipped with low-thrust propulsion systems are subject to the maximum thrust limit. If the decision results in a thrust exceeding this maximum limit, it is considered invalid, and the decision is executed with the maximum allowable thrust. To reduce the occurrence of such events, this embodiment introduces a penalty item for decisions that exceed the maximum thrust threshold.
[0114] Penalty item for the no-go area: The no-go area refers to a specific area where entry is prohibited. This item reflects the penalty for entering a specific area. To improve the efficiency of spacecraft exploration and avoid excessive exploration in invalid areas, this embodiment establishes a no-go area for the spacecraft that significantly deviates from the optimal trajectory. A large constant penalty is imposed on the spacecraft entering this area to limit excessive exploration in areas that deviate significantly from the orbit.
[0115] Penalty terms for position deviation and velocity deviation: In this embodiment, a reference trajectory will be used to accelerate the exploration of the agent. To improve the training speed of the agent in an uncertain environment, the optimal trajectory obtained in a deterministic environment is used as the reference trajectory in this embodiment. And a penalty is generated based on the deviation between the current position of the spacecraft and the reference trajectory, so as to guide the exploration of the agent. In practical applications, in order to avoid over-reliance on the reference trajectory resulting in sub-optimal solutions, the weight assigned to the penalty term established based on the reference trajectory constraint should not be too large.
[0116] In one implementation, calculating the reward value according to the reward term and all the penalty terms includes:
[0117] Calculating the reward value according to the weighted combination of the reward term and all the penalty terms;
[0118] Among them, the adjustment method of the weights corresponding to the reward term and each of the penalty terms includes:
[0119] Establishing a reward function according to the reward term and all the penalty terms, wherein the penalty terms established based on the final position error and the final velocity error include constraint values that change with the training process;
[0120] Training respectively using two sets of reward functions with different weight values;
[0121] Determining the training potential corresponding to the reward term and each of the penalty terms according to the relative magnitudes of the reward term and each of the penalty terms in the training results, and the difference between the training results of the two sets of reward functions;
[0122] Adjusting the weights according to the training potential of the reward term and each of the penalty terms.
[0123] Specifically, each data item of the reward function in this embodiment is weighted, and the weight requirements for each data item are different to reflect the relative importance of each data item. If the weight setting is not correctly calibrated, it may cause the target with a lower weight to be masked by the target with a higher weight. This imbalance will slow down the optimization of the target with a lower priority and have an adverse impact on the overall training efficiency. To solve this problem, this embodiment will Constraints are incorporated into the design of the reward function. By introducing constraints that change with the training process Constraints, the training process can be segmented, thereby reducing the sensitivity to the weights of the reward function. Secondly, considering that the reward function has multiple components, The effectiveness of the constraint is limited. In this embodiment, a self-tuning method for the reward weight is also designed to autonomously design an appropriate reward weight. In this embodiment, the way to achieve weight self-tuning is to train using two different reward functions. According to the relative magnitudes of each data item in the training results and the training potential of each data item measured by the difference between the two training results, the weights are adjusted autonomously to balance the relative importance and optimization potential between each data item.
[0124] In one implementation, the reward function for calculating the reward value is:
[0125] , formula (21);
[0126] Where represents the reward value; represents the weight of each term in the reward function; represents the constraint, which is used to reduce the sensitivity to the reward function weight; represents the penalty term corresponding to the no-go area; represents the penalty term corresponding to the maximum thrust threshold; represents the penalty term corresponding to the single-step fuel consumption; represents the final position error; represents the final velocity error; represents the penalty term corresponding to the position deviation; represents the penalty term corresponding to the velocity deviation; represents the reward term corresponding to the next state data.
[0127] Specifically, predict the next state based on the current observation and expected operation , , , respectively represent the predicted spacecraft position, velocity, and mass at the next state. The reward terms in the reward function related to the predicted state are defined as follows:
[0128] , formula (22);
[0129] Where is the remaining mass of the spacecraft at present, and are the position and velocity of the reference trajectory at the next time step .
[0130] The specific definitions of the penalty terms for the final position error and the final velocity error are as follows:
[0131] , formula (23);
[0132] , formula (24);
[0133] where , respectively represent the terminal position and velocity error, which are only defined at the end of the transfer (i.e., when ). Here and respectively represent the final position and velocity of the spacecraft at the end of the mission, while and correspond to the position and velocity of the target planet at the end of the rendezvous mission. and respectively represent the current time step and the total number of time steps.
[0134] The specific definition of the penalty term for single-step fuel consumption is as follows:
[0135] , formula (25);
[0136] where represents the final remaining mass of the spacecraft at the end of the mission; represents the remaining mass of the spacecraft at the th time step.
[0137] The specific definition of the penalty term for the maximum thrust threshold is as follows:
[0138] , formula (26).
[0139] The specific definition of the penalty term for entering the prohibited area is as follows:
[0140] , formula (27);
[0141] where, and are the upper and lower bounds of the established prohibited area.
[0142] The specific definitions of the penalty terms for the position deviation and velocity deviation established based on the reference trajectory constraint are as follows:
[0143] , formula (28);
[0144] , formula (29);
[0145] where and are the position and velocity of the reference trajectory at the time step .
[0146] Therefore, by integrating equations (22)-(29), the complete reward function, i.e., equation (21), is obtained.
[0147] To facilitate the understanding of the technical solution of this embodiment, taking the Earth-Mars orbit transfer mission as an example, the overall algorithm process is described as follows:
[0148] In the Earth-Mars orbit transfer mission, for a spacecraft equipped with a low-thrust electric propulsion device, the initial position , the initial velocity , and the initial mass are given. Through low-thrust orbit maneuvers, it is transferred to the target orbit within a specified time and reaches the same velocity as the target planet .
[0149] The kinematic and dynamic models of the spacecraft are constructed, i.e., equations (3)-(5) are obtained. Among them, represents the standard gravitational parameter, and is the specific impulse of the electric propulsion system.
[0150] After discretizing the low-thrust orbit, the maximum pulse value and the pulse value at the -th time step are defined as in equations (6) and (7). Among them, = 0.5 represents the maximum thrust amplitude that the spacecraft engine can provide, = 348.79 days represents the total transfer time, and is taken as 40.
[0151] The calculation formulas for the state transition matrix and the observation vector are constructed, i.e., equations (1) and (2) are obtained. Among them, during the construction process, the , , , represent the standard deviations of position and velocity in state uncertainty and observation uncertainty, and their values are = 1 km, = 0.05 km / s, = 1 km, = 0.05 km / s.
[0152] Through the sequential latent variable model and the reinforcement learning controller, a stochastic hidden state proximal policy optimization framework is established. Among them, the neural network describing the hidden state dynamics is set as follows: the hidden state dimension is set to , , all distributions consist of an input layer with respective input dimensions, two fully connected hidden layers with 64 nodes each, and an output layer with respective output dimensions. The learning rate for this part is set to linearly decrease throughout the training process starting from the initial value and continue to decrease linearly throughout the training process.
[0153] The evidence lower bound is derived through Equation (14), where in Equation (14) is the sequence length of the single-sequence latent variable model, taking = 4.
[0154] The objective formula for constructing the actor network is obtained, that is, Equation (17), where represents the hyperparameter controlling the pruning range, with an initial value of = 0.3 and linearly decreasing as the training process progresses. And the Generalized Advantage Estimation (GAE) is calculated through the following formula :
[0155] , Equation (30);
[0156] where is the adjustable hyperparameter in GAE, taking = 0.99; is the reward discount factor, taking = 0.9999.
[0157] The calculation formulas for the mean squared value function error term and the entropy term are constructed, that is, Equation (18) and Equation (19), and then the overall objective function of the reinforcement learning controller part is constructed, that is, Equation (16). , in Equation (16) represent the hyperparameters controlling the relative importance of the three objectives, taking = 0.5, = 4.75x10 -8 . The actor network and the critic network are set with an input layer of respective input dimensions, two fully connected hidden layers with 64 nodes each, and an output layer of respective dimensions. The learning rate is set to linearly decrease throughout the training process starting from the initial value = 2.5x10 -4 and continue to decrease linearly throughout the training process.
[0158] The reward function is constructed, that is, Equation (21), where the value of in Equation (21) is set to:
[0159] , Equation (31);
[0160] where represents the current training step, Denotes the total number of training steps, which is set to 10 here 8 .
[0161] Secondly, for the penalty term of the restricted area, i.e., formula (27), and are the upper and lower bounds of the established restricted area, taking = 0.8 AU and = 1.68 AU .
[0162] In addition, two different reward functions are used for training respectively, taking the initial weights and to perform self-tuning of the reward weights.
[0163] Finally, the weight tuning result is obtained . Figure 4 Shows the comparison of the training curves using the initial weights and the tuned weights. It can be seen that the reward curve obtained using the tuned weights is significantly better than the reward curve generated using the initial weights. When the training reaches steps, the agent trained using the adjusted weights has almost converged, while the agent trained using the initial weights still shows considerable fluctuations and has not converged. Near training steps, both curves seem to have converged. However, by locally magnifying the curves, it can be seen that the agent using the adjusted weights obtains higher and more stable rewards, while the reward curve using the initial weights still has small fluctuations. In summary, the above experimental data prove that the reward weight automatic adjustment method of this embodiment can effectively enable the agent to converge quickly and obtain better convergence results.
[0164] The present invention has at least the following four beneficial effects, and the four beneficial effects can be arbitrarily combined:
[0165] 1. The present invention explicitly represents and learns the uncertain environment through a sequential latent variable model, extracts the hidden information hidden under the observations of the uncertain environment, thereby accelerating the training of the agent in the uncertain environment and improving the agent's ability to handle uncertainty.
[0166] 2. Through the reinforcement learning framework that combines the sequential latent variable model and the proximal policy optimization algorithm, the agent can learn effective policies during continuous interaction with the environment, without a large amount of prior knowledge and accurate models, improving the adaptability to complex environments and uncertainties, and enhancing the robustness of orbit transfer design.
[0167] 3. The previous low-thrust orbit design method based on reinforcement learning may have the problem of sparse rewards when designing the reward function, resulting in low learning efficiency. The adaptive hidden state reinforcement learning method disclosed in the present invention combines the future state predicted by the sequential latent variable model and other relevant constraints in the orbit transfer process to design a dense reward function, which can more effectively guide the agent to learn, improve the learning efficiency, and accelerate the convergence to a better orbit transfer strategy. The introduced reward term related to the predicted state enables the agent to make decisions not only based on the current immediate reward but also taking into account the rewards that may be obtained in the future, thereby helping the agent to make more forward-looking and reasonable decisions and avoiding short-sighted behaviors.
[0168] 4. The previous reinforcement learning methods may lack flexibility when determining the weights of the reward function, unable to well adapt to different task scenarios and environmental changes, or require a large amount of engineering experience in the design process. The adaptive hidden state reinforcement learning method disclosed in the present invention adopts a reward function weight self-tuning algorithm, which can automatically adjust the weights of the reward function according to the environmental feedback, improve the flexibility and adaptability of the method, optimize the orbit transfer design effect, and further improve the training efficiency.
[0169] To prove the above technical effects, the present invention trains according to the neural network and environment established in the previous steps and conducts Monte Carlo simulation verification on the trained model. Figure 5 Shows the velocity increment distribution at each impulse point of the designed trajectory. Figure 6 Shows the obtained robust trajectory and Monte Carlo trajectories. One line depicts the robust trajectory generated by the policy. The arrows on the robust trajectory represent the thrust vector at each time step. The arrow direction reflects the thrust direction, and its length represents the thrust magnitude. The other two lines depict the trajectories of Mars and Earth during the rendezvous mission, and there is another line representing the trajectories generated by 1,000 Monte Carlo simulations to evaluate its performance in handling uncertainties. To improve the display, the error between the Monte Carlo simulation trajectories and the optimal robust trajectory (i.e., the robust trajectory generated by the policy) is magnified five times. As Figure 6 can be seen, the low-thrust orbit transfer design method proposed by the present invention can achieve robust orbit transfer in an uncertain environment. In an uncertain environment, the self-designed trajectory curve has only a little error in the middle part and can ultimately achieve the orbit transfer goal.
[0170] The possible technical variations of the present invention are as follows:
[0171] The variational autoencoder (VAE) can be used to replace the sequence latent variable model used in the present invention for representation learning of an uncertain environment. Compared with the sequence latent variable model used in the present invention, the variational autoencoder can achieve representation learning of data by mapping the input data to a latent space and encoding and decoding the data in the latent space. However, the sequence latent variable model focuses more on modeling the dynamic information in sequence data and can better capture the dependencies in time series. When dealing with sequence data, the variational autoencoder may require additional processing steps or model structures to consider the information in the time dimension. In addition, the variational autoencoder has a weak predictive generation ability, which may lead to inaccurate predicted states in the solution, thereby affecting model training.
[0172] In summary, the present invention discloses an interplanetary orbit transfer method based on hidden states and reinforcement learning, which relates to the field of spacecraft control technology. By establishing a sequence latent variable model, the present invention can explicitly perform representation learning of an uncertain environment and extract the hidden information hidden under the observations of the uncertain environment. And a reinforcement learning controller is composed of an actor network and a critic network. A reinforcement learning algorithm framework is constructed by using the sequence latent variable model and the reinforcement learning controller, thereby accelerating the training of the agent in the uncertain environment and improving the agent's ability to handle uncertainty. In addition, the embodiment of the present invention also predicts the next state based on the current observation and the expected operation and incorporates the quality of the predicted next state into the reward structure, so that the immediate reward can capture the effectiveness of the current and subsequent policies at the same time, thereby improving the learning efficiency of the algorithm.
[0173] It should be understood that the application of the present invention is not limited to the above examples. Those of ordinary skill in the art can make improvements or transformations according to the above description, and all such improvements and transformations should fall within the protection scope of the appended claims of the present invention.
Claims
1. An interplanetary orbit transfer method based on hidden states and reinforcement learning, characterized in that The method includes: Obtaining the observation data of the spacecraft; wherein, the observation data is determined based on the state data and the observation uncertainty; the state data is calculated based on the state transition matrix constructed for a preset orbit transfer mission; Inputting the observation data into a preset sequential latent variable model to obtain latent variables; Inputting the latent variables and the historical trajectory record of the spacecraft into a reinforcement learning controller to obtain the execution actions of the spacecraft; Predicting the next state data of the spacecraft according to the observation data and the execution actions, calculating a reward value according to the state data and the next state data, and enabling the reinforcement learning controller to update the orbit transfer strategy based on the reward value; Wherein, the reinforcement learning controller includes: A critic network, which is used to estimate the state value according to the input latent variables and its own value estimation function, and update the orbit transfer strategy according to the state value; An actor network, which is used to calculate the execution actions of the spacecraft according to the historical trajectory record and the orbit transfer strategy; The state transition matrix is: ; Among them, and represent the positions of the spacecraft at the k-th time step and the (k + 1)-th time step respectively; and represent the velocities of the spacecraft at the k-th time step and the (k + 1)-th time step respectively; and represent the masses of the spacecraft at the k-th time step and the (k + 1)-th time step respectively; and represent the Lagrangian coefficients at the k-th time step; and represent the derivatives of the Lagrangian coefficients at the k-th time step; The vector is the output of the reinforcement learning controller, representing the velocity change obtained by impulse inference ; represents the equivalent exhaust velocity of the thruster; represents the position deviation caused by state uncertainty at the (k + 1)-th time step; represents the velocity deviation caused by state uncertainty at the (k + 1)-th time step; The observation data is in vector form, and the calculation formula of the observation vector is: ; Among them, represents the observation vector; represents time, , represents the total number of time steps; represents the transpose of a matrix or vector; represents the position deviation caused by observation uncertainty at the k-th time step; represents the velocity deviation caused by observation uncertainty at the k-th time step.
2. The interplanetary orbit transfer method based on hidden state and reinforcement learning according to claim 1, characterized in that The sequential latent variable model consists of an action model, latent variables, an observation model, and latent variable dynamics, and the latent variable dynamics is used to establish the connection between the observation model, the action model, and the latent variables; The objective function of the sequential latent variable model is established based on the equivalence principle of maximizing the observation data probability and maximizing the evidence lower bound.
3. The interplanetary orbit transfer method based on hidden state and reinforcement learning according to claim 2, wherein The latent variable dynamics includes a generative model and an inference model; The generative model is used to sequentially generate the next latent variable and the prior distribution of each latent variable according to the previous latent variable information; The inference model is used to calculate the posterior distribution of the current latent variable according to the subsequent latent variable information and the observation data.
4. The interplanetary orbit transfer method based on hidden state and reinforcement learning according to claim 3, characterized in that The objective function of the sequential latent variable model is: ; Among them, represents the objective function of the sequence latent variable model; represents the sequence length of the single-sequence latent variable model; represents the relative entropy of two distributions; represents the probability of each observed data; represents the evidence lower bound; represents the observation vector; is the latent variable at the k-th time step, and respectively represent the first-layer latent variable and the second-layer latent variable at the k-th time step; represents the execution action at the k-th time step; represents the function with respect to the trajectory sampled from the probability distribution q of the expected value.
5. The interplanetary orbit transfer method based on hidden state and reinforcement learning according to claim 1, characterized in that The overall objective function of the reinforcement learning controller consists of the objective formula of the actor network, the mean square value function error term for value estimation based on the critic network, and the entropy term of the decision distribution generated by the policy function.
6. The interplanetary orbit transfer method based on hidden state and reinforcement learning according to claim 5, wherein The overall objective function of the reinforcement learning controller is: ; Among them, represents the overall objective function of the reinforcement learning controller; represents the objective formula of the actor network; represents the mean square value function error term; represents the entropy term; represents the trainable parameters of the neural network corresponding to the reinforcement learning controller; and represents the hyperparameter; The objective formula of the actor network is: ; Among them, represents the target formula of the actor network; represents along the policy the trajectory sampled expectation; represents the total number of time steps; represents the importance weight, represents the old policy used when sampling data, represents the execution of an action, represents the system state; represents to the value range limited to function, represents the importance weight at the represents the hyperparameter for controlling the pruning range; represents the generalized advantage estimation; The calculation formula of the mean square value function error term is: ; Among them, represents the mean square value function error term; represents the latent variable; represents the self - value estimation function established by a neural network with parameter ; represents the th power of the reward discount factor; represents the environmental reward at the th time step; The calculation formula of the entropy term is: ; Among them, represents the entropy term; represents the th time step and the step trajectories before it.
7. The interplanetary orbit transfer method based on hidden state and reinforcement learning according to claim 1, wherein Calculating the reward value according to the state data and the next state data includes: Obtaining the reference trajectory of the spacecraft, wherein the reference trajectory is the optimal trajectory obtained based on a deterministic environment; Calculating a reward term according to the next state data and the reference trajectory; Obtaining the target state data corresponding to the target planet, calculating the final position error and the final velocity error of the spacecraft according to the state data and the target state data, and calculating a penalty term according to the final position error and the final velocity error respectively; Calculating the single-step fuel consumption of the spacecraft according to the state data, and calculating a penalty term according to the single-step fuel consumption; Calculating a penalty term according to the state data and a preset maximum thrust threshold; Calculating a penalty term according to the state data and a preset no-entry area; Calculate the position deviation and velocity deviation according to the state data and the reference trajectory, and calculate a penalty term according to the position deviation and the velocity deviation respectively; Calculate the reward value according to the reward term and all the penalty terms.
8. The interplanetary orbit transfer method based on hidden state and reinforcement learning according to claim 7, wherein Calculating the reward value according to the reward term and all the penalty terms includes: Calculating the reward value according to the weighted combination of the reward term and all the penalty terms; Among them, the adjustment method of the weights corresponding to the reward term and each penalty term includes: Establish a reward function according to the reward term and all the penalty terms, wherein the penalty term established based on the final position error and the final velocity error includes a constraint value that changes with the training process; Train using two reward functions with different weight values respectively; Determine the training potential corresponding to the reward term and each penalty term according to the relative magnitudes of the reward term and each penalty term in the training results, and the difference between the training results of the two reward functions; Adjust the weights according to the training potential of the reward term and each penalty term.
9. The method for interplanetary orbit transfer based on hidden state and reinforcement learning according to claim 7, characterized in that The reward function for calculating the reward value is: ; Among them, represents the reward value; represents the weight of each term in the reward function; represents a constraint used to reduce the sensitivity to the weights of the reward function; represents the penalty term corresponding to the no-go area; represents the penalty term corresponding to the maximum thrust threshold; represents the penalty term corresponding to the single-step fuel consumption; represents the final position error; represents the final velocity error; represents the penalty term corresponding to the position deviation; represents the penalty term corresponding to the velocity deviation; represents the reward term corresponding to the next state data.
Citation Information
Patent Citations
Reinforcement learning model training method and device
CN117669650A
Track prediction intelligent optimization system and method
CN119415877A