Satellite orbit control method based on continuous time near-end strategy optimization reinforcement learning algorithm
By optimizing satellite orbit control through continuous-time proximal strategy optimization reinforcement learning algorithm and deep neural network, the problem of insufficient flexibility and responsiveness of traditional methods in dynamic environments is solved, and efficient and stable satellite orbit control is achieved.
Patent Information
- Application Number
- CN202511243345.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-02
- Publication Date
- 2025-09-30
- Estimated Expiration
- 2045-09-02
AI Technical Summary
Traditional satellite orbit control methods lack flexibility in dealing with unexpected dynamic changes and emergency situations, and the discrete-time decision model makes it impossible to respond to rapid environmental changes in a timely manner, affecting the effectiveness and stability of control.
A reinforcement learning algorithm based on continuous-time proximal policy optimization is adopted. The satellite orbit control is optimized through deep neural networks and reward functions. The state is updated in combination with the 8th-order Runge-Kutta method to realize action and time decision-making in continuous time.
It improves the real-time and accuracy of satellite orbit control, enhances the system's responsiveness and overall efficiency, ensures a high degree of synchronization and stability between decision-making and environmental conditions, and adapts to changes in complex orbital environments.
Smart Images

Figure CN120722768A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of satellite orbit control methods, and in particular to a satellite orbit control method based on a continuous-time proximal strategy optimization reinforcement learning algorithm. Background Art
[0002] In the field of satellite orbit control, achieving precise decision-making and control is crucial for the effective execution of various missions. Traditional control methods rely primarily on predetermined trajectories and pre-programmed instructions. These methods often lack sufficient flexibility to handle unexpected dynamic changes or respond to emergency situations. Furthermore, traditional methods are typically based on discrete-time decision models, where decision intervals are fixed. This can result in an inability to respond to rapid environmental changes in a highly dynamic orbital environment, thus compromising the effectiveness and stability of control.
[0003] As satellite missions become increasingly complex, the demand for real-time and adaptive control strategies in orbital pursuit environments is growing. These missions require control systems to analyze the current state in real time, make dynamic control decisions, and optimize orbital operations to adapt to changing environmental conditions. In this context, reinforcement learning, an artificial intelligence approach that learns optimal strategies through interaction with the environment, shows great potential. However, most current reinforcement learning methods are based on discrete time steps, with a fixed time interval between decision and execution. This can easily lead to a mismatch between decision and execution in orbital control tasks that require rapid and continuous decision-making, thereby reducing the effectiveness and safety of operations. Summary of the Invention
[0004] This paper presents a satellite orbit control method based on a continuous-time proximal policy optimization reinforcement learning algorithm. This method aims to achieve efficient, stable, and precise control of satellite orbits through advanced reinforcement learning techniques. By utilizing a continuous-time framework and the proximal policy optimization reinforcement learning algorithm (PPO), this method overcomes the limitations of traditional discrete-time control methods and demonstrates superior control performance in dynamic and complex orbital environments.
[0005] In order to achieve the above object, the technical solution adopted by the present invention is: A satellite orbit control method based on a continuous-time proximal policy optimization reinforcement learning algorithm includes the following steps: Step 1: Based on the continuous-time proximal strategy, the reinforcement learning algorithm is optimized to determine the action taken by the satellite, i.e., the acceleration, and the execution time of the action. The process is as follows: Establish a satellite agent, which defines the satellite agent's position state space and action space. The position state space includes the satellite agent's position state vector in a continuous time frame, and the action space includes the actions taken by the satellite agent in a continuous time frame, i.e., acceleration. A reward function is also established for the satellite agent. Build a policy network, train it using historical satellite trajectory data, and have the trained policy network output recommended actions and execution times to the satellite agent. After receiving the recommended action and execution time output by the trained policy network, the satellite agent evaluates the recommended action and execution time received through the reward function, thereby determining the optimal action, i.e., acceleration, and optimal execution time, which are used as the action, i.e., acceleration, and execution time, taken by the satellite; Step 2: Based on the actions taken by the satellite and the execution time obtained in step 1, the satellite orbit is controlled.
[0006] Furthermore, in step 1, the reward function of the satellite agent is Feedback to satellite agents based on execution time The acceleration magnitude, cost, and state deviation when executing the action, i.e., acceleration a, and the reward function As shown in the following formula: , in: Indicates that the satellite agent is based on execution time After executing the action, i.e. acceleration a, the position state vector s and the target position state vector Deviation between and are the cost coefficient and delay cost coefficient of the satellite agent’s action execution, i.e. acceleration a, respectively. is the magnitude of the action, i.e., acceleration a, which reflects the cost of executing the acceleration; Indicates movement, i.e. acceleration Execution time.
[0007] Furthermore, in step 1, the strategy network uses a deep neural network , as the deep neural network of the policy network, is composed of multiple layers of fully connected layers, and the output of each fully connected layer serves as the input of the next fully connected layer; represents the weight and bias parameters of the deep neural network, s is the current position state vector of the satellite, is the recommended action output by the policy network, i.e., acceleration, The recommended action output by the policy network is acceleration Execution time.
[0008] Furthermore, deep neural networks The activation function of each fully connected layer in is the ReLU function.
[0009] Furthermore, deep neural networks That is, when training the policy network, the position state vector and velocity vector in the historical satellite trajectory data are used as the deep neural network The input of the deep neural network After the multi-layer fully connected layers are calculated in sequence, the recommended action is obtained, i.e., the acceleration , and perform the recommended action i.e. acceleration Execution time .
[0010] Furthermore, the deep neural network The objective function during training is based on the advantage function Constructing the objective function for clipping , the objective function of clipping As shown in the following formula: , in: Indicates that the policy network after parameter update is a deep neural network and the old policy network before parameter update, i.e., the deep neural network The selection probability ratio of the output recommended action; represents the currently updated network parameters, and Represents the previous old strategy network parameters; is a clipping parameter used to limit the range of variation of the probability ratio; Used to clip probability ratios To limit the magnitude of policy updates, if Out of range , then Clipping to this range avoids excessive or unstable updates of the strategy; Represents the position state and action in the old policy network Expect the distribution under represents the advantage function, advantage function The calculation formula is as follows: , in: The satellite performs the action, i.e., acceleration In position The action value function under is the value function, representing the position expected returns; During training, the advantage function It is approximated by the generalized advantage estimation method based on empirical data, and the advantage estimate is adjusted by weighting the time series difference residuals at different time steps, thereby approximately solving the advantage function , the calculation formula is as follows: , in: A t express Advantage function at time The generalized advantage estimation method of is used to approximate the solution, that is, Differential estimation of the acceleration value relative to the position value at each moment; ,in Indicates the satellite at the historical moment The reward obtained after performing an action; is the position in the current policy network The estimated value of Expected returns under Is the next position state The estimated value of is the discount factor used to adjust the present value of future rewards; is a smoothing parameter that adjusts the trade-off between bias and variance of the advantage estimate; Minimize the clipping objective function during training , based on the objective function of clipping Gradient ascent method for deep neural networks Parameters in to update.
[0011] Furthermore, after step 2, the advantage function is solved by the generalized advantage estimation method according to the performance changes of the policy network. Discount factor , and the objective function of clipping The cropping parameters in , thereby further optimizing the policy network.
[0012] Furthermore, in step 2, the satellite state update equation is established, and then the satellite state update equation is solved by combining the 8th-order Runge-Kutta method and the action and execution time obtained in step 1, and the satellite orbit is controlled according to the solution.
[0013] Furthermore, in step 2, an integral form of the satellite state update equation is established, and then the 8th-order Runge-Kutta method is used to approximate the integral part of the integral form of the satellite state update equation based on the action and execution time obtained in step 1, thereby solving the satellite state update equation.
[0014] This paper proposes a satellite orbit control method based on a continuous-time proximal policy optimization reinforcement learning algorithm, thereby achieving satellite orbit control. The proposed continuous-time proximal policy optimization reinforcement learning algorithm fully considers the characteristics of a continuous-time environment and uses numerical integration techniques to achieve continuous state prediction and updating, ensuring real-time correspondence and high synchronization between decision-making and environmental states.
[0015] To overcome the limitations of discrete-time reinforcement learning models in dynamic environments, this paper employs a continuous-time proximal policy optimization reinforcement learning algorithm. This allows the algorithm to process time and state in a continuous manner, rather than relying on predefined time intervals. The introduction of a continuous-time framework enables reinforcement learning algorithms to more accurately simulate and predict satellite behavior in complex orbital environments.
[0016] Furthermore, the introduction of a planned acceleration execution mode further enhances the algorithm's applicability. Through this mode, the present invention not only determines the optimal acceleration selection but also calculates and schedules the optimal acceleration execution time, effectively eliminating the uncertainty and risk associated with acceleration execution delays.
[0017] This paper applies the Proximal Policy Optimization (PPO) algorithm to this continuous-time, planned acceleration framework and, through a truncated policy gradient, addresses the problem of excessive oscillations that can arise from policy updates in high-variability environments. This improvement in the PPO algorithm ensures the stability of the learning process and the reliability of the policy, making it more suitable for applications such as satellite orbit control, which require extremely high safety and accuracy.
[0018] This paper combines a continuous-time reinforcement learning framework with a planned execution acceleration model to address the dynamic and complex problems encountered in satellite orbit control. This approach improves the real-time and accuracy of operations, while also enhancing the overall efficiency and responsiveness of the system. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] Figure 1 It is a flow chart of a method according to an embodiment of the present invention. DETAILED DESCRIPTION
[0020] To help those skilled in the art better understand the present invention, the following detailed description of the embodiments of the present invention is provided in conjunction with the accompanying drawings and examples. This will help those skilled in the art to fully understand and implement the present invention by applying technical means to solve technical problems and achieve corresponding technical effects. The embodiments of the present invention and the various features therein may be combined with each other as long as they do not conflict with each other, and the resulting technical solutions are all within the scope of protection of the present invention.
[0021] Obviously, the embodiments described are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of the present invention.
[0022] It should be noted that the terms "include" and "have" in the specification, claims and drawings of the present invention and any variations thereof are intended to cover non-exclusive inclusions.
[0023] like Figure 1 As shown, this embodiment discloses a satellite orbit control method based on a continuous-time proximal policy optimization reinforcement learning algorithm, comprising the following steps: Step 1: Based on the continuous-time proximal strategy, the reinforcement learning algorithm is optimized to determine the action taken by the satellite, i.e., the acceleration, and the time when the action is executed. The specific process is described as follows: (1) For satellite orbit control, this embodiment establishes a satellite agent, in which the satellite agent defines the position state space and action space of the satellite agent. The position state space includes the position state vector of the satellite agent in the continuous time frame. ,in x , y , z Represents the x, y, and z axis coordinates of the satellite agent in the three-dimensional space coordinate system, and the position state vector A complete dynamic description of the satellite agent in three-dimensional space is provided, which provides a basis for the action selection, i.e., acceleration selection, in the decision-making process.
[0024] The action space includes the actions taken by the satellite intelligence in the continuous time frame, that is, the acceleration. Let the position of the satellite intelligence at time t be s(t) and the velocity be v(t). The action performed by the satellite intelligence at time t, that is, the acceleration, is , execution action is acceleration The execution time is , then the satellite agent executes at time t according to the execution time Execution action is acceleration After that, the satellite agent changes its velocity vector v(t) from time t to t+ The speed at the moment v(t+ ), and change the satellite agent's position s(t) from time t to t+ The position at the moment s(t+ ).
[0025] In this embodiment, a reward function is also established for the satellite agent , the reward function of the satellite agent As shown in the following formula: , in: Indicates that the satellite agent is based on execution time After executing the action, i.e. acceleration a, the position state vector s and the target position state vector Deviation between and are the cost coefficient and delay cost coefficient of the satellite agent’s action execution, i.e. acceleration a, respectively. is the magnitude of the action, i.e., acceleration a, which reflects the cost of executing the acceleration; Indicates movement, i.e. acceleration Execution time.
[0026] The reward function Feedback to satellite agents based on execution time The acceleration magnitude, cost, and state deviation when executing the action, i.e., acceleration a, are calculated through the reward function It aims to guide the satellite agent to optimize the cost and execution time of acceleration usage while minimizing position deviation, thereby achieving efficient orbit control.
[0027] (2) In this embodiment, a policy network is constructed to generate recommended actions for the satellite agent, i.e., recommended accelerations, and recommended action execution times.
[0028] The policy network uses a deep neural network ,in represents the weight and bias parameters of the deep neural network, s is the current position state vector of the satellite, The recommended action output by the policy network is the recommended acceleration. Recommended action output by the policy network, i.e., recommended acceleration Execution time.
[0029] The deep neural network as a policy network is composed of multiple layers of fully connected layers (also called dense layers). The output of each fully connected layer serves as the input of the next fully connected layer. The calculation of each fully connected layer of the deep neural network can be expressed as shown in the following formula: , in: It is The output of the fully connected layer, No. +1 fully connected layer output; and They are The weight matrix and bias vector of the fully connected layer; G( ) is the activation function, which is used to introduce nonlinearity so that deep neural networks can learn and simulate more complex functional relationships.
[0030] In the deep neural network of this embodiment, the ReLU (Rectified Linear Unit) function is used as the activation function G( ) of each fully connected layer, ,in x The ReLU function outputs the input value when the input is greater than zero, otherwise it outputs zero.
[0031] In this embodiment, the deep neural network as the policy network is trained by using the historical satellite trajectory data (the position state of the satellite at each moment in history - the action and the corresponding reward information). During training, the position state vector of each moment in the historical satellite trajectory data is used and velocity vector As a deep neural network The input, where v x 、 v y 、 v z Represents the speed of the satellite v The components of the x, y, and z axes in the three-dimensional space coordinate system. After the multi-layer fully connected layers are calculated in sequence, the recommended action is obtained, that is, the recommended acceleration , and perform the recommended action, i.e. the recommended acceleration Execution time , through training to realize the deep neural network as the policy network This training process involves calculating gradients and updating network parameters to improve prediction accuracy and policy effectiveness.
[0032] In this embodiment, based on the advantage function Constructing the objective function for clipping , to clip the objective function As a deep neural network Objective function during training. Objective function based on clipping during training To update the deep neural network Parameters , and through the clipping objective function Controlling deep neural networks during training The amplitude of the update and prevent performance collapse during the update process. The objective function of the clipping The expression is as follows: , in: Indicates that the policy network after parameter update is a deep neural network and the old policy network before parameter update, i.e., the deep neural network The selection probability ratio of the output recommended action; represents the currently updated network parameters, and Indicates the previous old policy network parameters. It is the key quantity in the algorithm to measure the difference between the new and old strategy networks. When it is close to 1, it means that the new policy network is very similar to the old policy network; when When the deviation is 1, it means that the preference of the new policy network for the current recommended action is significantly different from that of the old policy network.
[0033] It is a clipping parameter used to limit the range of variation of the probability ratio. The initial clipping parameter can be set to 0.1.
[0034] Used to clip probability ratios To limit the magnitude of policy updates, if Out of range , then Clipping to this range can avoid excessive or unstable updates of the strategy.
[0035] Represents the position state and action in the old policy network The expectation of the distribution under the old policy network is calculated, that is, the expectation calculation of the objective function is performed on the position state-action data collected during the operation of the old policy network; when the proximal policy optimization algorithm updates the policy network, it will use the trajectory data sampled by the old policy network, and then evaluate and optimize the new policy network based on this data.
[0036] represents the advantage function, advantage function It is a measure of the speed at which an action is taken, i.e. acceleration and execution time When, at position The additional value of the clipped objective function relative to the average policy is In the process of calculating the advantage function, we refine the learning of the strategy to achieve a more efficient and accurate control strategy. Advantage function It is the key metric for updating the policy network in the continuous-time proximal policy optimization reinforcement learning algorithm of this embodiment, and is used to guide the policy network to a more favorable recommended acceleration, i.e., acceleration update. Advantage function The calculation formula is as follows: , in: The satellite performs the action, i.e., acceleration In position The action value function under is the value function, representing the position expected return.
[0037] During training in this embodiment, the advantage function It does not directly follow from the definition Instead of calculating it precisely, it is approximated by the generalized advantage estimation method based on empirical data. It is usually calculated by sampling the reward and the value estimate of the future position state, and the value function Independent value network To reduce the variance and improve the accuracy of the advantage estimate, this embodiment adopts a generalized advantage estimation method, which adjusts the advantage estimate by weighting the time series difference residuals at different time steps, thereby approximating the advantage function , the calculation formula is as follows: , in: A t express Advantage function at time The generalized advantage estimation method of is used to approximate the solution, that is, Differential estimation of the acceleration value relative to the position value at each moment; ,in Indicates the satellite at the historical moment The reward obtained after performing an action; is the position in the current policy network The estimated value of Expected returns under Is the next position state The estimated value of is the discount factor used to adjust the present value of future rewards.
[0038] is a smoothing parameter that adjusts the trade-off between bias and variance of the advantage estimate; Advantage function After calculation, the policy network is a deep neural network The objective function can be clipped Iteratively update the network parameters. In each learning cycle, the deep neural network Use data obtained from the trajectory to adjust its parameters , the objective function for optimizing clipping , ensuring gradual improvement of the policy network while preventing over-updates.
[0039] In this embodiment, the objective function based on clipping , using the gradient ascent method to train the policy network, i.e., the deep neural network Parameters in Update as shown below: , in: is the learning rate; is the gradient of the objective function with respect to the policy network parameters; Represents the new parameters of the policy network after updating at the current time step; Represents the original parameters of the policy network before updating at the current time step.
[0040] Through this objective function calculation and policy network update mechanism based on advantage function clipping, the continuous-time proximal policy optimization reinforcement learning algorithm of this embodiment can effectively drive the policy toward optimization while maintaining learning stability.
[0041] In this embodiment, the process of training the policy network using historical satellite trajectory data is as follows: First, based on the historical satellite trajectory data, a Position status at the moment ,action , execution time ,award and the next state training samples.
[0042] Then, the training samples are input into the policy network, and the policy network parameters are optimized iteratively. , by minimizing the clipping objective function , improve the performance of the policy network in the track control task. After each iteration, the trained policy network is used to and speed , output recommended actions, i.e. acceleration, to the satellite agent and execution time .
[0043] In this embodiment, the policy network trained by historical satellite trajectory data can provide high-precision acceleration and execution time decisions in satellite orbit control. In particular, from the current time t to the execution time The continuous time integral form between enables the strategy network to more accurately describe the behavior and change process of the satellite, and the control signal can be adjusted more smoothly, avoiding the jumps and discontinuities that may occur in discrete time.
[0044] (3) In this embodiment, the satellite agent receives the recommended action output by the trained policy network, that is, the recommended acceleration. and execution time Then, through the reward function Evaluate and execute recommended actions, i.e. acceleration The effect of , including the size of the action, the cost and the state deviation. Based on the reward function Based on the feedback, the optimal action, i.e. acceleration, and optimal execution time, are determined from the multiple recommended actions and execution times output by the policy network, that is, the optimal solution that achieves the best balance between efficiency, cost, and state deviation.
[0045] Therefore, in this embodiment, based on the above continuous time proximal strategy optimization reinforcement learning algorithm, the optimal acceleration at time t is finally determined by the satellite agent. and optimal execution time (opt), as the action taken by the satellite at time t, i.e., the acceleration and the execution time.
[0046] Step 2: Based on the actions taken by the satellite and the execution time obtained in step 1, the satellite orbit is controlled.
[0047] Specifically, in this embodiment, a satellite state update equation is established. Then, the satellite state update equation is solved by combining the 8th-order Runge-Kutta method and the action and execution time obtained in step 1. The satellite orbit is controlled based on the solution. The process is as follows: (2.1) Establish the first-order differential dynamic equations of the satellite state, including the first-order differential dynamic equations of the position vector and the first-order differential dynamic equations of the velocity vector.
[0048] In this embodiment, in order to accurately simulate the orbital dynamics of the satellite, a dynamic equation of the satellite's motion state in orbit is first established. The satellite's motion state can be expressed by the position vector and velocity vector To fully describe, Represents the spatial position coordinates of the satellite relative to the center of the earth, Represents the velocity component, x, y, and z are the x-axis, y-axis, and z-axis coordinates of the satellite's spatial position, respectively. 、 、 They are x 、 y 、 z The first-order derivative of . And in this embodiment, taking into account the influence of perturbations such as the J2 effect, these states of the satellite can be described by the following second-order dynamic differential equations: , in, Represents the net force of all external forces, especially the main gravitational term ; is the position vector The first derivative of ; is the position vector The second derivative of or the velocity vector The first derivative of is the gravitational constant; is the mass of the Earth; l is the distance from the satellite to the center of the Earth.
[0049] The above second-order differential dynamics equation is transformed into an equivalent first-order differential dynamics equation by introducing the velocity vector As a new variable, the first-order differential dynamics equation is shown as follows: , , in, t For time; is the velocity vector The first-order derivative of .
[0050] (2.2) Based on the above first-order differential dynamics equation, a continuous time framework is adopted to establish the satellite state update equation in integral form, including the position state update equation and the velocity state update equation.
[0051] The current state of a satellite is represented by its position vector ( ) and the velocity vector ( ). To ensure the continuity and accuracy of state changes, this embodiment establishes an integral form of satellite state update equations, including position state update equations and velocity state update equations, as shown in the following formula: , , in: Indicates the satellite at time speed; represents the action exerted by the satellite at time t, i.e., acceleration; Indicates the satellite at time Position status; Indicates the satellite at time Speed state; Indicates the updated satellite position status; Indicates the updated satellite speed status.
[0052] This embodiment converts the first-order differential dynamics equation into this integral form, taking full account of the To the next time point The dynamic changes of the satellite provide a status update mechanism in a continuous time.
[0053] (2.3) The 8th-order Runge-Kutta numerical integration method (i.e., RK8 method) is used to approximate the integral part of the satellite state update equation in integral form by calculating the 8th-order slope, thereby solving the updated satellite position state and velocity state.
[0054] In this embodiment, the 8th-order Runge-Kutta numerical integration method is used to calculate the 8th-order position slope. and 8th order speed slope , approximating the satellite state update equation in the integral form in step 2, where k i s Indicates the location status is updated Step position slope, k i v The speed status is updated Step speed slope.
[0055] In the specific calculation process, Step position slope k i s By satellite at time Velocity vector The calculation formula is as follows: , Where: The first is determined by the 8th order Runge-Kutta numerical integration method. i The time coefficient of the order; Velocity vector is the velocity v of the satellite at time t( t ) and the motion performed by the satellite, i.e., acceleration Calculated, that is = v( t )+ ; The action performed by the satellite at time t is acceleration , which is the action at time t obtained by the continuous time proximal strategy optimization reinforcement learning algorithm in step 1, that is = ; To perform an action The time step is the execution time of the action at time t obtained by the continuous time proximal strategy optimization reinforcement learning algorithm in step 1, = (opt).
[0056] No. Step velocity slope k i v The satellite is calculated by the dynamic equation f( ) Location , satellite at time Speed The calculation formula is as follows: ( ), Where: The time step is , which is the execution time of the action at time t obtained by the continuous time proximal strategy optimization reinforcement learning algorithm in step 1, = (opt); 、 are respectively determined by the 8th-order Runge-Kutta numerical integration method. i The time coefficient of the order; Location is achieved by numerical integration in time The position vector obtained at the moment is ;in, It's in time The velocity vector at the moment can be obtained by the known velocity slope Perform approximate calculations; To perform an action The time step is the execution time of the action at time t obtained by the continuous time proximal strategy optimization reinforcement learning algorithm in step 1, = (opt); speed is the velocity v of the satellite at time t( t ) and the motion performed by the satellite, i.e., acceleration Calculated, that is = v( t )+ The action performed by the satellite at time t is the acceleration is the action at time t obtained by continuous-time proximal policy optimization reinforcement learning algorithm in step 1, that is = ; To perform an action The time step is the execution time of the action at time t obtained by the continuous time proximal strategy optimization reinforcement learning algorithm in step 1, = (opt).
[0057] Finally, by weighting all 8-order slopes by The weighted summation is used to accurately approximate the integral part of the satellite state update equation in integral form. These slopes reflect the changing trend of the state at different time steps. The updated satellite position state and velocity state are obtained from this solution, as shown in the following equation: , Where: is the weight coefficient, which is used to ensure the rationality of contributions at different stages and the accuracy of calculation results.
[0058] The above 8th-order Runge-Kutta method is used to approximate the integral part of the satellite state update equation in the integral form with high precision, thus achieving the satellite position and speed State at time step The precise update after the rotation ensures the accuracy and stability of track control.
[0059] After the calculation of each time step is completed, the time variable is updated, and the above process is continuously looped in the RK8 method until the predetermined simulation end time is reached.
[0060] In this embodiment, the optimal action, i.e., acceleration and optimal execution time, obtained by optimizing the reinforcement learning algorithm based on the continuous time proximal strategy in step 1 are applied to the satellite. The satellite calculates the 8th order position slope based on the optimal action, i.e., acceleration and optimal execution time, combined with the 8th order Runge-Kutta method. and 8th order speed slope , to approximate the satellite state update equation in integral form, thereby solving the updated satellite position state and velocity state to achieve control of the satellite orbit.
[0061] In this embodiment, after the satellite orbit control is performed at each moment, the advantage function is solved by the generalized advantage estimation method according to the performance change of the strategy network. Discount factor , and the objective function of clipping The cropping parameters in , thereby further optimizing the policy network. The specific instructions are as follows: (A) Discount Factor It is mainly responsible for controlling the degree of emphasis on future rewards in the proximal strategy optimization algorithm, thereby affecting the long-term goal positioning of the strategy. Appropriate adjustment of the discount factor can help the strategy better reflect the emphasis on long-term goals or immediate effects. The adjustment formula is as follows: , in: represents the adjusted discount factor; Indicates the discount factor currently used; It is a scaling factor adjusted based on the performance feedback of the policy network, used to fine-tune the discount factor to optimize the weighting of long-term rewards; Is the lower limit of the discount factor, ensuring that the discount factor will not be lower than this value; The upper limit of the discount factor is to ensure that the discount factor does not exceed this value. and Ensure that the adjustment is within a reasonable range to avoid extreme values affecting the stability of the learning process.
[0062] (B) Cropping parameters In the proximal policy optimization algorithm, it plays a key role in limiting the policy update step and trimming parameters. The key to optimization is to prevent excessive fluctuations during the learning process and ensure the smoothness and reliability of the strategy update. The adjustment formula for the clipping parameter is as follows: , in: is the adjusted cropping parameter; is the cropping parameter currently used; It is the adjustment sensitivity factor that determines how performance changes affect the adjustment of the clipping range; represents the observed performance changes of the policy network; and The lower and upper limits of the cropping parameters are respectively to ensure that the cropping parameters will not be lower or higher than the value. and The reasonable variation range of the cutting parameters is defined to ensure stability and effectiveness during the adjustment process.
[0063] Adjust discount factor and cropping parameters It can balance the relationship between the long-term goal and short-term feedback of the policy network, ensuring that the continuous-time proximal policy optimization reinforcement learning algorithm remains efficient and stable under changing task conditions. and cropping parameters , so that the continuous-time proximal policy optimization reinforcement learning algorithm in this embodiment can effectively push the strategy toward optimization while maintaining learning stability, ensuring efficient and stable execution of satellite orbit control under various mission conditions.
[0064] This embodiment, based on the application of a continuous-time framework proximal policy optimization reinforcement learning algorithm, demonstrates significant advancements and numerous advantages. By limiting the magnitude of each policy update, the algorithm ensures the stability and convergence of the policy update process. This is particularly important in satellite orbit control, as orbit adjustments require highly precise and stable control strategies to avoid system instability or orbit deviation caused by excessive adjustments. Furthermore, the algorithm has high sample efficiency and can learn effective control strategies within a limited number of interactions, reducing reliance on large amounts of training data. This is of great significance for the high-dimensional state and action spaces and high sample acquisition costs involved in satellite orbit control.
[0065] Furthermore, satellite orbit control requires fine-tuning in a continuous action space, such as the magnitude and direction of thrust. This embodiment, based on a continuous-time framework proximal policy optimization reinforcement learning algorithm, excels at processing continuous action spaces and can generate smooth and precise control commands to meet the needs of orbit adjustment. The satellite orbit control environment is typically highly uncertain and dynamic, with factors such as external disturbances and fuel consumption. The proximal policy optimization algorithm exhibits robustness, maintaining stable performance in the face of environmental changes and disturbances, and adapting to new environmental conditions through continuous learning.
[0066] This embodiment, based on a continuous-time framework proximal policy optimization reinforcement learning algorithm, demonstrates numerous advancements in satellite orbit control. First, the continuous-time framework is closer to real-world modeling and can more accurately describe the behavior and changes of continuous dynamic systems such as satellites. Compared to the discrete-time framework, the continuous-time architecture allows decisions and updates to be made at any time, without being restricted by a fixed time step. This is particularly important for tasks requiring high-precision time control. At the same time, the control signal in continuous time can be adjusted more smoothly, avoiding the jumps and discontinuities that may occur in discrete time, thereby improving the stability and reliability of the control strategy.
[0067] In terms of mathematical foundations, this embodiment's proximal policy optimization reinforcement learning algorithm, based on a continuous-time framework, can better integrate differential equations and dynamic systems theory from classical control theory, leveraging existing mathematical tools and theories to improve algorithm performance and stability. Continuous-time models often exhibit better properties in theoretical analysis, such as smoothness and differentiability, facilitating the design of more efficient learning algorithms and optimization methods. Furthermore, the continuous-time architecture enables adaptive sampling, adaptively selecting sampling time points based on the system's dynamics, avoiding unnecessary computations when the system changes slowly, and thus more efficiently utilizing computing resources. Continuous-time reinforcement learning also enhances the model's generalization capabilities, enabling better adaptation to changes at different time scales and improving its adaptability and flexibility in diverse environments and tasks. The policy can be fine-tuned at any time point, making the learned policy more flexible and adaptable. Furthermore, the continuous-time architecture more naturally integrates with the laws and phenomena of the physical world, such as dynamic modeling and energy conservation, enhancing the physical feasibility and effectiveness of the control strategy.
[0068] This embodiment, based on the application of a continuous-time proximal policy optimization reinforcement learning algorithm to satellite orbit control, not only overcomes the shortcomings of traditional control methods and existing reinforcement learning methods in dynamic environments, but also achieves higher-precision and smoother orbit control through its stability, efficiency, continuity, and robustness. These advantages make the continuous-time proximal policy optimization algorithm an advanced technology in the field of satellite orbit control, capable of meeting the stringent real-time, adaptability, and efficiency requirements of modern satellite missions.
[0069] The preferred embodiments of the present invention are described in detail above with reference to the accompanying drawings. The embodiments described in the present invention are merely descriptions of the preferred embodiments of the present invention and do not limit the concept and scope of the present invention. The various specific technical features described in the above specific embodiments can be combined in any suitable manner unless there is any contradiction. Such combinations should also be regarded as the contents disclosed in this disclosure as long as they do not violate the concept of the present invention. In order to avoid unnecessary repetition, the present invention will not further describe various possible combinations.
[0070] The present invention is not limited to the specific details of the above-mentioned embodiments. Within the scope of the technical concept of the present invention and without departing from the design concept of the present invention, various modifications and improvements made to the technical solution of the present invention by those skilled in the art should fall within the scope of protection of the present invention. The technical contents for which protection is sought in the present invention have been fully recorded in the claims.
Claims
1. A satellite orbit control method based on continuous-time proximal policy optimization reinforcement learning algorithm, characterized in that: The following steps are involved: Step 1: Based on the continuous-time proximal strategy, the reinforcement learning algorithm is optimized to determine the action taken by the satellite, i.e., the acceleration, and the execution time of the action. The process is as follows: Establish a satellite agent, which defines the satellite agent's position state space and action space. The position state space includes the satellite agent's position state vector in a continuous time frame, and the action space includes the actions taken by the satellite agent in a continuous time frame, i.e., acceleration. A reward function is also established for the satellite agent. Build a policy network, train it using historical satellite trajectory data, and have the trained policy network output recommended actions and execution times to the satellite agent. After receiving the recommended action and execution time output by the trained policy network, the satellite agent evaluates the recommended action and execution time received through the reward function, thereby determining the optimal action, i.e., acceleration, and optimal execution time, which are used as the action, i.e., acceleration, and execution time, taken by the satellite; Step 2: Based on the actions taken by the satellite and the execution time obtained in step 1, the satellite orbit is controlled.
2. The satellite orbit control method based on the continuous-time proximal policy optimization reinforcement learning algorithm according to claim 1 is characterized in that: In step 1, the reward function of the satellite agent is Feedback to satellite agents based on execution time The acceleration magnitude, cost, and state deviation when executing the action, i.e., acceleration a, and the reward function As shown in the following formula: , in: Indicates that the satellite agent is based on execution time After executing the action, i.e. acceleration a, the position state vector s and the target position state vector Deviation between and are the cost coefficient and delay cost coefficient of the satellite agent’s action execution, i.e. acceleration a, respectively. is the magnitude of the action, i.e., acceleration a, which reflects the cost of executing the acceleration; Indicates movement, i.e. acceleration Execution time.
3. The satellite orbit control method based on continuous-time proximal policy optimization reinforcement learning algorithm according to claim 1 is characterized in that: In step 1, the policy network uses a deep neural network , as the deep neural network of the policy network, is composed of multiple layers of fully connected layers, and the output of each fully connected layer serves as the input of the next fully connected layer; represents the weight and bias parameters of the deep neural network, s is the current position state vector of the satellite, is the recommended action output by the policy network, i.e., acceleration, The recommended action output by the policy network is acceleration Execution time.
4. The satellite orbit control method based on the continuous-time proximal policy optimization reinforcement learning algorithm according to claim 3 is characterized in that: Deep Neural Networks The activation function of each fully connected layer in is the ReLU function.
5. The satellite orbit control method based on continuous-time proximal policy optimization reinforcement learning algorithm according to claim 3 is characterized in that: Deep Neural Networks That is, when training the policy network, the position state vector and velocity vector in the historical satellite trajectory data are used as the deep neural network The input of the deep neural network After the multi-layer fully connected layers are calculated in sequence, the recommended action is obtained, i.e., the acceleration , and perform the recommended action i.e. acceleration Execution time .
6. The satellite orbit control method based on continuous-time proximal policy optimization reinforcement learning algorithm according to claim 3 is characterized in that: The deep neural network The objective function during training is based on the advantage function Constructing the objective function for clipping , the objective function of clipping As shown in the following formula: , in: Indicates that the policy network after parameter update is a deep neural network and the old policy network before parameter update, i.e., the deep neural network The selection probability ratio of the output recommended action; represents the currently updated network parameters, and Represents the previous old strategy network parameters; is a clipping parameter used to limit the range of variation of the probability ratio; Used to clip probability ratios To limit the magnitude of policy updates, if Out of range , then Clipping to this range avoids excessive or unstable updates of the strategy; Represents the position state and action in the old policy network Expect the distribution under represents the advantage function, advantage function The calculation formula is as follows: , in: The satellite performs the action, i.e., acceleration In position The action value function under is the value function, representing the position expected returns; During training, the advantage function It is approximated by the generalized advantage estimation method based on empirical data, and the advantage estimate is adjusted by weighting the time series difference residuals at different time steps, thereby approximately solving the advantage function , the calculation formula is as follows: , in: A t express Advantage function at time The generalized advantage estimation method of is used to approximate the solution, that is, Differential estimation of the acceleration value relative to the position value at each moment; ,in Indicates the satellite at the historical moment The reward obtained after performing an action; is the position in the current policy network The estimated value of Expected returns under Is the next position state The estimated value of is the discount factor used to adjust the present value of future rewards; is a smoothing parameter that adjusts the trade-off between bias and variance of the advantage estimate; Minimize the clipping objective function during training , based on the objective function of clipping Gradient ascent method for deep neural networks Parameters in to update.
7. The satellite orbit control method based on continuous-time proximal policy optimization reinforcement learning algorithm according to claim 6 is characterized in that: After step 2, adjust the advantage function by using the generalized advantage estimation method according to the performance changes of the policy network. Discount factor , and the objective function of clipping The cropping parameters in , thereby further optimizing the policy network.
8. The satellite orbit control method based on the continuous-time proximal policy optimization reinforcement learning algorithm according to any one of claims 1 to 7, characterized in that: In step 2, the satellite state update equation is established. Then, the satellite state update equation is solved by combining the 8th-order Runge-Kutta method and the actions and execution times obtained in step 1. The satellite orbit is controlled based on the solution.
9. The satellite orbit control method based on continuous-time proximal policy optimization reinforcement learning algorithm according to claim 8, characterized in that: In step 2, the satellite state update equation in integral form is established, and then the 8th-order Runge-Kutta method is used to approximate the integral part of the integral form of the satellite state update equation based on the action and execution time obtained in step 1, thereby solving the satellite state update equation.
Citation Information
Patent Citations
Satellite avoidance interception method based on deep reinforcement learning and navigation vector field
CN115659788A
Satellite intelligent robust approximate optimal orbit control method based on discrete time reinforcement learning
CN117762022A
Star group orbit pursuit decision-making method based on multi-near-end reinforcement learning
CN119962403A
Metareinforcement learning scheduling method and device for earth observation satellite task planning
CN120355195A
Cited By
Variable-scale satellite group orbit planning method based on graph near-end optimization algorithm
CN121032282A