A rocket landing guidance method and system based on deep reinforcement learning
By employing a rocket landing guidance method based on deep reinforcement learning, a six-degree-of-freedom dynamic model and Markov decision process for the rocket are constructed. The neural network is then trained to generate rocket landing control commands, which solves the problem of multi-factor uncertainty during the flight of the launch vehicle, improves its autonomy and adaptability, and reduces the risk of launch failure.
Patent Information
- Application Number
- CN202310615988.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-05-29
- Publication Date
- 2025-11-28
- Estimated Expiration
- 2043-05-29
AI Technical Summary
Existing launch vehicle guidance and control technologies are insufficient to comprehensively and holistically address the uncertainties of multiple factors during flight, resulting in a high risk of launch failure. Intelligent control technology is expected to enhance autonomy and adaptability.
A rocket landing guidance method based on deep reinforcement learning is adopted. By building a six-degree-of-freedom dynamic model of the rocket, a Markov decision process and a neural network, training and simulation verification are carried out to generate rocket landing flight control commands.
It has achieved autonomous and adaptive intelligent decision-making for rockets, improved successful landing capability and reduced fuel consumption, and enhanced adaptability to uncertainties.
Smart Images

Figure CN116697829B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of guidance and control of launch vehicles, in particular to a rocket landing guidance method and system based on deep reinforcement learning. BACKGROUND
[0002] Launch vehicles have unique properties such as uncertain flight environment, uncertain fault mode, uncertain external disturbance, uncertain self model, and uncertain flight task. After years of development, launch vehicle guidance and control technology has accumulated a number of methods, and has been put into practice in many major projects, effectively dealing with uncertainties in the flight process. However, the systematicness of these methods is not enough, and the ability to deal with multi-factor uncertainties is still insufficient, making it difficult to solve the problem comprehensively and holistically. Intelligent control is expected to provide a systematic and comprehensive solution, and in the history of spaceflight at home and abroad, rocket launch failures have occurred. According to statistics, launch failures of launch vehicles can be remedied by using advanced guidance and control technology to continue or downgrade the task. Therefore, intelligent control technology will be an inevitable choice for future space development, and developing launch vehicle intelligent control technology and creating a learning launch vehicle are effective ways to solve current difficulties. In view of the characteristics of launch vehicle flight control, how to complete efficient and generalizable control law design through offline interactive learning and automatic optimization, and how to use intelligent learning methods for rapid learning and adaptively optimize control law parameters in flight to improve the adaptability of launch vehicle control system to uncertain factors are important research topics. In order to achieve this goal, it is of great significance and value to effectively design a deep reinforcement learning algorithm framework suitable for the characteristics of launch vehicle flight control and realize simulation performance verification and evaluation of launch vehicles based on reinforcement learning.
[0003] In recent years, the rapid development of artificial intelligence has provided a new breakthrough for the realization of intelligent autonomous flight of aircraft. Deep learning mainly realizes the function mapping of data, while reinforcement learning is aimed at Markov decision process, and generates an optimal strategy for global decision-making through continuous interaction and iterative learning with the controlled object. Deep reinforcement learning method, which combines the advantages of both, is suitable for solving motion control problems and is expected to provide a feasible implementation approach for intelligent control methods. SUMMARY
[0004] The present application relates to the field of guidance and control of launch vehicles, in particular to a rocket landing guidance method and system based on deep reinforcement learning.
[0005] The technical solution for achieving the present application is as follows: a rocket landing guidance method based on deep reinforcement learning, comprising the following steps:
[0006] Step 1, according to the rocket six degree of freedom dynamics model, build a rocket landing guidance simulation environment;
[0007] Step 2, based on the rocket six degree of freedom dynamics, establish Markov decision process, including state space, action space, state transition equation and reward function;
[0008] Step 3, according to the deep reinforcement learning algorithm, build neural network;
[0009] Step 4, based on the state space, action space, state transition equation and reward function, through the interaction with the rocket landing guidance environment, train the neural network, obtain the trained neural network model;
[0010] Step 5, calling the trained neural network model for simulation verification;
[0011] Step 6, according to the neural network model after simulation test, generate rocket landing flight control command, complete the landing task of rocket.
[0012] A rocket landing guidance system based on deep reinforcement learning, the system is used for realizing the rocket landing guidance method based on deep reinforcement learning, the system includes environment building module, Markov decision module, algorithm module, training module, simulation test and control module, wherein:
[0013] The environment building module is used for constructing the rocket landing guidance simulation environment;
[0014] The Markov decision module is used for establishing the Markov decision process of rocket landing guidance, including state space, action space, state transition equation and reward function;
[0015] The algorithm module is used for building neural network according to the deep reinforcement learning algorithm;
[0016] The training module is used for training the neural network, obtaining the trained neural network model;
[0017] The simulation test module is used for calling the trained neural network model for simulation verification;
[0018] The control module is used for generating rocket landing flight control command.
[0019] Compared with the prior art, the present application has the following advantages: (1) a deep reinforcement learning PPO algorithm is used to design a deep reinforcement learning program for rocket landing guidance, a neural network is used to fit the mapping relationship between the environment and the agent, and the neural network is trained, so that the rocket can autonomously land using the trained neural network; (2) a six-degree-of-freedom dynamics model of the rocket is established, and a motion equation is applied to design and train the landing guidance model by using deep reinforcement learning and other methods, so as to realize rapid autonomous decision-making and improve the autonomous and adaptive capabilities of the rocket for typical scenarios; (3) a simulation environment is established to explore the ways and methods of applying deep reinforcement learning and other methods to develop rocket landing guidance decision-making technology, a simulation environment model for deep reinforcement learning is constructed, and the efficiency of the algorithm is improved by using the training of the deep reinforcement learning algorithm, so as to realize rapid decision-making, reduce fuel consumption, and improve the autonomous landing capability of the launch vehicle. BRIEF DESCRIPTION OF DRAWINGS
[0020] Figure 1 A flowchart of a rocket landing guidance method based on deep reinforcement learning according to the present application.
[0021] Figure 2 A flowchart of an Actor-Critic algorithm in the specific embodiment of the present application.
[0022] Figure 3 A structure diagram of a policy neural network in the specific embodiment of the present application.
[0023] Figure 4 A structure diagram of a value neural network in the specific embodiment of the present application.
[0024] Figure 5 A change curve diagram of a reward function in the embodiment of the present application.
[0025] Figure 6 A motion trajectory curve diagram of a rocket in the embodiment of the present application.
[0026] Figure 7 A change curve diagram of the acceleration of a rocket in the embodiment of the present application.
[0027] Figure 8 A change curve diagram of the thrust of a rocket in the embodiment of the present application. DETAILED DESCRIPTION
[0028] A rocket landing guidance method based on deep reinforcement learning according to the present application comprises the following steps:
[0029] Step 1: A rocket landing guidance simulation environment is built according to a six-degree-of-freedom dynamics model of the rocket.
[0030] Step 2, based on the six-degree-of-freedom dynamics of the rocket, a Markov decision process is established, including state space, action space, state transition equation and reward function;
[0031] Step 3, according to the deep reinforcement learning algorithm, a neural network is built;
[0032] Step 4, based on the state space, action space, state transition equation and reward function, the neural network is trained by interacting with the rocket landing guidance environment, and a trained neural network model is obtained;
[0033] Step 5, calling the trained neural network model for simulation verification;
[0034] Step 6, according to the neural network model after simulation test, the rocket landing flight control command is generated, and the landing task of the rocket is completed.
[0035] Further, the six-degree-of-freedom dynamics model of the rocket in step 1 is as follows:
[0036] The center of mass dynamics equation of the rocket in the inertial coordinate system is:
[0037]
[0038] In the formula: r is the position vector; v is the velocity vector; m is the mass of the rocket; g is the gravitational acceleration vector; T is the engine thrust vector; D is the aerodynamic drag vector; I sp represents the specific impulse of fuel, g0 represents the average gravitational acceleration at the sea level of the earth; p is the atmospheric density determined by the height; S ref is the reference cross-sectional area of the rocket; C D is the drag coefficient, which is a nonlinear function of the velocity v; Ma is the Mach number, which is determined by the velocity v and the sound speed;
[0039] The control variable is the engine thrust T, and the amplitude satisfies the constraint
[0040] T min ≤||T||≤T max (2) The motion equation of the rocket around the center of mass in the body coordinate system and the quaternion form is:
[0041]
[0042] In the formula: is the component of the attitude solution angular velocity in the body coordinate system, J is the inertia moment vector, ω x , ω y , ω z are the components of the rocket rotation angular velocity in the body coordinate system, M stx , M sty , Mstz , M dx , M dy , M dz , M cx , M cy , M cz are the components of aerodynamic stability moment, aerodynamic damping moment and control moment acting on the rocket on the three axes of the body coordinate system 3 respectively;
[0043] The kinematic equation of the rocket in quaternion form in the body coordinate system is:
[0044]
[0045] In the formula: q0, q1, q2 and q3 are the quaternions of the rocket.
[0046] Further, the inertial coordinate system and the body coordinate system are as follows:
[0047] The definition of the inertial coordinate system is that the origin O G of the inertial coordinate system is taken at the landing point of the rocket; the axis O G X G points to the north, the axis O G Y G points to the east, and the axis O G X G points to the south, the axis O G Y G points to the west, and the axis O G Z G points downward, satisfying the right-hand rule.
[0048] The definition of the body coordinate system is that the body coordinate system is fixed to the rocket, and the origin O T is at the center of mass of the rocket; the axis O T X T is located in the symmetry plane of the rocket, parallel to the axis of the rocket body pointing forward; the axis O T Y T is perpendicular to the symmetry plane of the rocket, i.e. the O T X T Z T plane, pointing to the right; and the axis O T Z T is located in the symmetry plane of the rocket, perpendicular to the X T axis pointing downward to the belly of the rocket.
[0049] Further, the Markov decision process based on the six-degree-of-freedom dynamics of the rocket in step 2 is established, including the state space, the action space, the state transition equation and the reward function, which are as follows:
[0050] The state space is:
[0051] S = [r, v, q0, q1, q2, q3, ω x , ω y , ω z , m] T (5)
[0052] where: r is the position vector; v is the velocity vector; m is the mass of the rocket; q0, q1, q2, q3 are the rocket quaternions, ω x , ω y , ω z are the components of the rocket angular velocity in the body coordinate system 3 axes respectively;
[0053] The action space is:
[0054] A = [δ y , δ z , ||T||] T (6)
[0055] where: δ y , δ z are the directions of the thrust, ||T|| is the magnitude of the engine thrust, and the value range of each action quantity is:
[0056]
[0057] The state transition equation is:
[0058]
[0059]
[0060]
[0061] The reward function design is divided into two parts: process cumulative return and terminal reward return, where the process cumulative return R1 is expressed as:
[0062]
[0063] where: a is the acceleration, a targ is the target acceleration, is the attitude angle, t go is the remaining flight time;
[0064] The terminal reward return R2 is expressed as:
[0065]
[0066] where: R is the terminal state reward, R r is the terminal position reward, R v is the terminal velocity reward, x is the rocket flight altitude, r targThe landing radius of the rocket;
[0067] The total reward is represented as:
[0068] reward = R1 + R2 (13)
[0069] Further, the deep reinforcement learning algorithm in step 3 is specifically the Proximal Policy Optimization (PPO) algorithm based on the Actor-Critic architecture.
[0070] Further, the neural network in step 3 includes an evaluation neural network and a policy neural network.
[0071] Both the policy neural network and the evaluation neural network are four-layer fully connected layers with 256, 256, 128, and 64 hidden layer neurons, respectively, using Relu as the activation function, with an initial value of 0.1 for the step size λ and a discount factor of 0.99.
[0072] Further, the neural network is built according to the deep reinforcement learning algorithm in step 3, specifically as follows:
[0073] Step 3.1, build an evaluation neural network to update the evaluation of each state-action pair according to the reward information at this moment, input the current and next state, and output the corresponding state-action pair evaluation value respectively;
[0074] Step 3.2, the policy neural network updates the rocket landing guidance policy according to the evaluation neural network, so that the selected rocket landing guidance policy always moves in the direction of high evaluation, inputting the current state of the environment including the position, velocity, mass, four elements, and angular velocity parameters of the rocket, and outputting the strategy that the rocket should take;
[0075] Step 3.3, design a loss function according to the reward feedback from the environment, which is used to update the evaluation neural network and the policy neural network.
[0076] Further, the neural network is trained based on the state space, action space, state transition equation, and reward function in step 4, and the trained neural network model is obtained by interacting with the rocket landing guidance environment, specifically as follows:
[0077] Step 4.1, initialize the parameters of the policy neural network and the evaluation neural network;
[0078] Step 4.2, initialize the state space to obtain the current state s t ;
[0079] Step 4.3, the rocket landing guidance simulation environment selects the behavior a based on the action space according to the strategy output by the strategy neural network t , the state transition equation is executed to obtain the next step state s t+1 , the reward r is obtained according to the reward function t , the advantage function A of this step is calculated t and saved;
[0080] Step 4.4, according to the loss function of the PPO algorithm, the gradient descent method is used to update the parameters of the strategy neural network and the parameters of the evaluation neural network;
[0081] Step 4.5, the strategy neural network outputs a new strategy;
[0082] Step 4.6, repeat steps 4.2-4.5 6x10 5 times to complete the training of the neural network model, and save the trained neural network model.
[0083] Further, the neural network model after simulation test is used to generate rocket landing flight control instructions to complete the landing task of the rocket, and the specific steps are as follows:
[0084] The neural network model after simulation test outputs the engine thrust size and angle of the rocket, and the rocket adjusts the guidance strategy according to these control quantities to achieve successful landing.
[0085] The application also provides a rocket landing guidance system based on deep reinforcement learning, which is used to realize the rocket landing guidance method based on deep reinforcement learning, and the system comprises an environment building module, a Markov decision module, an algorithm module, a training module, a simulation test and a control module, wherein:
[0086] The environment building module is used to build a rocket landing guidance simulation environment;
[0087] The Markov decision module is used to establish a Markov decision process for rocket landing guidance, including a state space, an action space, a state transition equation and a reward function;
[0088] The algorithm module is used to build a neural network according to a deep reinforcement learning algorithm;
[0089] The training module is used to train the neural network to obtain a trained neural network model;
[0090] The simulation test module is used to call the trained neural network model for simulation verification;
[0091] The control module is used to generate rocket landing flight control instructions.
[0092] The application will be further described in detail below with reference to the drawings and specific embodiments.
[0093] Embodiment
[0094] The rocket landing guidance system based on deep reinforcement learning comprises an environment building module, a Markov decision module, an algorithm module, a training module, a simulation test and control module.
[0095] The environment building module is used to build a rocket landing guidance simulation environment.
[0096] The Markov decision module is used to establish a Markov decision process of rocket landing guidance, including a state space, an action space, a state transition equation and a reward function.
[0097] The algorithm module is used to build a neural network according to a deep reinforcement learning algorithm.
[0098] The training module is used to train the neural network to obtain a trained neural network model.
[0099] The simulation test module is used to call the trained neural network model for simulation verification.
[0100] The control module is used to generate a rocket landing flight control instruction.
[0101] As shown in the figure, a rocket landing guidance method based on deep reinforcement learning comprises the following steps: Figure 1 Step 1: The environment building module builds a rocket landing guidance simulation environment according to rocket six-degree-of-freedom dynamics.
[0102] Step 2: The Markov decision module establishes a Markov decision process based on rocket six-degree-of-freedom dynamics, including a state space, an action space, a state transition equation and a reward function.
[0103] Step 3: The algorithm module builds a neural network according to a deep reinforcement learning algorithm.
[0104] Step 4: The training module trains the neural network based on the state space, the action space, the state transition equation and the reward function by interacting with the rocket landing guidance environment to obtain a trained neural network model.
[0105] Step 5: The simulation test module calls the trained neural network model for simulation verification.
[0106] Step 6: The control module generates a rocket landing flight control instruction according to the simulated neural network model to complete the landing task of the rocket.
[0107]
[0108] Further, the environment building module of step 1 builds a rocket landing guidance simulation environment according to the rocket six-degree-of-freedom dynamics, specifically as follows:
[0109] First, the dynamics model of the rocket is established, and various forces acting on it are analyzed to establish the motion and dynamics model under the complex force field environment of the launch vehicle, which is the basis for subsequent research on the model, specifically:
[0110] The mass center dynamics equation of the rocket in the inertial coordinate system is:
[0111]
[0112] In the formula: r is the position vector; v is the velocity vector; m is the mass of the rocket; g is the gravity acceleration vector; T is the engine thrust vector; D is the aerodynamic drag vector; I sp represents the specific impulse of the fuel, g0 represents the average gravity acceleration at the sea level of the earth; p is the atmospheric density determined by the height; S ref is the reference cross-sectional area of the rocket; C D is the drag coefficient, which is a nonlinear function of the speed v; Ma is the Mach number, which is determined by the speed v and the sound speed.
[0113] The control variable is the engine thrust T, and the amplitude satisfies the constraint
[0114] T min ≤||T||≤T max (2)
[0115] The motion equation of the rocket around the mass center in the body coordinate system and the kinematics equation in the quaternion form are:
[0116]
[0117] In the formula: is the attitude solution angular velocity, J is the inertia moment vector, ω x , ω y , ω z are the components of the rotation angular velocity of the rocket in the body coordinate system, M sty , M stz , M dx , M dy , M dz , M cy , M cz are the components of the aerodynamic stability moment, aerodynamic damping moment and control moment acting on the rocket in the body coordinate system.
[0118] The kinematics equation of the rocket in the quaternion form in the body coordinate system is:
[0119]
[0120] In the formula: q0, q1, q2 and q3 are the quaternions of the rocket.
[0121] Furthermore, the inertial coordinate system and the rocket body coordinate system are specifically as follows:
[0122] The definition of an inertial coordinate system is: the origin O of the inertial coordinate system. G Take it at the rocket landing point; axis O G X G and axis O G Y G In the horizontal plane, axis O G X G Pointing due north, axis O G Y G Pointing due east, axis O G Z G The right-hand rule applies, and the vertical is downward.
[0123] The rocket body coordinate system is defined as follows: the rocket body coordinate system is fixed to the rocket, and the origin of the coordinate system is at the rocket's center of mass O. T ; axis O T X T Located within the rocket's plane of symmetry, parallel to the rocket's axis and pointing forward; axis O T Y T Perpendicular to the rocket's plane of symmetry, i.e., O T X T Z T Plane, pointing to the right; axis O T Z T Located in the plane of symmetry of the rocket, perpendicular to X T The axis points downwards toward the belly of the rocket.
[0124] Furthermore, step 2 involves establishing a Markov decision process based on the six-degree-of-freedom dynamics of a rocket, including the state space, action space, state transition equations, and reward function, as detailed below:
[0125] The state space is:
[0126] S=[r, v, q0, q1, q2, q3, ω x ω y ω z m] T (5)
[0127] Where: r is the position vector; v is the velocity vector; m is the rocket mass; q0, q1, q2, q3 are rocket quaternions, ω x ω y ω z These represent the components of the rocket's rotational angular velocity along the three axes of the rocket's coordinate system.
[0128] The action space is:
[0129] A = [δ y , δ z , ||T||] T (6)
[0130] wherein δ y , δ z is the direction of the thrust, and ||T|| is the magnitude of the thrust of the engine. The value range of each action quantity is:
[0131]
[0132] The state transition equation is:
[0133]
[0134]
[0135]
[0136] The reward function design is divided into two parts: process cumulative return and terminal reward return. The process cumulative return R1 is expressed as:
[0137]
[0138] wherein a is the acceleration, a targ is the target acceleration, is the attitude angle, t go is the remaining flight time.
[0139] The terminal reward return R2 is expressed as:
[0140]
[0141] wherein R is the terminal state reward, R r is the terminal position reward, R v is the terminal speed reward, x is the rocket flight height, r targ is the landing radius of the rocket.
[0142] The total reward is expressed as:
[0143] reward = R1 + R2 (13)
[0144] Further, the algorithm module in step 3 is built according to the PPO algorithm, and the neural network is built as follows:
[0145] Step 3.1: Build an evaluation neural network to update the evaluation of each state-action pair based on the reward information at this moment. Input the current and next time-step states and output the corresponding state-action pair evaluation values respectively.
[0146] Step 3.2: Policy Neural Network. Update the rocket landing guidance strategy based on the evaluation neural network, so that the selected rocket landing guidance strategy always moves in the direction of the larger evaluation. The input environment is the current state, including the rocket's position, velocity, mass, quaternion, attitude calculation angular velocity parameters, and the output is the strategy that the rocket should take.
[0147] Step 3.3: Design a loss function based on the rewards from environmental feedback to update the valuation neural network and the policy neural network.
[0148] This method employs the classic Actor-Critic architecture in deep reinforcement learning, and its basic network structure is as follows: Figure 2 As shown.
[0149] After the observed state of the environment is input into the neural network, the parameters are updated. The Actor network generates the corresponding policy and produces the corresponding action output; the Critic network evaluates the current policy through the advantage function.
[0150] The neural network mentioned includes a valuation neural network and a policy neural network, combined with... Figure 3 , Figure 4 Both the policy neural network and the valuation neural network are four fully connected layers with 256, 256, 128 and 64 hidden layer neurons, respectively. ReLU is used as the activation function, the initial value of the step size λ is set to 0.1 and the discount factor is set to 0.99.
[0151] Furthermore, the training module described in step 4 trains the neural network based on the rocket's six-degree-of-freedom dynamics model, state space, action space, state transition equation, and reward function to obtain a trained neural network model, as detailed below:
[0152] Step 4.1: Initialize the parameters of the policy neural network and the evaluation neural network;
[0153] Step 4.2: Initialize the state space to obtain the current state s. t ;
[0154] Step 4.3: In the rocket landing guidance simulation environment, the strategy output by the strategy neural network is used to select behavior a based on the action space. t Execute the state transition equations (1) to (4) to obtain the next state s. t+1 The reward r is obtained according to the reward function. t, calculate the advantage function of this step and save it;
[0155] Step 4.4, update the parameters of the policy neural network and the parameters of the evaluation neural network using gradient descent method according to the loss function of the PPO algorithm;
[0156] Step 4.5, the policy neural network outputs a new strategy;
[0157] Step 4.6, repeat steps 4.2-4.5 6x10 5 times to complete the training of the neural network model, and save the trained neural network model.
[0158] Step 5, the simulation test module calls the trained neural network model for simulation verification;
[0159] Step 6, the control module generates rocket landing flight control instructions according to the neural network model after simulation test, and completes the landing task of the rocket.
[0160] The reward function convergence result of simulation is shown in Figure 5 , and the reward function can be converged. Figure 5 The rocket motion trajectory is shown in Figure 6 . Figure 7 The rocket acceleration curve is shown in Figure 8 , and the rocket thrust size change is shown. The results obtained by simulation show that the terminal position accuracy of the reinforcement learning guidance strategy is 5m, the velocity accuracy is 2m / s, and the fuel consumption is 4135kg, and the rocket autonomous landing is realized.
[0161] From Figure 6 , Figure 7 and Figure 8 , it can be seen that the present application is based on the deep reinforcement learning PPO algorithm, and a deep reinforcement learning program for rocket landing guidance is designed, the mapping relationship between the environment and the agent is fitted using a neural network, and the neural network is trained, so that the rocket can use the trained neural network to land autonomously; in addition, the present application researches to establish a six-degree-of-freedom dynamics model of the rocket, and applies deep reinforcement learning and other methods to carry out the design and training of the landing guidance model, realizes rapid autonomous decision-making, and improves the autonomous and adaptive ability of the rocket for typical scenarios.
Claims
1. A rocket landing guidance method based on deep reinforcement learning, characterized in that, The steps are as follows: Step 1: According to the six-degree-of-freedom dynamics model of the rocket, a rocket landing guidance simulation environment is established; The six-degree-of-freedom dynamics model of the rocket is as follows: The center-of-mass dynamics equation of the rocket in the inertial coordinate system is as follows: where r is the position vector; v is the velocity vector; m is the mass of the rocket; g is the gravitational acceleration vector; T is the engine thrust vector; D is the aerodynamic drag vector; I sp represents the specific impulse of the fuel, g0represents the average gravitational acceleration at sea level on Earth; p is the atmospheric density, which is determined by the altitude; S ref is the reference cross-sectional area of the rocket; C D is the drag coefficient, which is a nonlinear function of the velocity v; Ma is the Mach number, which is determined by the velocity v and the speed of sound; The control variable is the engine thrust T, and the amplitude satisfies the constraint T min ≤||T||≤T max (2) The motion equation of the rocket around the center of mass in the body coordinate system is as follows: In the formula: are the components of the angular velocity of the attitude solution in the three axes of the missile body coordinate system, J is the inertia moment vector, ω x , ω y , ω z are the components of the angular velocity of the rocket in the three axes of the missile body coordinate system, M stx , M sty , M stz , M dx , M dy , M dz , M cx , M cy , M cz are the components of the aerodynamic stability moment, the aerodynamic damping moment and the control moment acting on the rocket in the three axes of the missile body coordinate system; The motion equation of the rocket in the body coordinate system is as follows: In the formula: q0, q1, q2 and q3 are the quaternions of the rocket; Step 2: Based on the six-degree-of-freedom dynamics of the rocket, a Markov decision process is established, including state space, action space, state transition equation and reward function, which is as follows: The state space is as follows: S = [r, v, q0, q1, q2, q3, ω x , ω y , ω z , m] T (5) where: r is the position vector; v is the velocity vector; m is the mass of the rocket; q0, q1, q2, q3 are the rocket quaternions, ω x , ω y , ω z are the components of the rocket angular velocity in the body frame, respectively. The action space is as follows: A = [δ y , δ z , ||T||] T (6) wherein δ y , δ z is the direction of the thrust, ||T|| is the magnitude of the thrust of the engine, and the range of values of each of the motion quantities is: The state transition equation is as follows: The reward function design is divided into two parts: process cumulative return and terminal reward return, wherein the process cumulative return R1 is represented as: where: a is the acceleration, a targ is the target acceleration, is the attitude angle, t go is the remaining flight time; The terminal reward return R2 is represented as: where: R is the terminal state reward, R r is the terminal position reward, R v is the terminal velocity reward, x is the rocket's altitude, r targ is the rocket's landing radius; Then the total reward is represented as: reward=R1+R2 (13) Step 3: According to the deep reinforcement learning algorithm, a neural network is established; Step 4: Based on the state space, action space, state transition equation and reward function, the neural network is trained by interacting with the rocket landing guidance environment to obtain a trained neural network model; Step 5: The trained neural network model is called for simulation verification; Step 6: According to the neural network model after simulation test, the rocket landing flight control command is generated to complete the landing task of the rocket.
2. The deep-reinforcement learning based rocket landing guidance method according to claim 1, wherein, The inertial coordinate system and the body coordinate system are as follows: The definition of the inertial coordinate system is: the origin of the inertial system O G Take the rocket landing point; axis O G X G And axis O G Y G In the horizontal plane, axis O G X G Pointing to the north, axis O G Y G Pointing to the east, axis O G Z G Satisfy the right-hand rule plumb down; The definition of the arrow body coordinate system is that the arrow body coordinate system is fixed to the rocket, the coordinate origin is at the rocket mass center O T ; the axis O T X T is located in the rocket symmetry plane, is parallel to the arrow body axis and points forward; the axis O T Y T is perpendicular to the rocket symmetry plane, that is, O T X T Z T plane, points to the right; the axis O T Z T is located in the rocket symmetry plane, is perpendicular to the X T axis and points downward to the rocket belly.
3. The deep reinforcement learning-based rocket landing guidance method according to claim 1, wherein, The deep reinforcement learning algorithm in step 3 is specifically a proximal policy optimization algorithm based on the Actor-Critic architecture.
4. The deep reinforcement learning-based rocket landing guidance method according to claim 3, wherein, The neural network in step 3 includes an evaluation neural network and a policy neural network. Both the policy neural network and the evaluation neural network are four-layer fully connected layers, with hidden layer neuron numbers of 256, 256, 128 and 64 respectively, using Relu as the activation function, with the initial value of the step size λ set to 0.1 and the discount factor set to 0.
99.
5. The deep reinforcement learning-based rocket landing guidance method according to claim 4, wherein, According to the deep reinforcement learning algorithm in step 3, the neural network is established as follows: Step 3.1: Build an evaluation neural network to update the evaluation of each state-action pair according to the reward information at this moment, input the current and next state, and output the corresponding state-action pair evaluation value respectively; Step 3.2: The policy neural network updates the rocket landing guidance strategy according to the evaluation neural network, so that the rocket landing guidance strategy selected each time always moves towards the direction of high evaluation, and the input is the current state of the environment, including the position, velocity, mass, quaternion, and angular velocity parameters of the rocket, and the output is the strategy that the rocket should take; Step 3.3: According to the reward feedback from the environment, a loss function is designed to update the evaluation neural network and the policy neural network.
6. The deep reinforcement learning-based rocket landing guidance method according to claim 5, wherein, According to the state space, action space, state transition equation and reward function in step 4, the neural network is trained by interacting with the rocket landing guidance environment to obtain a trained neural network model, which is as follows: Step 4.1: Initialize the parameters of the policy neural network and the evaluation neural network; Step 4.2, initialize the state space to the current state s t ; Step 4.
3. The rocket landing guidance simulation environment selects the action a based on the action space according to the policy output by the policy neural network t , executes the state transition equation to obtain the next state s t+1 , obtains the reward r according to the reward function t , calculates the advantage function A of this step t and saves it; Step 4.4, updating the parameters of the policy neural network and the parameters of the value neural network using gradient descent method according to the loss function of the PPO algorithm; Step 4.5, the policy neural network outputs a new strategy; Step 4.6, repeat 6x10 5 Sub-step 4.2~step 4.5, the training of the neural network model is completed, and the trained neural network model is saved.
7. The deep reinforcement learning-based rocket landing guidance method according to claim 6, wherein, According to the neural network model after simulation test, the rocket landing flight control command is generated to complete the landing task of the rocket, and the specific steps are as follows: The neural network model after simulation test outputs the engine thrust size and angle of the rocket, and the rocket adjusts the guidance strategy according to these control quantities to achieve successful landing.
8. A deep reinforcement learning based rocket landing guidance system, characterized in that, The system is used to realize the rocket landing guidance method based on deep reinforcement learning in any one of claims 1-7, and the system comprises an environment building module, a Markov decision module, an algorithm module, a training module, a simulation test module and a control module, wherein: The environment building module is used to build a rocket landing guidance simulation environment; The Markov decision module is used to establish a Markov decision process for rocket landing guidance, including state space, action space, state transition equation and reward function; The algorithm module is used to build a neural network according to the deep reinforcement learning algorithm; The training module is used to train the neural network to obtain a trained neural network model; The simulation test module is used to call the trained neural network model for simulation verification; The control module is used to generate a rocket landing flight control command.
Citation Information
Patent Citations
Intelligent control method for vertical recovery of carrier rockets based on deep reinforcement learning
CN109343341A