Rocket landing control method based on multi-agent reinforcement learning

Through the rocket landing control method based on multi-agent reinforcement learning, the stability and convergence problems of traditional launch vehicle vertical landing control methods are solved, and intelligent control of the launch vehicle vertical landing process is realized, which simplifies the control logic and improves the stability and efficiency of the model.

CN120278007AActive Publication Date: 2025-07-08NAT UNIV OF DEFENSE TECH
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
CN202510343532.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-21
Publication Date
2025-07-08
Estimated Expiration
2045-03-21

AI Technical Summary

Technical Problem

In the prior art, the vertical landing control method of launch vehicles adopts traditional open-loop and closed-loop control, which requires a lot of ground interview experience, designs complex control logic, and poor stability and convergence of model training.

Method used

The rocket landing control method based on multi-agent reinforcement learning is adopted. By establishing a rocket engine and vertical landing model, state space, action space and reward functions are defined, and the MARL controller is designed and trained using the improved LCO-MADDPG algorithm to realize intelligent control of the vertical landing process of the launch vehicle.

Benefits of technology

Without complex ground interview experience and control logic, the model training stability and convergence are better, realizing nonlinear control of the vertical landing process of the launch vehicle and improving the intelligence level of control.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120278007A_ABST
    Figure CN120278007A_ABST
Patent Text Reader

Abstract

The invention discloses a rocket landing control method based on multi-agent reinforcement learning in the technical field of rocket landing control methods, and the method comprises the following steps: building a rocket engine model, and collecting and outputting state parameters necessary in the working process of a liquid rocket engine; establishing a carrier rocket vertical landing simulation model, and collecting and outputting necessary parameters in a carrier rocket vertical landing process; defining a state space, an action space and a reward function in the starting process of the engine; on the basis of an MADDPG algorithm, improvement is carried out, and an LCO-MADDPG algorithm is realized; and designing, training and evaluating the MARL model. According to the rocket landing control method, through an intelligent control method, complex control logic does not need to be designed, and nonlinear control over the vertical landing process of the carrier rocket is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of rocket landing control methods, and in particular, to a rocket landing control method based on multi-agent reinforcement learning. Background Art

[0002] With the development of artificial intelligence technology, multi-agent reinforcement learning (MARL) technology is becoming a new research hotspot. By decomposing control tasks into cooperative agents such as thrust adjustment and state stabilization, MARL can process high-dimensional state spaces in parallel, and its distributed decision-making mechanism enables the system to handle non-linear optimization problems. Therefore, multi-agent reinforcement learning provides a new idea and technical approach for the control of the vertical landing process of launch vehicles.

[0003] The vertical landing technology of launch vehicles is one of the core key technologies for building a reusable space transportation system. The high launch cost and resource waste problems of traditional disposable launch vehicles have prompted the global space community to explore the technical path of recycling and reuse since the 1990s. The vertical landing of a launch vehicle can achieve the recovery of the sub-stage structure by precisely controlling the re-entry trajectory of the rocket sub-stage, dynamic deceleration, and landing leg buffering. According to SpaceX engineering data, after a reusable medium-sized launch vehicle is reused 10 times, the single launch cost can be reduced by 80%. However, this technology involves interdisciplinary problems such as high-speed re-entry aerodynamic-thermodynamic coupling, transonic thrust vector control, and terminal high-precision soft landing. Its dynamic equation presents strong non-linearity, time-varying characteristics, and multi-constraint coupling characteristics, posing a severe challenge to traditional guidance, navigation, and control methods. Currently, the landing control method of launch vehicle vertical landing adopts traditional open-loop and closed-loop control methods, which need to face complex environments, require a large amount of ground test experience, need to design complex control logics, and have poor stability and convergence in model training. Summary of the Invention

[0004] The purpose of the present invention is to provide a rocket landing control method based on multi-agent reinforcement learning, which solves the technical problems that the existing landing control method of launch vehicle vertical landing adopts traditional open-loop and closed-loop control methods, needs to face complex environments, requires a large amount of ground test experience, needs to design complex control logics, and has poor stability and convergence in model training. Through an intelligent control method, it is not necessary to design complex control logics, and non-linear control of the vertical landing process of launch vehicles is realized.

[0005] In order to achieve the above purpose, the technical solution of the present invention is as follows:

[0006] The present invention provides a rocket landing control method based on multi-agent reinforcement learning, including the following steps:

[0007] S1. Establish a rocket engine model, collect and output the necessary state parameters during the operation of the liquid rocket engine;

[0008] S2. Establish a vertical landing simulation model for the launch vehicle, collect and output the necessary parameters during the vertical landing process of the launch vehicle;

[0009] S3. Define the state space, action space, and reward function during the engine startup process;

[0010] S4. Based on the MADDPG algorithm, make improvements to implement the LCO - MADDPG algorithm;

[0011] S5. Design, train, and evaluate the MARL model.

[0012] Furthermore, the S1 includes the following steps:

[0013] S11. Use commercial simulation software or programming language to establish a liquid rocket engine model to be analyzed;

[0014] S12. Collect the necessary state parameters during the operation of the liquid rocket engine, and enable the variables of the state parameters to be analyzed to be output.

[0015] Furthermore, the S2 includes the following steps:

[0016] S21. Use the same commercial simulation software or programming language as that used in step S11 to build a simulation model for the vertical landing process of the launch vehicle;

[0017] S22. Collect the necessary parameters during the vertical landing process of the launch vehicle, and enable the variables of the necessary parameters to be output.

[0018] Furthermore, the S3 includes the following steps:

[0019] S31. Define the state space during the engine startup process;

[0020] It is necessary to define the state space of each agent, and the set state space is as follows:

[0021] S agent =[MR GG ,P GG ,T GG ,MR CC ,P CC ,T CC ,F,H,V,a,Pos VGO ,Pos VGF ,Pos VCF

[0022] Among them, MR GG ,P​GG , T GG , MR CC , P CC , T CC , F are the mixture ratio of the engine gas generator, the chamber pressure of the gas generator, the temperature of the gas generator, the mixture ratio of the thrust chamber, the chamber pressure of the thrust chamber, the temperature of the thrust chamber, and the thrust magnitude respectively. H, V, a are the height, speed, and acceleration of the launch vehicle, and Pos VGO , Pos VGF , Pos VCF respectively represent the oxidizer valve of the gas generator, the fuel valve of the gas generator, and the fuel valve of the combustion chamber. The state space is normalized with the steady-state reference value;

[0023] S32. Define the action space during the engine startup process;

[0024] The action space A of the Agent consists of the opening degrees of three valves

[0025] A agent = [Pos VGO , Pos VGF , Pos VCF

[0026] At each time step, the MARL Agent receives the environmental observation results and sends control signals to the control valves of the engine;

[0027] S33. Define the reward function during the engine startup process.

[0028] Furthermore, the S33 includes the following steps:

[0029] S331. The rewards for training MARL and evaluating the startup timing consist of the following different parts:

[0030] Reward = R engine + R landing

[0031] R engine = r e1 + r e2 + r e3

[0032] The first reward is:

[0033]

[0034] where ε i ∈ [MR GG , P GG , T GG , MR CC , P CC , T​CC For the rewards close to the target value, the maximum value of each reward component in this item is clipped by 0.2 to improve the cumulative rewards during the training balance startup and steady state;

[0035] The second reward is:

[0036]

[0037] Here it is assumed that once the engine starts, it will not be shut down again. Therefore, after the engine generates thrust, the engine thrust should not be 0 later to prevent the engine from shutting down during landing;

[0038] The third reward is:

[0039]

[0040] Among them, Act i ∈[Pos VGO , Pos VGF , Pos VCF respectively represent the opening degrees of the three valves of the engine, encouraging to try to open the valves before the engine ignition; S represents the change in the valve position between the two steps before and after the valve, used to punish the reciprocating action of the valve. Through this item, the oscillation caused by the agent frequently actuating the valve during the vertical descent can be suppressed;

[0041] S332.

[0042] R landing =r l1 +r l2 +r l3

[0043] r l1 =0.2·exp(-abs(V) / 10)+0.2·(step max -step current ) / step max

[0044] Among them, step max is the maximum number of steps in each training episode, step current is the current training step. The exponential form of the reward is adopted to guide the agent to decelerate the launch vehicle;

[0045] In addition, step is introduced into the reward function to promote the agent to complete the landing of the launch vehicle in a shorter time,

[0046] r l2 =0.2·exp(H / 50)+0.2·(step max -step current) / step max

[0047] Similar to the above formula, its purpose is to enable the agent to complete a faster and better landing.

[0048]

[0049] After the launch vehicle successfully lands, the agent obtains a huge reward, encouraging the strategy of successful landing to train more strategies for successful landing.

[0050] Furthermore, the S4 includes the following steps:

[0051] S41. The LCO-MADDPG algorithm is improved and updated 10 times during the training process.

[0052] S42. The LCO-MADDPG algorithm is improved and an annealing learning rate is used.

[0053] Furthermore, the S5 includes the following steps:

[0054] S51. Combining the characteristics of the engine system and the nonlinearity of the vertical landing stage, a MARL controller based on the LCO-MADDPG algorithm is designed, trained, and evaluated.

[0055] S52. Using the defined state space as the input of the controller, the MARL agent, according to the trained strategy, controls the actuation of the valve during the engine startup process to enable each agent to obtain the maximum reward.

[0056] Adopting the above technical solutions, the present invention has the following advantages:

[0057] The present invention provides a rocket landing control method based on multi-agent reinforcement learning. By determining the state space, action space, and reward function on the launch vehicle engine and vertical landing model, using the LCO-MADDPG algorithm, the MARL control is designed, trained, and evaluated for controlling the vertical landing process of the launch vehicle. The rocket landing control method based on multi-agent reinforcement learning of the present invention realizes the intelligent control of the vertical landing process of the launch vehicle. Compared with the traditional open-loop and closed-loop control methods, the method of the present invention does not require a large amount of ground test run experience, does not require the design of complex control logic, can achieve complex goals by designing a suitable reward function, and compared with the MADDPG algorithm, the rocket landing control method based on multi-agent reinforcement learning of the present invention has better stability and convergence in model training for the control problem of the vertical landing process of the launch vehicle. Through the intelligent control method, without designing complex control logic, the nonlinear control of the vertical landing process of the launch vehicle is realized. Description of the Drawings

[0058] Figure 1 This is the overall flowchart of the rocket landing control method based on multi-agent reinforcement learning of the present invention;

[0059] Figure 2 It is a system diagram of a gas generator staged combustion cycle engine;

[0060] Figure 3 It is a diagram of the altitude change of the vertical landing simulation model of a launch vehicle;

[0061] Figure 4 It is a diagram of the speed change of the vertical landing simulation model of a launch vehicle;

[0062] Figure 5 It is a diagram of the thrust change of the gas generator during the start-up process based on open-loop control. Detailed Implementation Manner

[0063] The technical solution of the present invention will be specifically described below in conjunction with the drawings of the specification. It should be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variation thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements but also other elements not expressly listed, or also includes elements inherent to such process, method, article or device.

[0064] Figure 1 Shows the overall flowchart of the rocket landing control method based on multi-agent reinforcement learning of the present invention. A rocket landing control method based on multi-agent reinforcement learning is specifically as Figure 1 shown, and includes the following steps:

[0065] S1. Establish a rocket engine model, collect and output the state parameters necessary during the operation of a liquid rocket engine;

[0066] Among them, S1 includes the following specific steps:

[0067] S11. Use commercial simulation software or programming languages to establish a liquid rocket engine model to be analyzed; Commercial simulation software (such as Matlab-Simulink, Amesim, and EcosimPro, etc.) or programming languages (such as Python, C, or C++, etc.) can be used to establish the engine model to be analyzed.

[0068] S12. Collect the necessary state parameters during the operation of the liquid rocket engine so that the variables of the state parameters to be analyzed can be output.

[0069] In a specific embodiment, based on the Matlab-simulink platform, a model of a certain type of staged combustion cycle liquid engine of a gas generator was built, specifically as Figure 2 shown. Among them, RP1 represents fuel, LOX represents oxidizer, VCF represents the combustion chamber fuel valve, VRF represents the flow regulating valve, VGF represents the gas generator fuel valve, VGO represents the gas generator oxidizer valve, VGP represents the gas bypass valve, IGN represents the ignition device, GG represents the gas generator, CC represents the combustion chamber, and NE represents the nozzle.

[0070] S2. Establish a vertical landing simulation model of the launch vehicle, collect and output the necessary parameters during the vertical landing process of the launch vehicle;

[0071] S2 includes the following specific steps:

[0072] S21. The same as the commercial simulation software or programming language used in step S11, commercial simulation software (such as Matlab-Simulink, Amesim, and EcosimPro, etc.) or programming languages (such as Python, C, or C++, etc.) can be used to build a simulation model of the vertical landing process of the launch vehicle;

[0073] S22. Collect the necessary parameters during the vertical landing process of the launch vehicle so that the variables of the necessary parameters can be output.

[0074] In a specific embodiment, based on a python script, a simulation model of the vertical descent process of the launch vehicle was built. The specific content of this simulation model includes randomly setting the initial velocity and initial height of the launch vehicle from the range of the initial height and velocity at which the launch vehicle can achieve a safe landing; combining the thrust of the launch vehicle, assuming the diameter of the launch vehicle hull as a reference for the windward area calculation of air resistance; considering the variation of gravitational acceleration and air density with the height of the launch vehicle; considering that the mass of the launch vehicle will change with the consumption of propellant, and the changes in the height and velocity of the launch vehicle in the obtained vertical landing simulation model are specifically as Figure 3 and Figure 4 shown.

[0075] S3. Define the state space, action space, and reward function during the engine startup process;

[0076] According to the characteristics of the engine and the vertical landing process of the launch vehicle, select the state space, action space, and reward function for MARL training. The state space mainly includes engine state parameters, the height, speed, acceleration, etc. of the launch vehicle. The advantage of taking the engine state parameters as part of the state space is that since the vertical landing process is non-convex, MARL can better output the required thrust based on the engine state and the state of the launch vehicle, while preventing the engine state from exceeding the safety threshold. The action space mainly includes the valves of different engines during the vertical descent process; the reward function mainly includes the safety threshold of the engine state during the landing process and the hard constraint of the successful landing of the launch vehicle: factors that can cause engine damage or startup failure, such as deviations in the combustion chamber mixture ratio, factors that cause overheating, the speed and height that can cause the launch vehicle to crash during landing; soft constraints: factors that affect engine performance, such as thrust magnitude, combustion chamber pressure, etc.

[0077] S31. Define the state space during the engine startup process;

[0078] To train and use the MARL agent, it is necessary to define the state space and action space of each agent. The state space, that is, the variables received by the agent from the environment at each time step, is required to contain sufficient information to clearly define the current state of the system. The set state space is as follows:

[0079] S agent =[MR GG ,P GG ,T GG ,MR CC ,P CC ,T CC ,F,H,V,a,Pos VGO ,Pos VGF ,Pos VCF (1)

[0080] It contains 13 state parameters. Among them, MR GG ,P GG ,T GG ,MR CC ,P CC ,T CC ,F are respectively the mixture ratio of the engine gas generator, the chamber pressure of the gas generator, the temperature of the gas generator, the mixture ratio of the thrust chamber, the chamber pressure of the thrust chamber, the temperature of the thrust chamber, the thrust magnitude, H, V, a are the height, speed, and acceleration of the launch vehicle, Pos VGO ,Pos VGF ,Pos VCFThey respectively represent the oxidizer valve of the gas generator, the fuel valve of the gas generator, and the fuel valve of the combustion chamber. The observation space is normalized using the steady-state reference value. Therefore, the state space and method of the proposed engine are not limited to the simulation environment in the present invention. In other engine systems or environments, some parameters, such as variables like the engine turbine efficiency, cannot be directly measured.

[0081] S32. Define the action space during the engine startup process;

[0082] The action space A of the Agent consists of the opening degrees of three valves

[0083] A agent =[Pos VGO ,Pos VGF ,Pos VCF (2)

[0084] At each time step, the MARL Agent receives the environmental observation results and sends control signals to the control valves of the engine. The interaction frequency between the MARL Agent and the environment is 25 Hz.

[0085] S33. Define the reward function during the engine startup process.

[0086] S33 includes the following steps:

[0087] S331. The rewards for training MARL and evaluating the startup timing consist of the following different parts:

[0088] Reward = R engine +R landing (3)

[0089] R engine =r e1 +r e2 +r e3 (4)

[0090] The first reward is:

[0091]

[0092] Among them, ε i ∈[MR GG ,P GG ,T GG ,MR CC ,P CC ,T CC The reward for approaching the target value. Each reward component in this item has its maximum value clipped by 0.2 to improve the training balance of the cumulative rewards during startup and steady state.

[0093] The second reward is:

[0094]

[0095] Here, it is assumed that once the engine starts, it will not be shut down again. Therefore, after the engine generates thrust, the engine thrust should not be zero later to prevent the engine from shutting down during landing.

[0096] The third reward is:

[0097]

[0098] where Act i ∈[Pos VGO , Pos VGF , Pos VCF respectively represent the opening degrees of the three valves of the engine, encouraging to try opening the valves before the engine ignition; S represents the change in valve position between two adjacent steps before and after the valve, used to punish the reciprocating movement of the valve. Through this item, the oscillation caused by the agent frequently actuating the valve during vertical descent can be suppressed.

[0099] S332.

[0100] R landing =r l1 +r l2 +r l3 (8)

[0101] r l1 =0.2·exp(-abs(V) / 10)+0.2·(step max -step current ) / step max (9)

[0102] where step max is the maximum number of steps in each training episode, and step current is the current training step. Using an exponential form of reward makes the smaller the speed, the greater this reward, guiding the agent to decelerate the launch vehicle. In addition, introducing step into the reward function can promote the agent to complete the landing of the launch vehicle in a shorter time.

[0103] r l2 =0.2·exp(H / 50)+0.2·(step max -step current ) / step max (10)

[0104] Same as the above formula, the purpose is to enable the agent to complete a faster and better landing.

[0105]

[0106] After the launch vehicle successfully lands, the agent receives a huge reward, encouraging the strategy of successful landing to train more strategies for successful landing.

[0107] S4. Based on the MADDPG algorithm, improvements are made, including 10 iterations of update and the introduction of an annealing algorithm to implement the LCO-MADDPG (Lossless Convex Optimized-MADDPG) algorithm.

[0108] To make the MADDPG algorithm better applicable to the vertical landing control process of the launch vehicle, the LCO-MADDPG algorithm is proposed with two main improvements, including: 10 iterations of update during the training process and the use of an annealing learning rate.

[0109] S4 includes the following steps:

[0110] S41. Improve the LCO-MADDPG algorithm with 10 iterations of update during the training process;

[0111] S42. Improve the LCO-MADDPG algorithm by using an annealing learning rate.

[0112] For the non-linear and non-convex characteristics during the vertical landing process of the launch vehicle, this improvement can further enhance the stability and efficiency of the algorithm, especially suitable for dealing with environments with complex or multi-modal reward structures designed for the vertical landing process of the launch vehicle.

[0113] The pseudo-code of the LCO-MADDPG algorithm is shown in Table 1 below:

[0114] Table 1: Pseudo-code of the LCO-MADDPG algorithm

[0115]

[0116]

[0117] In a specific embodiment, the code is implemented in the Python language, and real-time interaction between Matlab and Python is achieved through the API with an interaction frequency of 25Hz. The change diagram of the thrust during the open-loop control process is obtained as Figure 5 shown, where Figure 5 represents the change of the engine thrust under open-loop control;

[0118] S5. Design, train, and evaluate the MARL model.

[0119] S5 includes the following steps:

[0120] S51. After transforming the vertical landing problem of the launch vehicle into a MARL problem, a MARL controller based on the LCO-MADDPG algorithm is designed, trained, and evaluated in combination with the characteristics of the engine system and the nonlinearity of the vertical landing stage;

[0121] S52. Taking the defined state space as the input of the controller, the MARL agent obtains the maximum reward for each Agent by controlling the actuation of the valve during the starting process of the engine according to the trained policy.

[0122] In a specific embodiment, based on the Matlab-simulink simulation platform, a MARL controller based on the LCO-MADDPG algorithm is implemented using Python code to control VGO, VGF, and VCF respectively.

[0123] Finally, it should be noted that although the present invention has been described with reference to the current specific embodiments, those of ordinary skill in the art in this technical field should recognize that the above embodiments are only used to illustrate the present invention and are not used to limit the present invention. Various equivalent changes or substitutions can be made without departing from the concept of the present invention. Therefore, as long as the changes and modifications to the above embodiments are within the scope of the spirit of the present invention, they will fall within the scope of the claims of the present invention.

Claims

1. A rocket landing control method based on multi-agent reinforcement learning, characterized in that, It includes the following steps: S1. Establish a rocket engine model, collect and output the state parameters necessary during the operation of a liquid rocket engine; S2. Establish a vertical landing simulation model for a launch vehicle, collect and output the parameters necessary during the vertical landing process of the launch vehicle; S3. Define the state space, action space, and reward function during the engine startup process; S4. Based on the MADDPG algorithm, make improvements to implement the LCO-MADDPG algorithm; S5. Design, train, and evaluate the MARL model.

2. The rocket landing control method based on multi-agent reinforcement learning according to claim 1, characterized in that, The said S1 includes the following steps: S11. Use commercial simulation software or programming language to establish a liquid rocket engine model to be analyzed; S12. Collect the state parameters necessary during the operation of the liquid rocket engine to enable the output of the variables of the state parameters to be analyzed.

3. The rocket landing control method based on multi-agent reinforcement learning according to claim 2, wherein, The said S2 includes the following steps: S21. Use the same commercial simulation software or programming language as that used in step S11 to build a simulation model for the vertical landing process of the launch vehicle; S22. Collect the parameters necessary during the vertical landing process of the launch vehicle to enable the output of the variables of the necessary parameters.

4. A rocket landing control method based on multi-agent reinforcement learning according to claim 3, characterized in that, The said S3 includes the following steps: S31. Define the state space during the engine startup process; It is necessary to define the state space of each agent, and the set state space is as follows: S agent = [MR GG , P GG , T GG , MR CC , P CC , T CC , F, H, V, a, Pos VGO , Pos VGF , Pos VCF ​ Among them, MR GG , P GG , T GG , MR CC , P CC , T CC , F are respectively the mixture ratio of the engine gas generator, the chamber pressure of the gas generator, the temperature of the gas generator, the mixture ratio of the thrust chamber, the chamber pressure of the thrust chamber, the temperature of the thrust chamber, and the thrust magnitude. H, V, a are the height, speed, and acceleration of the launch vehicle. Pos VGO , Pos VGF , Pos VCF respectively represent the oxidizer valve of the gas generator, the fuel valve of the gas generator, and the fuel valve of the combustion chamber. The state space is normalized with the steady-state reference value; S32. Define the action space during the engine startup process; The action space A of the agent consists of the opening degrees of three valves A agent = [Pos VGO , Pos VGF , Pos VCF ​ At each time step, the MARL Agent receives the environmental observation results and sends a control signal to the control valve of the engine; S33. Define the reward function during the engine startup process.

5. The rocket landing control method based on multi-agent reinforcement learning according to claim 4, characterized in that, The said S33 includes the following steps: S331. The rewards for training MARL and evaluating the startup timing consist of the following different parts: Reward=R engine +R landing R engine =r e1 +r e2 +r e3 The first reward is: Among them, ε i ∈ [MR GG , P GG , T GG , MR CC , P CC , T CC Rewards for approaching the target value. In this item, the maximum value of 0.2 is trimmed from each reward component to improve the cumulative reward during the training balance start-up and steady state; The second reward is: Here it is assumed that once the engine starts, it will not be shut down again. Therefore, after the engine generates thrust, the engine thrust should not be 0 later to prevent the engine from shutting down during the landing process; The third reward is: Among them, Act i ∈[Pos VGO , Pos VGF , Pos VCF respectively represent the opening degrees of the three valves of the engine, and it is encouraged to try to open the valves before the engine ignition; S represents the change in the valve position between two adjacent steps before and after the valve, which is used to punish the reciprocating movement of the valve. Through this term, the oscillation caused by the agent frequently actuating the valve during the vertical descent can be suppressed; S332. R landing =r l1 +r l2 +r l3 r l1 = 0.2·exp(-abs(V) / 10) + 0.2·(step max - step current ) / step max Among them, step max is the maximum number of steps in each training episode, and step current is the current training step. An exponential reward is adopted to guide the agent to decelerate the launch vehicle; In addition, introduce step into the reward function to encourage the agent to complete the landing of the launch vehicle in a shorter time, r l2 = 0.2·exp(H / 50)+0.2·(step max - step current ) / step max The same as the above formula, the purpose is to enable the agent to complete the landing faster and better, After the launch vehicle successfully lands, the agent obtains a huge reward to encourage the successful landing strategy and train more successful landing strategies.

6. The rocket landing control method based on multi-agent reinforcement learning according to claim 5, characterized in that, The said S4 includes the following steps: S41. Make improvements to the LCO-MADDPG algorithm and perform 10 iterations of updates during the training process; S42. Make improvements to the LCO-MADDPG algorithm and use an annealing learning rate.

7. The rocket landing control method based on multi-agent reinforcement learning according to claim 6, characterized in that, The said S5 includes the following steps: S51. Combining the characteristics of the nonlinearity of the engine system and the vertical landing stage, design, train, and evaluate a MARL controller based on the LCO-MADDPG algorithm; S52. Use the defined state space as the input of the controller. The MARL agent, according to the trained strategy, controls the actuation of the valves during the engine startup process to enable each Agent to obtain the maximum reward.

Citation Information

Patent Citations

  • Deep space probe soft landing path planning method based on multi-task deep reinforcement learning

    CN113408796A

  • Intelligent planning and decision-making method for landing behavior of multi-node detector

    CN115374933A

  • Rocket landing real-time robust guidance method and system based on reinforcement learning

    CN115524964A

  • Unmanned aerial vehicle cluster distributed cooperative guidance law based on multi-agent reinforcement learning method

    CN116610139A

  • Rocket landing guidance method and system based on deep reinforcement learning

    CN116697829A