A Multi-Agent Collaboration Method and System for Large Liquid Rocket Operations at Sea
By employing a multi-agent collaborative approach and deep reinforcement learning, the problem of dynamic pose uncertainty in rocket launches and recoveries at sea was solved, enabling autonomous decision-making and action planning, and improving the accuracy and safety of rocket launches and recoveries.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-11-23
- Publication Date
- 2026-03-13
AI Technical Summary
Rocket launches and recoveries at sea face complex and ever-changing marine environments, leading to uncertainties in dynamic attitude and movement, which affect the accuracy and safety of rocket launches and recoveries.
By employing a multi-agent collaborative approach, combining deep reinforcement learning and graph neural networks, autonomous decision-making and action planning are achieved through the perception, communication, and decision-making of the floating launch and recovery agents and the rocket agents, thus meeting the conditions for rocket launch and recovery.
It enables autonomous decision-making and planning for rocket launch and recovery in complex marine environments, improving the accuracy and safety of launch and recovery, avoiding the shortcomings of human experience, and featuring flexible adjustment and systematization.
Smart Images

Figure CN115729108B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of marine engineering technology, and in particular to a multi-agent cooperative method and system for large liquid rocket operations at sea. Background Technology
[0002] In recent years, competition for space resources has intensified globally. Sea-based rocket launches and recoveries offer numerous advantages and represent a promising future. Firstly, their mobility provides flexibility, avoiding political and geopolitical constraints, facilitating the selection of uninhabited landing sites, allowing for flexible choice of launch time and location based on launch windows, and enabling rapid satellite network formation at any inclination. Secondly, equatorial launches provide the rocket with maximum initial linear velocity, significantly saving fuel and increasing the payload capacity into equatorial orbit, effectively extending satellite lifespan. Major spacefaring nations are seeking launch sites as close to the equator as possible for sea-based operations, while suitable land-based launch sites are extremely limited.
[0003] A major challenge in rocket operations at sea is making decisions based on a constantly changing environment, specifically determining the ignition timing and recovery process. For a long time, land-based rocket launches have required consideration of numerous coupled factors; at sea, new uncertainties and environmental interactions are added. For example, the dynamic posture of the launch platform under the influence of complex and ever-changing ocean waves can cause the rocket to sway after the traction release mechanism is activated, affecting the alignment of the dynamic base and even leading to collapse. Sea recovery also presents new challenges. On one hand, there's the issue of landing timing caused by the dynamic posture of the recovery platform under the influence of complex and ever-changing ocean waves—when will the rocket land? On the other hand, there's the accuracy of the landing point—whether the rocket can successfully land on the recovery platform in the vast ocean. Summary of the Invention
[0004] To address the aforementioned problems and technical needs, the inventors have proposed a multi-agent collaborative method and system for large liquid rocket sea operations. This method, based on the complete process and key technologies of sea-based rocket launches, incorporates the entire process of perception, decision-making, planning, and evaluation. It systematically integrates multi-agent communication optimization technology, providing feasible and high-performance sea-based rocket launch and recovery schemes from environmental perception and communication to intelligent decision-making output. This method features intelligent characteristics and advantages, enabling autonomous decision-making during sea launches and first-stage rocket recovery, completing the planning, design, and optimization of schemes; it automates the entire design and analysis process without human intervention. It is suitable for intelligent perception and decision-making during sea launches and recovery under different environmental conditions in engineering projects.
[0005] The technical solution of the present invention is as follows:
[0006] Firstly, this application provides a multi-agent cooperative method for large liquid rocket operations at sea, comprising the following steps:
[0007] The rocket agent perceives its current stage and its own position and attitude, while the agent that collaborates and communicates with the rocket agent perceives the environmental state and its own position and attitude; among them, the agent that collaborates and communicates is a floating launch agent or a floating recovery agent.
[0008] The rocket agent processes the perceived information and the first cooperative message sent by the cooperative communication agent through deep reinforcement learning to obtain the rocket agent's action and the second cooperative message.
[0009] The cooperative communication agent processes the perceived information and the second cooperative message sent by the rocket agent through deep reinforcement learning to obtain the action of the cooperative communication agent and the first cooperative message.
[0010] The rocket agent repeatedly senses the current launch phase and its own position and attitude to complete the sea launch or sea recovery operation of large liquid rockets.
[0011] Its further technical solution is that the rocket's intelligent agent perceives its current stage and its own position and attitude, including:
[0012] The rocket intelligence agent performs temporal perception of position and attitude, and at the same time, the rocket intelligence agent perceives the current stage and determines whether it is the waiting stage for launch or the flight waiting to be recovered stage.
[0013] The rocket agent selects and establishes communication connections with other agents based on its current stage.
[0014] A further technical solution involves an intelligent agent that collaborates with the rocket's intelligent agent to perceive the environmental state and its own pose state, including when the intelligent agent collaborating with the intelligent agent is a floating launch intelligent agent at sea:
[0015] After the floating launch agent arrives at the operational sea area, its position and attitude are initially set.
[0016] The floating intelligent agent at sea performs temporal perception of position and attitude, and simultaneously perceives wave environment parameters, including short-peak waves, long waves, and solitary waves.
[0017] A further technical solution involves an intelligent agent that collaborates with the rocket's intelligent agent to perceive the environmental state and its own pose state, including when the intelligent agent collaborating with the intelligent agent is a floating recovery agent at sea:
[0018] After the floating recovery agent arrives at the operating area, its position and attitude are initially set.
[0019] The floating recovery agent at sea performs temporal perception of position and attitude, and simultaneously perceives wave environment parameters, including short-peak waves, long waves, and solitary waves.
[0020] A further technical solution involves the rocket agent processing perceived information and the first cooperative message sent by the cooperative communication agent through deep reinforcement learning to obtain the rocket agent's action and the second cooperative message, including:
[0021] The rocket agent learns the actions it will take based on the perceived information and the first collaborative message sent by the agent in collaborative communication. It uses deep reinforcement learning to obtain the actions the rocket agent will take, and employs attention mechanisms and graph neural networks to accelerate the learning speed. It learns the action mode based on the result of taking the action.
[0022] The rocket agent sends the actions it takes and the information it perceives as a second collaborative message to the collaborating agent with the communication connection.
[0023] During the launch waiting phase, the actions taken include the ignition of the first-stage rocket. In the process of deep reinforcement learning, the optimization goal of the rocket agent is to make the initial launch posture motion of the rocket agent meet the aiming conditions at takeoff and its own stability.
[0024] During the flight recovery phase, the actions taken include adjusting the thrust and direction of the first-stage rocket engine; in the process of using deep reinforcement learning, the optimization goal of the rocket agent is to ensure that the recovery posture motion of the rocket agent meets its own stability during landing and that there is no collision with the floating recovery agent at sea.
[0025] A further technical solution involves the cooperative communication agent processing the perceived information and the second cooperative message sent by the rocket agent through deep reinforcement learning to obtain the actions of the cooperative communication agent and the first cooperative message, including, when the cooperative communication agent is a floating launch agent at sea:
[0026] The floating launch agent at sea learns the actions it will take based on the perceived information and the second cooperative message sent by the rocket agent through deep reinforcement learning. It uses attention mechanisms and graph neural networks to accelerate the learning speed and learns the action mode based on the result of the action taken.
[0027] The floating launch agent at sea will send the actions it takes and the information it senses as the first collaborative message to the rocket agent with the communication connection.
[0028] The actions taken include dynamic positioning and active roll reduction of the floating intelligent agent at sea;
[0029] In the process of using deep reinforcement learning, the optimization goal of the floating launch agent at sea is to make its own pose motion meet the launch conditions of the rocket agent.
[0030] A further technical solution involves the cooperative communication agent processing the perceived information and the second cooperative message sent by the rocket agent through deep reinforcement learning to obtain the actions of the cooperative communication agent and the first cooperative message, including, when the cooperative communication agent is a floating recovery agent at sea:
[0031] The floating recovery agent at sea learns the actions it will take based on the perceived information and the second collaborative message sent by the rocket agent through deep reinforcement learning. It also uses an attention mechanism and graph neural network to accelerate the learning speed and learns the action mode based on the result of the action taken.
[0032] The floating recovery agent at sea will send the actions it takes and the information it senses as the first collaborative message to the rocket agent connected to the communication link.
[0033] The actions taken include the autonomous navigation path and active roll reduction of the floating recovery smart body at sea;
[0034] In the process of using deep reinforcement learning, the optimization goal of the floating recovery agent at sea is to make its own pose motion satisfy the recovery conditions of the rocket agent.
[0035] Secondly, this application also provides a multi-agent cooperative system for large liquid rocket operations at sea, including cooperative communication:
[0036] The floating launch agent at sea is used to perceive the environmental state and its own pose state. It processes the perceived information and the second cooperative message sent by the rocket agent through deep reinforcement learning to obtain the actions of the floating launch agent at sea and the first cooperative message to complete the sea launch operation of the large liquid rocket.
[0037] The floating recovery agent at sea is used to sense the environmental state and its own pose state, and processes the sensed information and the second cooperative message sent by the rocket agent through deep reinforcement learning to obtain the actions of the floating recovery agent at sea and the first cooperative message, so as to complete the sea recovery operation of the large liquid rocket.
[0038] The rocket agent is used to perceive its current stage and its own pose state, and processes the perceived information and the first cooperative message sent by the floating launch agent or the floating recovery agent at sea through deep reinforcement learning to obtain the rocket agent's actions and the second cooperative message.
[0039] The further technical solution is that both the floating launch intelligent body and the floating recovery intelligent body are floating structures, and both include an active roll reduction device, a wave sensing module, a first attitude sensor and a first intelligent decision-making module; the wave sensing module is used to obtain quantitative information of wave environment parameters in real time, including short-peak waves, long waves and solitary waves; the first attitude sensor is used to obtain information on the position and attitude changes of the platform in real time.
[0040] The floating launch agent also includes a dynamic positioning device. The first intelligent decision-making module I of the floating launch agent is used to obtain the actions to be taken by the floating launch agent through deep reinforcement learning based on the perceived information and the second cooperative message sent by the rocket agent. Attention mechanism and graph neural network are used to accelerate the learning speed and learn the action mode based on the result of the action taken. It is also used to send the taken action and the perceived information as the first cooperative message to the rocket agent with which it is connected. The actions taken include dynamic positioning and active roll reduction of the floating launch agent.
[0041] The floating recovery agent also includes a self-propulsion route planning module. The first intelligent decision-making module II of the floating recovery agent is used to obtain the actions to be taken by the floating recovery agent through deep reinforcement learning based on the perceived information and the second cooperative message sent by the rocket agent. Attention mechanism and graph neural network are used to accelerate the learning speed, and the action mode is learned based on the result of the action taken. It is also used to send the taken action and perceived information as the first cooperative message to the rocket agent with communication connection. The actions taken include the floating recovery agent's self-propulsion route and active roll reduction.
[0042] The further technical solution is as follows: the rocket intelligent agent includes a first-stage rocket, a stage perception module, a second pose sensor, and a second intelligent decision-making module; the stage perception module is used to obtain the current stage of the first-stage rocket, determining whether it is the waiting-for-launch stage or the flight-to-recovery stage; the second pose sensor is used to obtain the position and attitude change information of the first-stage rocket in real time; the second intelligent decision-making module is used to obtain the actions to be taken by the rocket intelligent agent through deep reinforcement learning based on the perceived information and the first cooperative message sent by the floating launch intelligent agent or the floating recovery intelligent agent, and to learn the action mode based on the result of the action taken; it is also used to select the intelligent agent to cooperate with according to the current stage, and send the taken action and perceived information as the second cooperative message to the cooperative intelligent agent; wherein, for the waiting-for-launch stage, the actions taken include the ignition operation of the first-stage rocket; for the flight-to-recovery stage, the actions taken include the thrust and directional adjustment of the first-stage rocket engine.
[0043] The beneficial technical effects of this invention are:
[0044] First, this application creatively utilizes deep reinforcement learning to achieve decision-making and planning for the intelligent launch and recovery of large liquid rockets at sea, avoiding the shortcomings of human experience-based judgment. It cleverly employs multi-agent perception, communication, and decision-making to represent the uncertainties and planning / decision-making processes involved in sea-based rocket launches, thus providing the correct actions.
[0045] Second, this application features collaborative decision-making. The three agents are partially observable and can fully utilize information from other agents when making decisions through mutual communication, enabling the entire launch and recovery system to have a globally consistent cognitive representation.
[0046] Third, this application features flexible adjustment and high generalization. This is mainly reflected in the learning process of time-series decision-making, which does not require consideration of all environments at the outset. Based on optimization objectives and the constantly changing environment, and tailored to different operational characteristics, it obtains scalar-form rewards through environmental perception, enabling launch and first-stage rocket recovery decisions in different sea areas.
[0047] Fourth, this application has the characteristics of being systematic, and can coordinate and control multiple intelligent agents to make decisions. Each intelligent agent interacts with the environment and interacts with the corresponding collaborative messages of the collaborating intelligent agents, and maximizes its utility based on the obtained state. On the other hand, it continuously trains its own model to improve the decision-making ability of the intelligent agent based on the data generated by the interaction. Attached Figure Description
[0048] Figure 1 This is a flowchart of a large liquid rocket sea launch operation in the multi-agent cooperative method provided in this application.
[0049] Figure 2 This is a flowchart of a large liquid rocket sea recovery operation in the multi-agent cooperative method provided in this application.
[0050] Figure 3 This is a complete flowchart of another large liquid rocket sea launch operation in the multi-agent cooperative method provided in this application.
[0051] Figure 4 This is a complete flowchart of another large-scale liquid rocket sea recovery operation in the multi-agent cooperative method provided in this application.
[0052] Figure 5 This is a schematic diagram of the multi-agent cooperative system provided in this application. Detailed Implementation
[0053] The specific embodiments of the present invention will be further described below with reference to the accompanying drawings.
[0054] This embodiment discloses a multi-agent collaborative method and system for large liquid rocket maritime operations. The floating launch agent, the floating recovery agent, and the rocket agent can complete given tasks based on perceived environmental information and mutual communication, and can influence and change the environmental state through the actions they perform. The method and system rely on multi-agent collaborative perception, decision-making, and planning to better achieve maritime launch and recovery tasks. The method and system mainly address three aspects: (1) environmental perception, (2) intelligent decision-making and planning, and (3) agent actions. The latitude and longitude requirements of rocket launch necessitate that the platform's horizontal position be controlled within a certain range. This requires a corresponding positioning system to make intelligent decisions and take actions, namely perception, decision-making, planning, and control, to keep the platform within a certain range. Rocket recovery requires control of the landing point, necessitating mutual communication, decision-making, and actions between the floating recovery agent and the rocket agent, so that the rocket can accurately land on the recovery agent. Furthermore, stability is involved throughout the rocket launch and recovery process, requiring attitude control during launch and recovery, ultimately leading to decisions and planning regarding whether to launch or recover the rocket, and real-time actions of the rocket attitude and platform attitude.
[0055] This embodiment discloses a multi-agent cooperative method for large liquid rocket sea launch operations. Please refer to [link / reference]. Figure 1 As shown, it includes the following steps:
[0056] Step 1: The rocket agent perceives its current stage and its own position and posture.
[0057] Step 2: The floating launch agent at sea, which communicates with the rocket agent, perceives the environmental state and its own position and attitude.
[0058] Step 3: The rocket agent processes the perceived information and the first cooperative message sent by the floating launch agent at sea through deep reinforcement learning to obtain the rocket agent's action and the second cooperative message.
[0059] Step 4: The rocket agent sends a second cooperative message to the floating launch agent at sea.
[0060] Step 5: The floating launch agent at sea processes the perceived information and the second cooperative message sent by the rocket agent through deep reinforcement learning to obtain the actions of the floating launch agent at sea and the first cooperative message.
[0061] Step 6: The floating launch agent at sea sends the first cooperation message to the rocket agent.
[0062] Repeat steps 1-6, with the floating launch agent and the rocket agent sensing and communicating with each other to complete the multi-agent collaborative decision-making for the sea launch operation of large liquid rockets.
[0063] Specifically, the floating launch agent's perception of the environment mainly involves acquiring various wave environment parameters, including short-peak waves, long waves, and solitary waves, using methods such as computer vision or multi-parameter buoys. The floating launch agent's perception of its own position and attitude involves methods such as inertial navigation and positioning systems. The floating launch agent's actions include dynamic positioning and active roll reduction. The rocket agent's perception of its current stage mainly involves confirming whether it is in the waiting-for-launch, flight-to-recovery, or recovered on the floating recovery agent, using methods such as signal-based measurement and accelerometers. The rocket agent's perception of its own position and attitude mainly involves obtaining the initial launch attitude to prepare for alignment with the launch platform, using methods such as inertial navigation and positioning systems. At this stage, the rocket agent's actions include ignition and non-ignition.
[0064] This embodiment discloses a multi-agent cooperative method for large liquid rocket sea recovery operations. Please refer to [link / reference]. Figure 2 As shown, it includes the following steps:
[0065] Step 1: The rocket agent perceives its current stage and its own position and posture.
[0066] Step 2: The floating recovery agent at sea, which communicates with the rocket agent, senses the environmental state and its own pose.
[0067] Step 3: The rocket agent processes the perceived information and the first cooperative message sent by the floating recovery agent at sea through deep reinforcement learning to obtain the rocket agent's action and the second cooperative message.
[0068] Step 4: The rocket agent sends a second cooperative message to the floating recovery agent at sea.
[0069] Step 5: The floating recovery agent at sea processes the perceived information and the second cooperative message sent by the rocket agent through deep reinforcement learning to obtain the actions of the floating recovery agent at sea and the first cooperative message.
[0070] Step 6: The floating recovery agent at sea sends the first cooperation message to the rocket agent.
[0071] Repeat steps 1-6, with the floating recovery agent at sea and the rocket agent sensing and communicating with each other to complete the multi-agent collaborative decision-making for the sea recovery operation of the large liquid rocket.
[0072] Specifically, the floating recovery aerospace system at sea primarily perceives various wave environment parameters, including short-peak waves, long waves, and solitary waves, using methods such as computer vision or multi-parameter buoys. Its own position and attitude are perceived using inertial navigation and positioning systems. The aerospace system's actions include autonomous route planning and active roll reduction to fine-tune its position. The rocket aerospace system perceives its current stage, confirming whether it is in the waiting-for-launch, flight-to-recovery, or recovered phase on the floating launch aerospace system, using methods such as signal-based measurement and accelerometers. Its position and attitude are perceived during the recovery phase using methods such as inertial navigation and positioning systems. At this stage, the rocket aerospace system's actions include adjusting the first-stage rocket's engine to ensure it lands on the deck of the floating recovery aerospace system.
[0073] This embodiment discloses another multi-agent cooperative method for large liquid rocket sea launch operations. Please refer to [link / reference]. Figure 3 As shown, it includes the following steps:
[0074] Step 1: After the floating launch agent arrives at the operational sea area, its position and attitude are initially set. The floating launch agent then performs temporal perception of its position and attitude.
[0075] Step 2: The floating intelligent agent at sea senses wave environment parameters, including short-peak waves, long waves, and solitary waves.
[0076] Steps 1 and 2 together complete two aspects of environmental perception for the floating intelligent agent at sea.
[0077] Step 3: Based on the perceived information and the second cooperative message sent by the rocket agent, the floating launch agent at sea uses deep reinforcement learning to obtain the actions it will take. Attention mechanisms and graph neural networks are employed to accelerate the learning process, and the action patterns are learned based on the results of the actions taken. These actions include dynamic positioning and active roll reduction.
[0078] Optionally, during the deep reinforcement learning process, the optimization objective of the floating launch agent is to ensure that its own pose and motion satisfy the launch conditions of the rocket agent. Specifically, when the pose conditions reach a set threshold, the reward is maximized, and the next stage of action can proceed.
[0079] Step 4: The floating launch agent operates according to the planned dynamic positioning and active roll reduction mode, and sends the actions taken and the perceived information as the first cooperative message to the rocket agent with which it is connected.
[0080] Step 5: The rocket agent senses the current stage, determines it to be the waiting-for-launch stage, and selects to establish a communication connection with the floating launch agent at sea based on the current stage.
[0081] Step 6: The rocket agent performs temporal perception of position and attitude.
[0082] Steps 5 and 6 together complete two aspects of the rocket agent's environmental perception.
[0083] Step 7: Based on the perceived information and the first collaborative message sent by the floating launch agent at sea, the rocket agent uses deep reinforcement learning to obtain the actions it will take. Attention mechanisms and graph neural networks are employed to accelerate the learning process, and the action pattern is learned based on the results of the actions taken. These actions include the ignition of the first-stage rocket.
[0084] Optionally, during the deep reinforcement learning process, the optimization objective of the rocket agent is to ensure that its initial launch posture meets the aiming conditions and stability requirements at takeoff. Specifically, the ignition and launch action is executed when the posture meets the threshold requirements for both the aiming conditions and stability at takeoff.
[0085] Step 8: The rocket agent sends the actions it has taken and the information it has sensed as a second cooperative message to the floating launch agent at sea.
[0086] Step 9: Based on the attitude change behavior of the floating launch agent and the motion behavior of the rocket agent obtained through deep reinforcement learning, the final sea launch decision is made, specifically including:
[0087] Based on the attitude and motion perception results of the floating launch agent and the wave environment perception, the rocket agent makes ignition and launch decisions according to the attitude change behavior patterns of the rocket and the floating launch platform obtained through deep reinforcement learning. If the predicted attitude and motion of the floating launch platform meets the launch requirements, ignition and launch occur; if the predicted attitude and motion of the launch pad do not meet the launch requirements, ignition does not occur, and the process returns to step 1 to begin the next cycle of dynamic decision-making. This is an iterative temporal decision-making process. On the one hand, the actions of the floating launch agent in step 4 affect the environment; on the other hand, the actions of the rocket agent in step 7 also change the environment, ultimately completing the intelligent decision-making for the maritime launch operation.
[0088] This embodiment discloses another multi-agent cooperative method for large liquid rocket sea recovery operations. Please refer to [link / reference]. Figure 4 As shown, it includes the following steps:
[0089] Step 1: After arriving at the operational area, the floating recovery agent performs initial position and attitude settings. The floating recovery agent then performs temporal perception of its position and attitude.
[0090] Step 2: The floating recovery agent at sea senses wave environment parameters, including short-peak waves, long waves, and solitary waves.
[0091] Steps 1 and 2 together complete two aspects of environmental perception for the marine floating recovery agent.
[0092] Step 3: Based on the perceived information and the second collaborative message sent by the rocket agent, the floating recovery agent at sea uses deep reinforcement learning to obtain the actions it will take. Attention mechanisms and graph neural networks are employed to accelerate the learning process, and the action patterns are learned based on the results of the actions taken. These actions include self-propulsion and active roll reduction.
[0093] Optionally, during the deep reinforcement learning process, the optimization objective of the floating recovery agent at sea is to ensure that its own pose and motion satisfy the recovery conditions of the rocket agent. Specifically, when the pose conditions reach a set threshold, the reward is maximized, and the next stage of action can proceed.
[0094] Step 4: The floating recovery agent at sea operates according to the planned self-propulsion route and active roll reduction mode, and sends the actions taken and the information sensed as the first collaborative message to the rocket agent with which it is connected.
[0095] Step 5: The rocket's intelligent agent senses the current stage, determines it to be in the flight recovery stage, and selects to establish a communication connection with the floating recovery intelligent agent at sea based on the current stage.
[0096] Step 6: The rocket agent performs temporal perception of position and attitude.
[0097] Steps 5 and 6 together complete two aspects of the rocket agent's environmental perception.
[0098] Step 7: Based on the perceived information and the first collaborative message sent by the floating recovery agent at sea, the rocket agent uses deep reinforcement learning to obtain the actions it will take. Attention mechanisms and graph neural networks are employed to accelerate the learning process, and the action patterns are learned based on the results of the actions taken. These actions include adjusting the thrust and direction of the first-stage rocket engine.
[0099] Optionally, during the deep reinforcement learning process, the optimization objective of the rocket agent is to ensure that its recovery pose motion satisfies its own stability during landing and avoids collisions with the floating recovery agent at sea. Specifically, the recovery action is performed when the pose motion meets the relative motion threshold of satisfying its own stability during landing and avoiding collisions with the floating recovery agent at sea.
[0100] Step 8: The rocket agent sends the actions it has taken and the information it has sensed as a second collaborative message to the floating recovery agent at sea.
[0101] Step 9: Based on the attitude change behavior of the floating recovery agent and the motion behavior of the rocket agent obtained through deep reinforcement learning, the final sea recovery decision is made, specifically including:
[0102] Based on the attitude and motion perception results of the floating recovery agent at sea and the wave environment perception, and according to the attitude change behavior patterns of the rocket and the floating recovery platform obtained through deep reinforcement learning, the rocket agent makes recovery decisions. If the predicted landing point and attitude motion conditions meet the requirements for first-stage rocket recovery, the rocket is recovered; if the predicted landing point of the first-stage rocket or the swaying characteristics of the recovery platform do not meet the recovery requirements, the rocket is not recovered, and the process returns to step 1 to begin the next cycle of dynamic decision-making. This is an iterative temporal decision-making process. On the one hand, the actions of the floating recovery agent in step 4 affect the environment; on the other hand, the actions of the rocket agent in step 7 also change the environment. Ultimately, the intelligent decision-making for the sea recovery operation is completed.
[0103] Based on the same inventive concept, in one embodiment, a multi-agent collaborative system for large liquid rocket maritime operations is also disclosed. This system interacts with the environment to complete a specific launch mission. The characteristics of the mission environment constituted by this process can be conceptually described from the following dimensions: (1) Due to the incompleteness of environmental observation, the situation in this application belongs to a partially observable Markov decision process (POMDP); (2) Due to the uncertainty of actions, the maritime platform agent needs to perform corresponding positioning control, and the rocket agent also needs to perform attitude control; (3) The three agents need to dynamically evaluate their respective environmental states at all times; (4) The perception of the environment is a continuous process. The system that interacts with the environment to make decisions can be regarded as a closed-loop feedback system composed of "floating launch agent 1 - rocket agent 3 - floating recovery agent 2" interacting with the environment, such as Figure 5 As shown.
[0104] The floating launch agent 1 is used to sense the environmental state and its own position and attitude. Through deep reinforcement learning, it processes the sensed information and the second cooperative message sent by the rocket agent 3 to obtain the actions of the floating launch agent 1 and the first cooperative message, thus completing the sea launch operation of a large liquid rocket. Specifically, agent 1 is a semi-submersible floating structure with good seaworthiness and stability. It has a large deck area and an adjustable draft according to operating conditions. This platform is responsible for the assembly and inspection of launch vehicles, propulsion systems, and spacecraft, as well as rocket fueling, automatic pre-launch preparation, and rocket maintenance, providing a favorable launch environment. Agent 1 includes a dynamic positioning device and active roll reduction device, a wave sensing module, a first attitude sensor, and a first intelligent decision-making module I. The wave sensing module is used to obtain quantitative information on wave environment parameters in real time, including short-peak waves, long waves, and solitary waves. The first attitude sensor is used to obtain real-time information on the platform's position and attitude changes. The first intelligent decision-making module I is used to obtain the actions to be taken by the floating launch agent 1 at sea based on the perceived information and the second collaborative message sent by the rocket agent 3, using deep reinforcement learning. It employs an attention mechanism and graph neural network to accelerate the learning speed and learns the action mode based on the results of the actions taken. It is also used to send the taken actions and perceived information as the first collaborative message to the rocket agent 3, which is connected to the communication link, to determine whether the rocket will ignite and launch. The actions taken include dynamic positioning and active roll reduction of the floating launch agent 1 at sea.
[0105] The floating recovery agent 2 is used to sense the environmental state and its own attitude state. Through deep reinforcement learning, it processes the sensed information and the second cooperative message sent by the rocket agent 3 to obtain the actions of the floating recovery agent 2 and the first cooperative message, thus completing the sea recovery operation of the large liquid rocket. Specifically, agent 2 is a floating structure with a large deck area, possessing good seaworthiness and stability, self-propulsion capability, and active roll reduction device. This platform undertakes the recovery task of the first-stage rocket. Agent 2 includes a self-propulsion route planning module and active roll reduction device, a wave sensing module, a first attitude sensor, and a first intelligent decision-making module II. The wave sensing module is used to obtain quantitative information on wave environment parameters in real time, including short-peak waves, long waves, and solitary waves. The first attitude sensor is used to obtain real-time information on the platform's position and attitude changes. The first intelligent decision-making module II is used to obtain the actions to be taken by the floating recovery agent 2 at sea based on the perceived information and the second collaborative message sent by the rocket agent 3. It employs deep reinforcement learning, using attention mechanisms and graph neural networks to accelerate the learning process, and learns the action mode based on the results of the actions taken. It is also used to send the taken actions and perceived information as the first collaborative message to the rocket agent 3, which is connected to the communication link, to determine the recovery of the first stage rocket. The actions taken include the autonomous navigation route and active roll reduction of the floating recovery agent 2 at sea.
[0106] Rocket agent 3 is used to perceive its current stage and its own pose state. It processes the perceived information and the first cooperative message sent by either the floating launch agent 1 or the floating recovery agent 2 through deep reinforcement learning to obtain the rocket agent 3's action and a second cooperative message. Specifically, rocket agent 3 includes a first-stage rocket, a stage perception module, a second pose sensor, and a second intelligent decision-making module. The stage perception module determines the current stage of the first-stage rocket, identifying it as either the launch-awaiting stage or the flight-to-recovery stage. The second pose sensor obtains real-time information on the position and attitude changes of the first-stage rocket. The second intelligent decision-making module, based on the perceived information and the first cooperative message sent by either the floating launch agent 1 or the floating recovery agent 2, uses deep reinforcement learning to determine the action the rocket agent will take. It employs attention mechanisms and graph neural networks to accelerate the learning process and learns the action pattern based on the results of the actions taken. It also selects a cooperating agent to establish communication connections based on the current stage and sends the taken action and perceived information as the second cooperative message to the cooperating agent to determine whether to launch or recover the first-stage rocket. During the launch-waiting phase, actions include igniting the first-stage rocket; during the recovery phase, actions include adjusting the thrust and direction of the first-stage rocket engine.
[0107] The innovation of this application lies primarily in providing a novel design and convenient operational process for launching / recovering rockets at sea. This mainly involves the perception, communication, and decision-making of the floating launch agent 1, the floating recovery agent 2, and the rocket agent 3. The agent perceives the environment, addresses system uncertainty through deep reinforcement learning, defines goals (i.e., optimization goals) based on rewards, and improves decision-making capabilities through communication between agents, ultimately achieving optimal action output.
[0108] It should be noted that the solution provided by the above system is similar to the solution described in the above method. Therefore, the specific limitations of the above embodiment of a multi-agent cooperative system for large liquid rocket maritime operations can be found in the limitations of the multi-agent cooperative method for large liquid rocket maritime operations described above, and will not be repeated here.
[0109] The above descriptions are merely preferred embodiments of this application, and the present invention is not limited to the above embodiments. It is understood that other improvements and variations directly derived or conceived by those skilled in the art without departing from the spirit and concept of the present invention should be considered to be included within the protection scope of the present invention.
Claims
1. A large liquid rocket sea operation multi-agent cooperation method, characterized in that, The method comprises: The rocket agent perceives the current stage and its own pose state, and the agent in cooperation with the rocket agent perceives the environment state and its own pose state; wherein the agent in cooperation is a sea floating launch agent or a sea floating recovery agent; The rocket agent processes the perceived information and the first cooperation message sent by the agent in cooperation through deep reinforcement learning, to obtain the action of the rocket agent and the second cooperation message; The agent in cooperation processes the perceived information and the second cooperation message sent by the rocket agent through deep reinforcement learning, to obtain the action of the agent in cooperation and the first cooperation message; The rocket agent perceives the current stage and its own pose state repeatedly, to complete the sea launch operation or the sea recovery operation of the large liquid rocket; The rocket agent processes the perceived information and the first cooperation message sent by the agent in cooperation through deep reinforcement learning, to obtain the action of the rocket agent and the second cooperation message, comprising: The rocket agent acquires the action to be taken by the rocket agent through the method of deep reinforcement learning for the perceived information and the first cooperation message sent by the agent in cooperation, adopts the attention mechanism and the graph neural network to speed up the learning speed, and learns the action mode according to the result of taking the action; The rocket agent sends the taken action and the perceived information to the cooperation agent connected in communication as the second cooperation message; When the rocket agent is in the waiting launch stage, the taken action comprises the ignition operation of the first-stage rocket; in the process of the method of deep reinforcement learning, the optimization target of the rocket agent is to make the launch initial pose motion of the rocket agent meet the aiming condition at the time of take-off and the self-stability; When the rocket agent is in the flight waiting recovery stage, the taken action comprises the thrust and direction adjustment of the first-stage rocket engine; in the process of the method of deep reinforcement learning, the optimization target of the rocket agent is to make the recovery pose motion of the rocket agent meet the self-stability at the time of landing and have no collision with the sea floating recovery agent.
2. The method of claim 1, wherein, The rocket agent perceives the current stage and its own pose state, comprising: The rocket agent performs time sequence perception of the position and the attitude, and simultaneously perceives the current stage to determine whether it is the waiting launch stage or the flight waiting recovery stage; The rocket agent selects the agent in cooperation to perform communication connection according to the current stage.
3. The method of claim 1, wherein, The agent in cooperation with the rocket agent perceives the environment state and its own pose state, comprising, when the agent in cooperation is the sea floating launch agent: The sea floating launch agent performs initial setting of the pose state after reaching the operation sea area; The sea floating launch agent performs time sequence perception of the position and the attitude, and simultaneously performs perception of the wave environment parameters, the wave environment parameters comprising short peak wave, long wave and isolated wave.
4. The method of claim 1, wherein, The agent in cooperation with the rocket agent perceives the environment state and the pose state of itself, including, when the agent in cooperation is the sea floating recovery agent: After the sea floating recovery agent arrives at the operation sea area, the pose state is initially set; The sea floating recovery agent performs time sequence perception of position and attitude, and simultaneously performs wave environment parameter perception, the wave environment parameters including short peak wave, long wave and solitary wave.
5. The method of claim 1, wherein, The agent in cooperation processes the perceived information and the second cooperation message sent by the rocket agent through deep reinforcement learning, obtains the action of the agent in cooperation and the first cooperation message, including, when the agent in cooperation is the sea floating launch agent: The sea floating launch agent obtains the action to be taken by the sea floating launch agent through the method of deep reinforcement learning for the perceived information and the second cooperation message sent by the rocket agent, adopts the attention mechanism and the graph neural network to accelerate the learning speed, and learns the action mode according to the result of taking the action; The sea floating launch agent sends the taken action and the perceived information as the first cooperation message to the rocket agent connected by communication; Wherein, the taken action includes the dynamic positioning and active roll reduction of the sea floating launch agent; In the process of the method of deep reinforcement learning, the optimization target of the sea floating launch agent is to make the pose motion of the sea floating launch agent meet the launch condition of the rocket agent.
6. The method of claim 1, wherein, The agent in cooperation processes the perceived information and the second cooperation message sent by the rocket agent through deep reinforcement learning, obtains the action of the agent in cooperation and the first cooperation message, including, when the agent in cooperation is the sea floating recovery agent: The sea floating recovery agent obtains the action to be taken by the sea floating recovery agent through the method of deep reinforcement learning for the perceived information and the second cooperation message sent by the rocket agent, adopts the attention mechanism and the graph neural network to accelerate the learning speed, and learns the action mode according to the result of taking the action; The sea floating recovery agent sends the taken action and the perceived information as the first cooperation message to the rocket agent connected by communication; Wherein, the taken action includes the self-propelled route and active roll reduction of the sea floating recovery agent; In the process of the method of deep reinforcement learning, the optimization target of the sea floating recovery agent is to make the pose motion of the sea floating recovery agent meet the recovery condition of the rocket agent.
7. A large liquid rocket sea operation multi-agent cooperation system characterized by, The sea floating launch agent for perceiving the environment state and the pose state of itself, and processing the perceived information and the second cooperation message sent by the rocket agent through deep reinforcement learning, obtaining the action of the sea floating launch agent and the first cooperation message, to complete the sea launch operation of large liquid rocket; The offshore floating recovery intelligent agent is used for perceiving an environmental state and a self pose state, and processing perceived information and a second cooperation message sent by a rocket intelligent agent through deep reinforcement learning to obtain an action and a first cooperation message of the offshore floating recovery intelligent agent, so as to complete offshore recovery work of a large liquid rocket; The rocket intelligent agent is used for perceiving a current stage and a self pose state, and processing perceived information and a first cooperation message sent by the offshore floating launch intelligent agent or the offshore floating recovery intelligent agent through deep reinforcement learning to obtain an action and a second cooperation message of the rocket intelligent agent; The rocket intelligent agent comprises a first-stage rocket, a stage perception module, a second pose sensor and a second intelligent decision module. The stage perception module is used for obtaining a current stage of the first-stage rocket, and determining whether the current stage is a waiting launch stage or a flight waiting recovery stage. The second pose sensor is used for obtaining position and attitude change information of the first-stage rocket in real time. The second intelligent decision module is used for obtaining an action to be taken by the rocket intelligent agent through a deep reinforcement learning method based on perceived information and the first cooperation message sent by the offshore floating launch intelligent agent or the offshore floating recovery intelligent agent, using an attention mechanism and a graph neural network to speed up the learning speed, and learning an action mode according to a result of taking the action. The second intelligent decision module is also used for selecting a cooperative intelligent agent according to the current stage, and connecting the cooperative intelligent agent in communication, and sending the taken action and the perceived information as a second cooperation message to the cooperative intelligent agent. For the waiting launch stage, the taken action comprises a firing operation of the first-stage rocket. For the flight waiting recovery stage, the taken action comprises thrust and direction adjustment of an engine of the first-stage rocket.
8. The large liquid rocket marine operation multi-agent collaboration system according to claim 7, wherein, The offshore floating launch intelligent agent and the offshore floating recovery intelligent agent are both floating structures, and each comprises a active roll damping device, a wave perception module, a first pose sensor and a first intelligent decision module. The wave perception module is used for obtaining quantized information of wave environmental parameters in real time. The wave environmental parameters comprise short-crested waves, long waves and solitary waves. The first pose sensor is used for obtaining position and attitude change information of a platform in real time. The offshore floating launch intelligent agent further comprises a dynamic positioning device. The first intelligent decision module I of the offshore floating launch intelligent agent is used for obtaining a taken action of the offshore floating launch intelligent agent through a deep reinforcement learning method based on perceived information and a second cooperation message sent by the rocket intelligent agent, using an attention mechanism and a graph neural network to speed up the learning speed, and learning an action mode according to a result of taking the action. The first intelligent decision module I is also used for sending the taken action and the perceived information as a first cooperation message to the rocket intelligent agent connected in communication. The taken action comprises dynamic positioning and active roll damping of the offshore floating launch intelligent agent. The offshore floating recovery intelligent agent further comprises a self-navigation route planning module, and a first intelligent decision module II of the offshore floating recovery intelligent agent is configured to, for the perceived information and the second cooperation message sent by the rocket intelligent agent, acquire an action to be taken by the offshore floating recovery intelligent agent through a deep reinforcement learning method, accelerate the learning speed by using an attention mechanism and a graph neural network, and learn an action mode according to a result of taking the action; and further configured to send the taken action and the perceived information to the rocket intelligent agent connected in communication as a first cooperation message; wherein the taken action comprises a self-navigation route and active roll reduction of the offshore floating recovery intelligent agent.
Citation Information
Patent Citations
Stable control platform and stable control method for marine rocket launching
CN114114918A
Offshore rocket recovery method for compensating drop point deviation through autonomous navigation
CN114152152A