Deep reinforcement learning-based anti-tilting control method for marine floating launching device
By constructing a multi-source state feature set and a multi-modal control strategy through deep reinforcement learning, and combining marine environmental disturbance parameters and control priority weights, multi-agent collaborative control of the floating launch device was realized. This solved the problem of insufficient control in complex environments by traditional methods and improved the stability and energy efficiency of the device.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-24
- Publication Date
- 2026-03-31
AI Technical Summary
Existing floating launchers at sea lack sufficient control precision and robustness in the face of multi-source disturbances and complex environments. Traditional methods lack adaptive dynamic optimization mechanisms, resulting in slow attitude adjustment speed, high energy consumption, and insufficient stability.
Based on deep reinforcement learning, this method collects the basic performance parameters of the actuator unit, constructs a multi-source state feature set, introduces marine environmental disturbance parameters, generates a multi-modal control strategy distribution map, and introduces control priority weights to achieve multi-agent collaborative control, generates the final anti-tilt control command sequence, and performs adaptive adjustment through real-time reward signals.
It improves the stability and energy efficiency of the floating launcher in complex sea conditions, enhances the safety and robustness of the device, has good scalability and engineering verifiability, and significantly improves the device's stable operation capability.
Smart Images

Figure CN121763744A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of deep reinforcement learning technology, specifically to an anti-tilt control method for a floating launcher at sea based on deep reinforcement learning. Background Technology
[0002] With the continuous development of marine engineering equipment and space launch technology, floating launch vehicles are gradually becoming an important choice for new space launch modes. Floating launch vehicles offer flexible site selection, low dependence on ground infrastructure, and ease of multi-dimensional launches, effectively reducing launch risks and improving launch window utilization efficiency. However, the complex and variable environmental characteristics at sea, such as waves, wind currents, and tides, significantly affect the attitude stability of the floating vehicle, thus impacting the accuracy and safety of rocket launches. Existing research primarily employs traditional control methods, including PID-based attitude control, fuzzy control, and model predictive control. While these methods are effective in handling single or short-term disturbances, their control accuracy and robustness are often insufficient in complex environments with coupled multi-source disturbances and drastic dynamic changes. Furthermore, existing methods generally rely on precise mathematical modeling, while the marine environment is highly uncertain and random, making it difficult to accurately characterize using deterministic models.
[0003] However, existing research on directly applying deep reinforcement learning to the control of floating launch vehicles at sea remains limited, mainly focusing on the simulation level and lacking systematic fusion of actuator characteristic parameters and design of dynamic response mechanisms for multi-source disturbances. At the actuator level, current methods often only consider single leveling or propulsion control, failing to achieve multi-agent cooperative control of multiple actuators, resulting in slow attitude adjustment speed, high energy consumption, and insufficient stability under complex disturbances. Especially in terms of control strategy scheduling and priority weight introduction, traditional methods lack adaptive dynamic optimization mechanisms, making it impossible to flexibly adjust to different disturbance intensities and coupling effects. Therefore, existing technologies still have significant shortcomings in terms of real-time performance, cooperativeness, and adaptability in anti-roll control. Summary of the Invention
[0004] The purpose of this invention is to provide an anti-tilt control method for a floating launcher at sea based on deep reinforcement learning, so as to solve the problems mentioned in the background art.
[0005] To solve the above-mentioned technical problems, the present invention provides the following technical solution:
[0006] A deep reinforcement learning-based anti-tilt control method for a floating launcher at sea includes the following steps: Step S1: Collect basic performance parameters of different actuator units from the backend of the floating launcher control system and construct a multi-source state feature set; Step S2: Preset marine environmental disturbance parameters and construct a dynamic disturbance label attitude response map; construct a stability index matrix and match it with the stability index matrix to generate a corrected dynamic disturbance label attitude response map; Step S3: Based on the corrected dynamic disturbance label attitude response map, construct an actuator action space set; introduce a multi-modal anti-disturbance strategy into the actuator action space set, construct a control strategy scheduling matrix, and generate a multi-modal control strategy distribution map; Step S4: Based on the multi-modal control... The strategy distribution map maps all actuator units in the corrected dynamic disturbance label attitude response map to their corresponding control strategy entries, constructs a multi-agent control strategy correspondence matrix, and introduces control priority weights to generate a multi-agent cooperative control map; Step S5: Based on the multi-agent cooperative control map, a final anti-tilt control command sequence is generated; the final anti-tilt control command sequence is sent to the corresponding actuator units via the control system communication bus; an instant reward signal is generated; Step S6: Based on the final anti-tilt control command sequence, current ocean environment disturbance parameters, new motion state parameters, and the instant reward signal, an experience tuple is constructed; the policy network and value network parameters of the deep reinforcement learning network are updated, and feedback is provided for adaptive control and adjustment.
[0007] As a preferred embodiment of the anti-tilt control method for a floating marine launcher based on deep reinforcement learning described in this invention, basic performance parameters of different actuator units are collected from the backend of the floating marine launcher control system. The actuator units include a ballast tank actuator unit, a thruster actuator unit, and an attitude adjustment actuator unit. The basic performance parameters of the ballast tank actuator unit include ballast water tank capacity, leveling response time, and energy consumption parameters. The basic performance parameters of the thruster actuator unit include thrust magnitude, steering angle, and endurance. The basic performance parameters of the attitude adjustment actuator unit include pitch angle adjustment range, roll angle adjustment accuracy, and response delay parameters.
[0008] As a preferred embodiment of the anti-tilt control method for a floating launcher based on deep reinforcement learning described in this invention, a multi-source state feature set is constructed based on the basic performance parameters of the balance cabin actuator, the thruster actuator, and the attitude adjustment actuator, as detailed below:
[0009] The numbers are assigned according to a unified status identifier code, which includes a three-segment combination code consisting of a unit category prefix, a device number, and a performance serial number. A correspondence is established based on the motion state parameters and control attribute parameters of the actuator unit. The motion state parameters include the roll angle, pitch angle, and heave displacement of the actuator unit. The control attribute parameters include the control mode, available time period, and action priority of the actuator unit.
[0010] The multi-source state feature set is input into a deep reinforcement learning network. An initial attitude response map is generated based on the correlation between the actuator units. The correlation between the actuator units is calculated by normalization and Euclidean distance to determine the matching degree of the basic performance parameters and control attribute parameters between the actuator units. The correlation is then weighted by a preset distance weight factor of the motion state parameters to calculate the coordination degree between different actuator units.
[0011] As a preferred embodiment of the anti-tilt control method for a floating launcher based on deep reinforcement learning described in this invention, ocean environmental disturbance parameters are preset, including wave height parameters, wind speed and direction parameters, and ocean current intensity parameters; the ocean environmental disturbance parameters are introduced into all actuator units in the initial attitude response diagram and mapped to the unified state identifier code to construct a dynamic disturbance label attitude response diagram;
[0012] A stability index matrix is constructed, which includes attitude recovery time, maximum tilt angle limit, and energy consumption limit. The dynamic disturbance tag attitude response map is matched with the stability index matrix, and the basic performance parameters of the actuator unit are corrected by weighted matching to generate a corrected dynamic disturbance tag attitude response map.
[0013] As a preferred embodiment of the anti-tilt control method for a floating launcher based on deep reinforcement learning described in this invention, the leveling response time parameter of the balance cabin actuator and the response delay parameter of the attitude adjustment actuator are extracted based on the corrected dynamic disturbance tag attitude response map; the motion capability parameters are uniformly quantified to construct the actuator motion space set.
[0014] A multimodal anti-disturbance strategy is introduced into the action space set of the actuator, which includes a main control strategy, an auxiliary strategy, and an emergency recovery strategy; and the multimodal anti-disturbance strategy is bound to the unified state identifier code to construct a control strategy scheduling matrix.
[0015] The control strategy scheduling matrix is input into a deep reinforcement learning network, and a multimodal control strategy distribution map is generated based on the correlation between motion state parameters and disturbance intensity parameters between actuator units.
[0016] As a preferred embodiment of the anti-tilt control method for a floating launcher based on deep reinforcement learning described in this invention, based on the multimodal control strategy distribution map, all actuator units in the corrected dynamic disturbance tag attitude response map are mapped to the corresponding control strategy entries to construct a multi-agent control strategy correspondence matrix. The multi-agent control strategy correspondence matrix is used to characterize the controllability of different actuator units under the control strategy.
[0017] A control priority weight is introduced into the matrix corresponding to the multi-agent control strategy. The control priority weight represents the control attribute parameter in the stability index matrix. The control priority weight is then weighted and fused with the matrix corresponding to the multi-agent control strategy to construct a priority-weighted control table.
[0018] The priority weighted control table is input into a deep reinforcement learning network, and a multi-agent cooperative control graph is generated based on the correlation between the actuator units.
[0019] As a preferred embodiment of the anti-tilt control method for a floating launcher based on deep reinforcement learning described in this invention, based on the multi-agent cooperative control graph, the cooperative action instruction set of all actuator units under the current marine environmental disturbance parameters is analyzed by utilizing the degree of cooperation and control priority weights among different actuator units.
[0020] The coordinated action instruction set is compared and verified with the priority weighted control table. Instruction conflicts are identified by a conflict detection algorithm based on a rule base, and arbitration is performed according to the control priority weights in the priority weighted control table to generate the final anti-tilt control instruction sequence.
[0021] The final anti-tilt control command sequence is sent to the corresponding actuator unit via the control system communication bus, as follows:
[0022] The balance tank actuator allocates ballast water, the thruster actuator provides active thrust, and the attitude adjustment actuator performs attitude fine-tuning.
[0023] By deploying inertial measurement units and sensor networks on the hull, the new motion state parameters of the floating launcher at sea after command execution are collected in real time, including the actual roll angle, actual pitch angle and actual heave displacement. These parameters are compared with the stability index matrix, and an instant reward signal is generated based on the attitude recovery time, maximum tilt angle limit and energy consumption upper limit index. The instant reward signal is calculated by weighting the attitude deviation penalty term, energy consumption penalty term and stability reward term.
[0024] As a preferred embodiment of the anti-tilt control method for a floating launcher based on deep reinforcement learning described in this invention, an experience tuple is constructed based on the final anti-tilt control command sequence, the current marine environment disturbance parameters, the new motion state parameters, and the instantaneous reward signal. The experience tuple is then stored in the experience replay pool of the deep reinforcement learning network based on a first-in-first-out strategy.
[0025] When the training cycle is triggered, a small batch of experience tuples are randomly sampled from the experience replay pool. The deep reinforcement learning algorithm is used to update the policy network and value network parameters of the deep reinforcement learning network with the goal of minimizing the temporal difference error between the predicted reward and the target reward.
[0026] The updated value network parameters are hot-updated online to the strategy generation engine deployed in the control system, providing feedback and enabling adaptive control and adjustment.
[0027] Compared with existing technologies, the beneficial effects achieved by this invention are as follows: The anti-tilt control method for a floating launcher based on deep reinforcement learning provided by this invention, by collecting basic performance parameters of actuators such as the balance chamber, thrusters, and attitude adjustment mechanisms and constructing a multi-source state feature set, achieves a quantitative description and unified encoding of the actuator capabilities, enabling subsequent strategy training to have reliable input and shortening the convergence cycle. Subsequently, by introducing ocean disturbance parameters such as waves, wind speed, and ocean currents, and combining them with stability indicators such as attitude recovery time, maximum tilt angle, and energy consumption limit for matching and correction, it ensures that the generated dynamic disturbance label attitude response map can accurately reflect the real working conditions and safety constraints, thereby improving the robustness of the strategy in extreme environments. Furthermore, by quantifying the motion capability parameters and combining them with three types of multimodal anti-disturbance strategies—main control, auxiliary, and emergency recovery—the invention constructs a multimodal anti-disturbance system. A control strategy scheduling matrix was constructed to achieve layered responses to different disturbance levels, improving the system's fault tolerance and emergency response capabilities. Based on this, the execution mechanism units were mapped to control strategy entries, and priority weights were introduced to generate a multi-agent cooperative control graph, effectively avoiding conflicting commands and improving the coordination and execution efficiency of multi-mechanism collaboration. Subsequently, the final anti-tilt control command sequence was generated by parsing the action command set based on the cooperative control graph and verifying it with a rule base. An instant reward signal was generated by the inertial measurement unit and sensor network, realizing closed-loop control of command-execution-evaluation, improving the safety and transparency of control. Finally, by constructing experience tuples from the control results, disturbance, state, and reward information and updating the deep reinforcement learning network, adaptive iterative optimization of the strategy was achieved, enabling the system to continuously adapt to complex and changing sea conditions. Overall, this invention not only improves the safety, robustness, and energy efficiency of the floating launch device, but also possesses good scalability and engineering verifiability, thus significantly enhancing the device's stable operation capability in actual complex sea conditions. Attached Figure Description
[0028] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used together with the embodiments of the invention to explain the invention and do not constitute a limitation thereof.
[0029] Figure 1 This is a schematic diagram illustrating the steps of the anti-tilt control method for a floating launcher at sea based on deep reinforcement learning, as described in this invention. Detailed Implementation
[0030] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0031] Please see Figure 1 In this first embodiment: a method for anti-tilt control of a floating launcher at sea based on deep reinforcement learning is provided, which includes the following steps:
[0032] Step S1: Collect basic performance parameters of different actuator units from the back end of the control system of the floating launcher at sea, and construct a multi-source state feature set.
[0033] Specifically, basic performance parameters of different actuator units are collected from the back end of the control system of the floating launcher at sea. The actuator units include a ballast tank actuator unit, a thruster actuator unit, and an attitude adjustment actuator unit. The basic performance parameters of the ballast tank actuator unit include ballast water tank capacity, leveling response time, and energy consumption parameters. The basic performance parameters of the thruster actuator unit include thrust magnitude, steering angle, and endurance. The basic performance parameters of the attitude adjustment actuator unit include pitch angle adjustment range, roll angle adjustment accuracy, and response delay parameters.
[0034] Furthermore, based on the basic performance parameters of the balance chamber actuator, the thruster actuator, and the attitude adjustment actuator, a multi-source state feature set is constructed, as follows:
[0035] The numbers are assigned according to a unified status identifier code, which includes a three-segment combination code consisting of a unit category prefix, a device number, and a performance serial number. A correspondence is established based on the motion state parameters and control attribute parameters of the actuator unit. The motion state parameters include the roll angle, pitch angle, and heave displacement of the actuator unit. The control attribute parameters include the control mode, available time period, and action priority of the actuator unit.
[0036] The multi-source state feature set is input into a deep reinforcement learning network. An initial attitude response map is generated based on the correlation between the actuator units. The correlation between the actuator units is calculated by normalization and Euclidean distance to determine the matching degree of the basic performance parameters and control attribute parameters between the actuator units. The correlation is then weighted by a preset distance weight factor of the motion state parameters to calculate the coordination degree between different actuator units.
[0037] It should be noted that by calculating the matching degree between actuators based on normalization and Euclidean distance and combining motion state distance weighting factors, a quantitative representation of the degree of cooperation between units is realized, thereby generating an initial attitude response map. This allows the strategy to identify natural cooperative pairs and inefficient combinations in the initial stage, providing engineering priors for subsequent action space construction and strategy initialization. Furthermore, it can accelerate strategy convergence, reduce the exploration of ineffective control combinations, and reduce energy consumption and stability losses caused by cooperative misjudgment.
[0038] Step S2: Preset marine environmental disturbance parameters and construct a dynamic disturbance tag attitude response map; construct a stability index matrix and match it with the stability index matrix to generate a corrected dynamic disturbance tag attitude response map.
[0039] Specifically, preset marine environmental disturbance parameters, including wave height parameters, wind speed and direction parameters, and ocean current intensity parameters; introduce the marine environmental disturbance parameters into all actuator units in the initial attitude response diagram, and correspond them with the unified state identifier code to construct a dynamic disturbance label attitude response diagram;
[0040] A stability index matrix is constructed, which includes attitude recovery time, maximum tilt angle limit, and energy consumption limit. The dynamic disturbance tag attitude response map is matched with the stability index matrix, and the basic performance parameters of the actuator unit are corrected by weighted matching to generate a corrected dynamic disturbance tag attitude response map.
[0041] It should be noted that by pre-setting ocean disturbance parameters (wave height, wind speed / direction, and current intensity) and mapping the disturbances to the initial attitude response map to form a dynamic disturbance label map, and then using a stability index matrix (attitude recovery time, maximum tilt angle, and energy limit) for weighted matching correction, coupled modeling of disturbance conditions and safety constraints is achieved. Embedding engineering safety boundaries into the scenario generation and parameter correction process before training / simulation and online decision-making allows for the screening of action combinations that could lead to exceeding limits or excessive energy consumption before policy generation, and feeds real-world constraints back into subsequent policy generation. This significantly reduces the probability of exceeding limits (over-tilting or excessive energy consumption) events, improves the robustness and verifiability of the policy under extreme or unseen sea conditions, and facilitates risk verification and compliance testing.
[0042] Step S3: Based on the corrected dynamic disturbance label attitude response map, construct the actuator action space set; introduce a multimodal disturbance rejection strategy into the actuator action space set, construct a control strategy scheduling matrix, and generate a multimodal control strategy distribution map.
[0043] Specifically, based on the corrected dynamic disturbance tag attitude response map, the leveling response time parameter of the balance cabin execution unit and the response delay parameter of the attitude adjustment execution unit are extracted; the motion capability parameters are uniformly quantified to construct the motion space set of the execution mechanism;
[0044] A multimodal anti-disturbance strategy is introduced into the action space set of the actuator, which includes a main control strategy, an auxiliary strategy, and an emergency recovery strategy; and the multimodal anti-disturbance strategy is bound to the unified state identifier code to construct a control strategy scheduling matrix.
[0045] The control strategy scheduling matrix is input into a deep reinforcement learning network, and a multimodal control strategy distribution map is generated based on the correlation between motion state parameters and disturbance intensity parameters between actuator units.
[0046] It should be noted that by extracting key action capability parameters (such as leveling response time and response delay) from the corrected dynamic disturbance label map and uniformly quantifying the action capabilities, the action space of the actuator is constructed; three types of multimodal disturbance rejection strategies—main control, auxiliary, and emergency recovery—are introduced and bound to unified coding to generate a control strategy scheduling matrix and a multimodal strategy distribution map, thereby realizing the hierarchical structure of the action space and the structured backup of strategies.
[0047] By stratifying control strategies by function and urgency level, it is easy to prioritize the use of efficient main control under normal conditions and automatically switch to auxiliary or recovery strategies in the event of equipment degradation or emergencies; at the same time, it provides clear strategy boundaries for simulation verification, playback review and fault diagnosis.
[0048] This step improves the speed and reliability of response to sudden disturbances, reduces the probability of task interruption caused by single point of failure, and makes the strategy easier to verify in engineering and formally review.
[0049] Step S4: Based on the multimodal control strategy distribution map, map all actuator units in the corrected dynamic disturbance tag attitude response map to the corresponding control strategy entries, construct a multi-agent control strategy correspondence matrix, introduce control priority weights, and generate a multi-agent cooperative control map.
[0050] Specifically, based on the multimodal control strategy distribution map, all actuator units in the corrected dynamic disturbance tag attitude response map are mapped to the corresponding control strategy entries to construct a multi-agent control strategy correspondence matrix. The multi-agent control strategy correspondence matrix is used to characterize the controllability of different actuator units under the control strategy.
[0051] A control priority weight is introduced into the matrix corresponding to the multi-agent control strategy. The control priority weight represents the control attribute parameter in the stability index matrix. The control priority weight is then weighted and fused with the matrix corresponding to the multi-agent control strategy to construct a priority-weighted control table.
[0052] The priority weighted control table is input into a deep reinforcement learning network, and a multi-agent cooperative control graph is generated based on the correlation between the actuator units.
[0053] It should be noted that by mapping all actuators and control strategy entries in the revised response diagram, a multi-agent control strategy correspondence matrix is constructed, and control priority weights based on stability indicators are introduced. These are then integrated into a priority-weighted control table and a multi-agent collaborative control diagram is generated, realizing a collaborative decision-making framework based on "controllability-importance" ranking. This provides a quantitative basis for conflict detection, arbitration, and task allocation when multiple actuators operate in parallel; the priority weights ensure that key stability actions are executed first, while incorporating engineering indicators such as energy consumption constraints and recovery timing into the decision-making logic.
[0054] This step reduces the frequency of issuing conflicting commands simultaneously, improves collaboration efficiency and security margins, and maintains overall platform stability when resources are limited or some equipment degrades.
[0055] Step S5: Based on the multi-agent cooperative control diagram, generate the final anti-tilt control command sequence; send the final anti-tilt control command sequence to the corresponding actuator unit via the control system communication bus; generate an instant reward signal.
[0056] Specifically, based on the multi-agent cooperative control diagram, the cooperative action instruction set of all actuator units under the current marine environmental disturbance parameters is analyzed by utilizing the degree of cooperation and control priority weights among different actuator units.
[0057] The coordinated action instruction set is compared and verified with the priority weighted control table. Instruction conflicts are identified by a conflict detection algorithm based on a rule base, and arbitration is performed according to the control priority weights in the priority weighted control table to generate the final anti-tilt control instruction sequence.
[0058] The final anti-tilt control command sequence is sent to the corresponding actuator unit via the control system communication bus, as follows:
[0059] The balance tank actuator allocates ballast water, the thruster actuator provides active thrust, and the attitude adjustment actuator performs attitude fine-tuning.
[0060] By deploying inertial measurement units and sensor networks on the hull, the new motion state parameters of the floating launcher at sea after command execution are collected in real time, including the actual roll angle, actual pitch angle and actual heave displacement. These parameters are compared with the stability index matrix, and an instant reward signal is generated based on the attitude recovery time, maximum tilt angle limit and energy consumption upper limit index. The instant reward signal is calculated by weighting the attitude deviation penalty term, energy consumption penalty term and stability reward term.
[0061] It should be noted that the cooperative action instruction set is obtained by parsing the multi-agent cooperative control graph. After verification and arbitration by comparing the rule-based conflict detection algorithm with the priority weighted control table, the final anti-roll control instruction sequence is generated and sent to the actuators in real time via the communication bus. At the same time, the ship's inertial measurement unit and sensor network are used to collect the motion state after execution and calculate the real-time reward signal based on attitude deviation, energy consumption, and stability reward, thus realizing a closed loop of execution-perception-evaluation. An engineered conflict check is added before the policy is issued to ensure that the command is executable and does not violate safety constraints. The execution result is converted into a reinforcement learning feedback signal in real time for online policy optimization.
[0062] This step avoids issuing unexecutable or dangerous commands, improves execution transparency and auditability, and optimizes the strategy towards reasonable energy consumption and rapid recovery of stability through real-time reward guidance.
[0063] Step S6: Based on the final anti-tilt control command sequence, current ocean environment disturbance parameters, new motion state parameters, and the instantaneous reward signal, construct an experience tuple; update the policy network and value network parameters of the deep reinforcement learning network, and provide feedback for adaptive control and adjustment.
[0064] Specifically, based on the final anti-tilt control command sequence, current ocean environment disturbance parameters, new motion state parameters, and the instantaneous reward signal, an experience tuple is constructed, and the experience tuple is stored in the experience replay pool of the deep reinforcement learning network based on a first-in-first-out strategy.
[0065] When the training cycle is triggered, a small batch of experience tuples are randomly sampled from the experience replay pool. The deep reinforcement learning algorithm is used to update the policy network and value network parameters of the deep reinforcement learning network with the goal of minimizing the temporal difference error between the predicted reward and the target reward.
[0066] The updated value network parameters are then hot-updated online to the strategy generation engine deployed in the control system and fed back to step S3 for adaptive control and adjustment.
[0067] It should be noted that by forming experience tuples from the final instruction sequence, current disturbance, new motion state, and immediate reward, and storing them in the experience replay pool according to a first-in-first-out strategy, and periodically sampling small batches of data to minimize temporal difference error to update the strategy and value network, the system is then updated online to the deployment engine and fed back to step S3. This achieves a closed loop of online-offline hybrid training and adaptive deployment. While maintaining the continuity of on-site control, strategy learning and parameter updates are completed, enabling the system to continuously adapt to non-stationary sea states, equipment aging, or changing operating conditions, reducing reliance on manual parameter readjustment.
[0068] This step improves long-term robustness and adaptability, shortens recovery time in distressed states, reduces maintenance and debugging costs, and enables rapid iteration and verification of strategy quality through a hot update mechanism.
[0069] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus.
[0070] Finally, it should be noted that the above descriptions are merely preferred embodiments of the present invention and are not intended to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments or make equivalent substitutions for some of the technical features. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A method for anti-tilting control of a floating launcher at sea based on deep reinforcement learning, characterized in that, The method includes the following steps: Step S1: Collect basic performance parameters of different actuator units from the back end of the control system of the floating launcher at sea, and construct a multi-source state feature set; Step S2: Preset marine environmental disturbance parameters and construct a dynamic disturbance tag attitude response map; construct a stability index matrix and match it with the stability index matrix to generate a corrected dynamic disturbance tag attitude response map; Step S3: Based on the corrected dynamic disturbance label attitude response map, construct the actuator action space set; introduce a multimodal disturbance rejection strategy into the actuator action space set, construct a control strategy scheduling matrix, and generate a multimodal control strategy distribution map; Step S4: Based on the multimodal control strategy distribution map, map all actuator units in the corrected dynamic disturbance tag attitude response map to the corresponding control strategy entries, construct a multi-agent control strategy correspondence matrix, introduce control priority weights, and generate a multi-agent cooperative control map. Step S5: Based on the multi-agent cooperative control diagram, generate the final anti-tilt control command sequence; send the final anti-tilt control command sequence to the corresponding actuator unit via the control system communication bus; generate an instant reward signal; Step S6: Based on the final anti-tilt control command sequence, current ocean environment disturbance parameters, new motion state parameters, and the instantaneous reward signal, construct an experience tuple; update the policy network and value network parameters of the deep reinforcement learning network, and provide feedback for adaptive control and adjustment.
2. The anti-tilting control method for a floating launcher based on deep reinforcement learning according to claim 1, characterized in that, The specific implementation process of step S1 includes: Basic performance parameters of different actuator units are collected from the back end of the control system of the floating launcher at sea. The actuator units include a ballast tank actuator unit, a thruster actuator unit, and an attitude adjustment actuator unit. The basic performance parameters of the ballast tank actuator unit include ballast water tank capacity, leveling response time, and energy consumption parameters. The basic performance parameters of the thruster actuator unit include thrust magnitude, steering angle, and endurance. The basic performance parameters of the attitude adjustment actuator unit include pitch angle adjustment range, roll angle adjustment accuracy, and response delay parameters.
3. The anti-tilting control method for a floating launcher based on deep reinforcement learning according to claim 2, characterized in that, The specific implementation process of step S1 also includes: Based on the basic performance parameters of the balance cabin actuator, the thruster actuator, and the attitude adjustment actuator, a multi-source state feature set is constructed, as follows: The numbers are assigned according to a unified status identifier code, which includes a three-segment combination code consisting of a unit category prefix, a device number, and a performance serial number. A correspondence is established based on the motion state parameters and control attribute parameters of the actuator unit. The motion state parameters include the roll angle, pitch angle, and heave displacement of the actuator unit. The control attribute parameters include the control mode, available time period, and action priority of the actuator unit. The multi-source state feature set is input into a deep reinforcement learning network. An initial attitude response map is generated based on the correlation between the actuator units. The correlation between the actuator units is calculated by normalization and Euclidean distance to determine the matching degree of the basic performance parameters and control attribute parameters between the actuator units. The correlation is then weighted by a preset distance weight factor of the motion state parameters to calculate the coordination degree between different actuator units.
4. The anti-tilting control method for a floating launcher based on deep reinforcement learning according to claim 3, characterized in that, The specific implementation process of step S2 includes: Preset marine environmental disturbance parameters, including wave height parameters, wind speed and direction parameters, and ocean current intensity parameters; introduce the marine environmental disturbance parameters into all actuator units in the initial attitude response diagram, and correspond them with the unified state identifier code to construct a dynamic disturbance label attitude response diagram; A stability index matrix is constructed, which includes attitude recovery time, maximum tilt angle limit, and energy consumption limit. The dynamic disturbance tag attitude response map is matched with the stability index matrix, and the basic performance parameters of the actuator unit are corrected by weighted matching to generate a corrected dynamic disturbance tag attitude response map.
5. The anti-tilting control method for a floating launcher based on deep reinforcement learning according to claim 4, characterized in that, The specific implementation process of step S3 includes: Based on the corrected dynamic disturbance tag attitude response map, the leveling response time parameter of the balance cabin execution unit and the response delay parameter of the attitude adjustment execution unit are extracted; the motion capability parameters are uniformly quantified to construct the motion space set of the execution mechanism; A multimodal anti-disturbance strategy is introduced into the action space set of the actuator, which includes a main control strategy, an auxiliary strategy, and an emergency recovery strategy; and the multimodal anti-disturbance strategy is bound to the unified state identifier code to construct a control strategy scheduling matrix. The control strategy scheduling matrix is input into a deep reinforcement learning network, and a multimodal control strategy distribution map is generated based on the correlation between motion state parameters and disturbance intensity parameters between actuator units.
6. The anti-tilting control method for a floating launcher based on deep reinforcement learning according to claim 5, characterized in that, The specific implementation process of step S4 includes: Based on the multimodal control strategy distribution map, all actuator units in the corrected dynamic disturbance tag attitude response map are mapped to the corresponding control strategy entries to construct a multi-agent control strategy correspondence matrix. The multi-agent control strategy correspondence matrix is used to characterize the controllability of different actuator units under the control strategy. A control priority weight is introduced into the matrix corresponding to the multi-agent control strategy. The control priority weight represents the control attribute parameter in the stability index matrix. The control priority weight is then weighted and fused with the matrix corresponding to the multi-agent control strategy to construct a priority-weighted control table. The priority weighted control table is input into a deep reinforcement learning network, and a multi-agent cooperative control graph is generated based on the correlation between the actuator units.
7. The anti-tilting control method for a floating launcher based on deep reinforcement learning according to claim 6, characterized in that, The specific implementation process of step S5 includes: Based on the multi-agent cooperative control diagram, the cooperative action instruction set of all actuators under the current marine environmental disturbance parameters is analyzed by utilizing the degree of cooperation and control priority weights among different actuator units. The coordinated action instruction set is compared and verified with the priority weighted control table. Instruction conflicts are identified by a conflict detection algorithm based on a rule base, and arbitration is performed according to the control priority weights in the priority weighted control table to generate the final anti-tilt control instruction sequence. The final anti-tilt control command sequence is sent to the corresponding actuator unit via the control system communication bus, as follows: The balance tank actuator allocates ballast water, the thruster actuator provides active thrust, and the attitude adjustment actuator performs attitude fine-tuning. By deploying inertial measurement units and sensor networks on the hull, the new motion state parameters of the floating launcher at sea after command execution are collected in real time, including the actual roll angle, actual pitch angle and actual heave displacement. These parameters are compared with the stability index matrix, and an instant reward signal is generated based on the attitude recovery time, maximum tilt angle limit and energy consumption upper limit index. The instant reward signal is calculated by weighting the attitude deviation penalty term, energy consumption penalty term and stability reward term.
8. The anti-tilting control method for a floating launcher based on deep reinforcement learning according to claim 7, characterized in that, The specific implementation process of step S6 includes: Based on the final anti-tilt control command sequence, current ocean environment disturbance parameters, new motion state parameters, and the instantaneous reward signal, an experience tuple is constructed, and the experience tuple is stored in the experience replay pool of the deep reinforcement learning network based on a first-in-first-out strategy. When the training cycle is triggered, a small batch of experience tuples are randomly sampled from the experience replay pool. The deep reinforcement learning algorithm is used to update the policy network and value network parameters of the deep reinforcement learning network with the goal of minimizing the temporal difference error between the predicted reward and the target reward. The updated value network parameters are then hot-updated online to the strategy generation engine deployed in the control system and fed back to step S3 for adaptive control and adjustment.
9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that: When the processor executes the computer program, it implements the steps of the anti-tilt control method for a floating launcher based on deep reinforcement learning as described in any one of claims 1 to 8.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by the processor, it implements the steps of the anti-tilt control method for a floating marine launcher based on deep reinforcement learning as described in any one of claims 1 to 8.