Marine rocket launching platform surge interference resisting method based on reinforcement learning

By optimizing control strategies through real-time multi-source data fusion and reinforcement learning algorithms, the problem of longitudinal surge interference on marine rocket launch platforms under complex sea conditions was solved, achieving high-precision suppression and improved adaptability, thus ensuring the stability of the platform and the safety of rocket launches.

CN121956544APending Publication Date: 2026-05-01LUDONG UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
LUDONG UNIVERSITY
Filing Date
2026-01-26
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

Existing sea-based rocket launch platforms struggle to effectively suppress longitudinal surge interference in complex sea conditions, exhibiting insufficient control precision and poor adaptability. Traditional control methods are ill-suited to dynamically adapt to rapidly changing sea conditions, thus impacting launch accuracy and safety.

Method used

By fusing real-time multi-source data to construct a platform state feature vector, using reinforcement learning algorithms to generate control action suggestions, and combining safety rule constraints, the control strategies of the thrusters, ballast systems, and mooring devices are optimized to suppress longitudinal surge interference.

Benefits of technology

It achieves high-precision, adaptive longitudinal surge interference suppression, improves the platform's position and attitude stability in complex sea conditions, enhances the robustness and adaptability of the control strategy, and improves the safety and reliability of rocket launches.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121956544A_ABST
    Figure CN121956544A_ABST
Patent Text Reader

Abstract

The invention discloses an anti-surge interference method for an offshore rocket launching platform based on reinforcement learning. The anti-surge interference method comprises the following steps of collecting multi-source data of the platform in real time; carrying out fusion processing on the collected data to obtain a platform position, a posture and a motion state; estimating a longitudinal surge disturbance state of the platform; a real-time state feature vector is constructed after a disturbance state estimation result is processed; inputting the state feature vector into a soft actor commentator algorithm controller to obtain a platform control action suggestion value; checking and adjusting a control action suggestion value according to a safety rule; and executing a final control action to suppress longitudinal surge interference in real time. According to the invention, the precision and adaptability of surge disturbance resistance of the marine rocket launching platform are effectively improved, and the stability of the position and attitude of the platform is remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

A Surge Resistance Method for Marine Rocket Launch Platforms Based on Reinforcement Learning Technical Field

[0001] This invention relates to the field of shipbuilding and marine engineering technology, and in particular to a method for resisting surge interference on a marine rocket launch platform based on reinforcement learning. Background Technology

[0002] Sea-based rocket launch platforms are important space launch facilities, offering advantages such as being far from land, reducing launch safety risks, and improving orbital entry efficiency. In recent years, they have received increasing attention and application in the aerospace field. Sea-based rocket launch platforms typically consist of floating or semi-submersible structures. Due to the complex and variable sea conditions, the platforms are highly susceptible to interference from external environmental factors such as wind and waves. In particular, longitudinal wave surges can significantly affect the platform's position and attitude, thereby posing a serious threat to the accuracy, safety, and success rate of rocket launches.

[0003] To address the aforementioned sea state disturbances, feedback control methods are currently the primary approach, including proportional-integral-derivative (PID) control based on control theory and linear quadratic control based on modern control theory. This method mainly involves real-time measurement of the platform's attitude and position errors, using feedback control laws to calculate the necessary adjustment actions, and then applying these actions to the thrusters, ballast system, or mooring devices to resist disturbances. However, this approach exhibits significant limitations when sea conditions are relatively complex and disturbances change frequently. It typically requires sophisticated mathematical models and parameter calibration, struggles to dynamically adapt to changes in sea conditions, and suffers from prominent issues such as control response hysteresis and insufficient control precision. It is also ill-suited for effectively handling real-time surge disturbances in actual maritime launch scenarios, demonstrating significant technical limitations.

[0004] In recent years, with the rapid development of artificial intelligence and machine learning technologies, reinforcement learning has been gradually applied to the control optimization of complex systems and has initially shown good performance advantages. However, the application research of reinforcement learning-based methods in the anti-surge disturbance of marine rocket launch platforms is still in its infancy. Existing solutions generally suffer from problems such as low data utilization efficiency, insufficient exploration of control actions, and insufficient robustness of control decisions. In particular, under the conditions of rapidly changing longitudinal surge disturbances and complex interference scenarios, a mature and stable anti-interference control solution has not yet been formed.

[0005] Therefore, how to provide a method for resisting surge interference on a marine rocket launch platform based on reinforcement learning is a problem that urgently needs to be solved by those skilled in the art. Summary of the Invention

[0006] One objective of this invention is to propose a surge interference mitigation method for marine rocket launch platforms based on reinforcement learning. Addressing the problems of insufficient control precision, poor adaptability, and difficulty in stably suppressing longitudinal surge interference under complex sea conditions, this invention proposes an anti-interference control scheme that constructs a platform state feature vector through real-time multi-source data fusion, generates control action suggestion values ​​in real-time based on a reinforcement learning algorithm, and executes them after adjustment by safety rules. This invention possesses advantages such as high control precision, strong adaptability, and high stability of platform position and attitude.

[0007] The reinforcement learning-based anti-surge interference method for a marine rocket launch platform according to embodiments of the present invention includes:

[0008] By collecting multi-source data on the environment in which the marine rocket launch platform is located in real time through environmental sensors, the current real-time motion status and environmental interference information can be obtained.

[0009] Data fusion processing is performed on the current real-time motion state and environmental interference information to obtain fused observation results;

[0010] Based on the fusion observation results, the amplitude and trend of the current longitudinal surge disturbance of the marine rocket launch platform are estimated, and the longitudinal surge disturbance state estimation results are obtained.

[0011] Data processing is performed on the longitudinal surge disturbance state estimation results to obtain the real-time state feature vector;

[0012] The real-time state feature vector is input into the pre-trained soft actor critic algorithm controller, and the control action suggestion value of the sea rocket launch platform at the current moment is obtained based on the output real-time longitudinal surge disturbance control strategy.

[0013] The recommended values ​​of the control actions are tested for launch safety constraints to obtain the final control actions that meet the requirements of rocket launch safety rules;

[0014] Based on the final control action, the thrust distribution strategy distributes execution commands to the propulsion system, ballast system, and mooring device of the sea-based rocket launch platform to suppress longitudinal surge interference of the sea-based rocket launch platform.

[0015] Optionally, the multi-source data includes differential GPS measurement data, inertial measurement unit measurement data, wind speed data, and wave data.

[0016] Optionally, the step of performing data fusion processing on the current real-time motion state and environmental interference information to obtain fused observation results specifically involves:

[0017] Based on real-time received differential GPS measurement data and inertial measurement unit measurement data, dynamic joint correction of clock drift and random noise is performed through extended Kalman filtering to obtain fused observations of position and velocity status.

[0018] The angular velocity and linear acceleration data in the inertial measurement unit (IMU) measurement data are initially integrated, and the quaternion attitude calculation method is used to compensate for the IMU drift error caused by the platform motion to obtain the initial attitude estimate.

[0019] Based on the initial attitude estimate, a nonlinear coupling relationship between wind speed data, wave data and platform attitude is established, and a nonlinear state observer is constructed to correct the attitude angle and angular velocity to obtain the platform attitude correction observation.

[0020] Adaptive weighted fusion is performed on the position and velocity state fusion observations and the platform attitude correction observations to obtain the fused observation results.

[0021] Optionally, based on the fused observation results, the current longitudinal surge disturbance amplitude and variation trend of the marine rocket launch platform are estimated to obtain the longitudinal surge disturbance state estimation result, specifically as follows:

[0022] Based on the fusion observation results, the real-time displacement and velocity data of the longitudinal motion direction of the sea-based rocket launch platform were determined;

[0023] Using real-time displacement and velocity data in the longitudinal motion direction, a data sequence containing historical longitudinal motion displacement and velocity is constructed through a dynamic sliding time window method;

[0024] Empirical mode decomposition is performed on the historical longitudinal motion displacement and velocity data sequence to extract intrinsic mode components representing characteristics at different time scales. For the extracted low-frequency trend term, an autoregressive moving average model is used to predict the amplitude of longitudinal surge disturbance at the next moment.

[0025] For the high-frequency fluctuation components in the intrinsic mode components, the instantaneous frequency and amplitude changes are calculated in real time based on the Hilbert-Huang transform to determine the real-time trend of the current longitudinal surge disturbance.

[0026] The predicted longitudinal surge disturbance amplitude and real-time change trend are fused using a weighted least squares dynamic estimation method to construct a longitudinal surge disturbance state estimation equation.

[0027] Adaptive parameter updates are performed on the longitudinal surge disturbance state estimation equation to obtain the longitudinal surge disturbance state estimation result.

[0028] Optionally, the process of processing the longitudinal surge disturbance state estimation results to obtain a real-time state feature vector specifically involves:

[0029] Based on the real-time longitudinal surge disturbance state estimation results, the longitudinal surge disturbance amplitude, change trend, platform longitudinal displacement, and platform longitudinal motion velocity at the current moment are extracted as raw data.

[0030] The real-time difference between the longitudinal displacement of the platform and the amplitude of the longitudinal surge disturbance is calculated based on the raw data, and the real-time difference between the longitudinal motion velocity of the platform and the changing trend of the longitudinal surge disturbance is calculated as a preliminary data feature.

[0031] The initial data features are extended to a time series, and combined with the initial data features from continuous historical moments, to construct enhanced data features containing information from multiple moments.

[0032] An adaptive segmented dynamic normalization method is used to perform real-time normalization on the enhanced data features to obtain stable data features;

[0033] Stable data features are arranged and combined dimensionally to form a high-dimensional feature array;

[0034] Feature filtering is performed on the high-dimensional feature array, and high-dimensional feature arrays with feature importance higher than a preset threshold are selected to construct the real-time state feature vector.

[0035] Optionally, obtaining the suggested control actions for the sea-based rocket launch platform at the current moment specifically involves:

[0036] An experience replay dataset for training the soft actor critic algorithm was established based on multiple sets of pre-collected sea state environmental data and platform longitudinal surge disturbance data.

[0037] Constructing the policy network and value network of the soft actor critic algorithm;

[0038] When training the soft actor critic algorithm controller, batches of sample data are randomly selected from the experience replay dataset and input into the policy network and value network. The probability distribution of the output action of the policy network and the real-time evaluation value and uncertainty variance estimate of the output of the value network are calculated for each batch of sample data.

[0039] Based on the calculation results of batch sample data in the experience playback dataset, the gradient descent method is used to optimize the policy network parameters. The objective function jointly defined by maximizing the policy entropy of the policy network output action probability distribution and the real-time evaluation value output by the value network is obtained to obtain the soft actor critic algorithm controller.

[0040] The real-time state feature vector of the sea-based rocket launch platform is input into the soft actor critic algorithm controller in real time to obtain the real-time control action probability distribution;

[0041] Based on the probability distribution of real-time control actions, the combination of control actions with the highest probability is selected using the maximum expected probability method to obtain the suggested control action values ​​for the sea-based rocket launch platform at the current moment.

[0042] Optionally, the policy network and value network are specifically:

[0043] The strategy network is a deep neural network structure, including an input layer for receiving the real-time state feature vector of the marine rocket launch platform, three fully connected hidden layers using rectified linear unit activation functions, and an output layer for outputting the probability distribution of the current control actions of the platform's thrusters, ballast system, and mooring device. The number of neurons in the input layer is equal to the number of feature dimensions of the real-time state feature vector. The number of neurons in the first hidden layer is twice the number of neurons in the input layer, the number of neurons in the second hidden layer is half the number of neurons in the first hidden layer, and the number of neurons in the third hidden layer is half the number of neurons in the second hidden layer. The feature data output by each hidden layer is batch normalized and passed to the next layer. The number of neurons in the output layer is the same as the number of control actions corresponding to the platform's thrusters, ballast system, and mooring device. The neurons in the output layer obtain the probability value corresponding to each control action through the Softmax function, thereby outputting the probability distribution of each control action at the current moment.

[0044] The value network is a deep neural network structure, including an input layer for receiving the real-time state feature vector and current control action of the sea-based rocket launch platform, three fully connected hidden layers using rectified linear unit activation functions, and an output layer with a dual-branch structure. The number of neurons in the input layer is the sum of the number of feature dimensions of the real-time state feature vector and the number of control action dimensions. The number of neurons in the first hidden layer is twice the number of neurons in the input layer, the number of neurons in the second hidden layer is half the number of neurons in the first hidden layer, and the number of neurons in the third hidden layer is half the number of neurons in the second hidden layer. The feature data output by each hidden layer is batch normalized and passed to the next layer. The first branch of the output layer includes a fully connected linear layer, which outputs the real-time evaluation value corresponding to the longitudinal surge disturbance suppression effect. The second branch of the output layer includes a fully connected linear layer, which outputs the uncertainty variance estimate corresponding to the real-time evaluation value.

[0045] Optionally, the step of performing a launch safety constraint check on the suggested control action values ​​to obtain the final control action that meets the requirements of rocket launch safety rules specifically involves:

[0046] Obtain recommended values ​​for control actions of the sea-based rocket launch platform, and establish safety constraints required by rocket launch safety rules based on the performance parameters and working limits of the platform's thrusters, ballast system, and mooring devices.

[0047] Based on the control action recommendation values, calculate the thrust output of the platform thruster, the load adjustment of the ballast system, and the mooring force adjustment of the mooring device, and determine whether the thrust output, load adjustment, and mooring force adjustment are within the safe range allowed by the performance parameters of the corresponding devices.

[0048] If the thrust output, load adjustment, and mooring force adjustment all meet the safety range, the control action recommendation value is determined to comply with the rocket launch safety rules, and the control action recommendation value is directly output as the final control action.

[0049] If any of the thrust output, load adjustment, or mooring force adjustment exceeds the safe range allowed by the corresponding device performance parameters, the recommended control action value is determined to be inconsistent with the rocket launch safety rules, and the difference in the value exceeding the safe range is calculated.

[0050] Based on the differences in magnitude and the safe range allowed by the performance parameters of the corresponding devices, the thrust output of the platform booster, the load adjustment of the ballast system and the mooring force adjustment of the mooring device are proportionally limited and adjusted in accordance with the rocket launch safety rules until the parameters of each device meet the safe range allowed by the performance parameters.

[0051] The control actions corresponding to the thrust output of the platform thruster, the load adjustment of the ballast system, and the mooring force adjustment of the mooring device after real-time proportional amplitude limiting adjustment are taken as the final control actions.

[0052] Optionally, the distribution of execution commands to the thrusters, ballast system, and mooring devices of the sea-based rocket launch platform through the thrust distribution strategy specifically includes:

[0053] The real-time thrust requirements of the platform thruster, the real-time load adjustment requirements of the ballast system, and the real-time anchoring force requirements of the anchoring device are determined based on the final control actions.

[0054] Based on the real-time thrust demand, the thrust distribution coefficient of the platform thrusters is determined according to the rated power, maximum response rate and current load status of the platform thrusters, and the thrust distribution of each thruster is calculated.

[0055] Based on the real-time load adjustment demand, the ballast water distribution ratio is determined according to the capacity limit of each ballast compartment of the ballast system and the current load status, and the ballast water injection or discharge volume of each ballast compartment is calculated.

[0056] Based on the real-time anchoring force demand, the anchoring force distribution ratio of each anchoring device is determined according to the rated anchoring force, anchor chain tension limit and current anchoring status, and the anchor chain tension adjustment of each anchoring device is calculated.

[0057] Based on the thrust distribution of the propeller, the amount of ballast water injected or discharged from the ballast compartment, and the adjustment amount of the anchor chain tension of the mooring device, corresponding execution commands are generated.

[0058] The thruster, ballast system, and mooring device receive and execute the execution commands in real time, generating a longitudinal torque acting on the sea-based rocket launch platform to suppress longitudinal surge interference on the platform.

[0059] The beneficial effects of this invention are:

[0060] (1) This invention achieves high-precision adaptive suppression of longitudinal surge disturbances by fusing multi-source data from the platform in real time and constructing dynamic feature vectors, and by using reinforcement learning algorithms to optimize control strategies in real time. This effectively improves the anti-surge interference control accuracy and response efficiency of the marine rocket launch platform and enhances the position and attitude stability of the platform under complex sea conditions.

[0061] (2) By setting up a value network uncertainty assessment and policy entropy dynamic adjustment mechanism in the reinforcement learning controller, the present invention realizes real-time optimization and adjustment of control action decisions, significantly improves the robustness of the platform control strategy, and exhibits better adaptability and control effect in complex and ever-changing marine environments.

[0062] (3) In terms of longitudinal surge disturbance suppression in complex sea conditions, this invention effectively solves the shortcomings of the prior art in control response lag and poor adaptive capability by real-time state feature extraction, soft actor commentator algorithm decision-making and safety constraint dynamic adjustment mechanism. It breaks through the dependence of traditional control methods on the accuracy of dynamic model and parameter calibration, realizes specific and significant improvement of platform anti-surge interference technology, and effectively improves the safety and application reliability of marine rocket launch platform. Attached Figure Description

[0063] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used in conjunction with embodiments of the invention to explain the invention and do not constitute a limitation thereof. In the drawings:

[0064] Figure 1 is a flowchart of the anti-surge interference method for marine rocket launch platforms based on reinforcement learning proposed in this invention. Detailed Implementation

[0065] The present invention will now be described in further detail with reference to the accompanying drawings. These drawings are simplified schematic diagrams, illustrating only the basic structure of the invention, and therefore only show the components relevant to the invention.

[0066] Referring to Figure 1, the reinforcement learning-based method for resisting surge interference on a marine rocket launch platform includes:

[0067] By collecting multi-source data on the environment in which the marine rocket launch platform is located in real time through environmental sensors, the current real-time motion status and environmental interference information can be obtained.

[0068] The current real-time motion status and environmental interference information of the marine rocket launch platform are fused to obtain the fused observation results of the position, attitude and motion status of the marine rocket launch platform.

[0069] Based on the fused observation results, the current longitudinal surge disturbance amplitude and variation trend of the marine rocket launch platform are estimated, and the real-time longitudinal surge disturbance state estimation results are obtained.

[0070] Data processing is performed on the longitudinal surge disturbance state estimation results to construct data features that meet the input requirements of the soft actor critic algorithm, and the data features are normalized to obtain the real-time state feature vector of the sea rocket launch platform.

[0071] The real-time state feature vector of the marine rocket launch platform is input into the pre-trained soft actor critic algorithm controller. Based on the real-time longitudinal surge disturbance control strategy output by the soft actor critic algorithm controller, the control action suggestion value of the marine rocket launch platform at the current moment is obtained.

[0072] The suggested control action value is subjected to launch safety constraint verification to determine whether it complies with rocket launch safety rules. If the suggested control action value complies with rocket launch safety rules, it is output as the final control action. If the suggested control action value does not comply with rocket launch safety rules, it is adjusted according to the rocket launch safety rules to obtain the final control action that meets the requirements of rocket launch safety rules.

[0073] According to the final control action, the thrust distribution strategy distributes execution commands to the thrusters, ballast system and mooring device of the sea-based rocket launch platform. The execution commands are applied to the sea-based rocket launch platform to suppress longitudinal surge interference in real time, so as to ensure the stability of the position and attitude of the sea-based rocket launch platform.

[0074] In this embodiment, the multi-source data includes differential GPS measurement data, inertial measurement unit measurement data, wind speed data, and wave data.

[0075] In this embodiment, the data fusion processing of the current real-time motion state and environmental interference information to obtain the fused observation result specifically includes:

[0076] Based on real-time received differential GPS measurement data and inertial measurement unit measurement data, clock drift and random noise are dynamically and jointly corrected by extended Kalman filtering to obtain real-time position and velocity state fusion observations after clock error dynamic compensation.

[0077] The angular velocity and linear acceleration data from the inertial measurement unit (IMU) measurement data are initially integrated, and a quaternion attitude calculation method is used to compensate for the IMU drift error caused by platform motion in real time, thereby obtaining the initial attitude estimate of the marine rocket launch platform. The quaternion attitude calculation method involves initial integration of the three-axis angular velocity data and three-axis linear acceleration data from the IMU measurement data. Based on the three-axis angular velocity data, the real-time attitude change rate of the platform is determined. A quaternion differential equation is established according to the attitude change rate and numerically integrated and solved. The projection of the gravity vector in the platform coordinate system is determined using the three-axis linear acceleration data. Through the normalization of the quaternion attitude and the dynamic correction calculation of the gravity vector, real-time dynamic compensation for the IMU drift error caused by platform motion is achieved, thereby obtaining the initial attitude estimate of the marine rocket launch platform.

[0078] Based on the initial attitude estimation, a nonlinear coupling relationship between wind speed data, wave data, and platform attitude is established. A nonlinear state observer is constructed to dynamically and in real-time correct the attitude angle and angular velocity, obtaining the platform attitude correction observation value after compensating for sea state interference. The nonlinear coupling relationship is obtained by calculating the nonlinear interference torques of wind-induced and wave-induced interferences on the platform using real-time collected wind speed data and wave data, respectively. The real-time dynamic correction amount of the platform attitude angle is obtained according to the nonlinear coupling function relationship between the nonlinear interference torque and the platform attitude angle. The nonlinear state observer is based on an extended Kalman filter structure, using the platform attitude angle and angular velocity as the state variables to be estimated. The real-time dynamic correction amount obtained by the nonlinear coupling relationship is used as the observation input of the nonlinear state observer. The platform attitude correction observation value after compensating for sea state interference is obtained by real-time linearization of the nonlinear state equation and the nonlinear observation equation and recursive iterative calculation.

[0079] The real-time position and velocity state fusion observations and platform attitude correction observations are subjected to real-time adaptive weighted fusion. The fusion weights are dynamically determined by the real-time error covariance of the differential GPS measurement data, inertial measurement unit measurement data, wind speed data, and wave data collected by the current sensors, so as to obtain a high-confidence fusion observation result of platform position, attitude, and motion state.

[0080] In this embodiment, the step of estimating the current longitudinal surge disturbance amplitude and its changing trend on the sea-based rocket launch platform based on the fused observation results, and obtaining the longitudinal surge disturbance state estimation result, specifically involves:

[0081] Based on the fusion observation results, the real-time displacement and velocity data of the longitudinal motion direction of the sea-based rocket launch platform were determined;

[0082] Using real-time displacement and velocity data in the longitudinal motion direction, a data sequence containing historical longitudinal motion displacement and velocity is constructed through a dynamic sliding time window method. The length of the sliding time window is dynamically and adaptively determined based on the real-time estimate of the current wave cycle. The specific method for constructing the data sequence containing historical longitudinal motion displacement and velocity is as follows: zero-crossing detection is performed on the real-time longitudinal velocity data to determine the single motion cycle of the current platform's longitudinal motion; the dynamic length of the current sliding time window is calculated based on the single motion cycle; historical longitudinal displacement and velocity data that are continuously adjacent to the current moment are selected in the data storage buffer according to the dynamic length, and combined and arranged according to the time order to form the data sequence of historical longitudinal motion displacement and velocity.

[0083] Empirical mode decomposition is performed on the historical longitudinal motion displacement and velocity data sequence to extract intrinsic mode components representing characteristics at different time scales. For the extracted low-frequency trend terms, an autoregressive moving average model is used to predict the amplitude of the longitudinal surge disturbance at the next moment.

[0084] The empirical mode decomposition is as follows: for the data sequence of historical longitudinal motion displacement and velocity, all local maxima and local minima are determined respectively, and their corresponding upper and lower envelopes are constructed by cubic spline interpolation.

[0085] The average envelope is calculated based on the upper and lower envelopes, and the average envelope is subtracted from the historical longitudinal motion displacement and velocity data sequence to obtain the intermediate signal;

[0086] The intermediate signal is treated as a new processing object and the aforementioned steps of determining local extrema, constructing upper and lower envelopes and calculating average envelopes are repeated until the obtained intermediate signal satisfies the intrinsic mode component definition conditions, thereby obtaining an intrinsic mode component of the historical longitudinal motion displacement and velocity data sequence.

[0087] Remove the corresponding intrinsic mode components from the historical longitudinal motion displacement and velocity data sequence to obtain the residual signal;

[0088] The above decomposition steps are repeated by replacing the historical longitudinal motion displacement and velocity data sequence with the residual signal until the obtained residual signal meets the predetermined decomposition termination condition, and finally multiple intrinsic mode components of different orders are obtained.

[0089] For the high-frequency fluctuation component in the intrinsic mode component, the instantaneous frequency and amplitude changes are calculated in real time based on the Hilbert-Huang transform to determine the real-time trend of the current longitudinal surge disturbance.

[0090] The predicted longitudinal surge disturbance amplitude and real-time change trend are fused using a weighted least squares dynamic estimation method to construct a longitudinal surge disturbance state estimation equation that takes into account both the longitudinal surge disturbance amplitude and change trend in real time.

[0091] Real-time adaptive parameter updates are performed on the longitudinal surge disturbance state estimation equation to obtain high-precision and high-stability real-time longitudinal surge disturbance state estimation results.

[0092] In this embodiment, the data processing of the longitudinal surge disturbance state estimation results to obtain the real-time state feature vector specifically involves:

[0093] Based on the real-time longitudinal surge disturbance state estimation results, the longitudinal surge disturbance amplitude, change trend, platform longitudinal displacement, and platform longitudinal motion velocity at the current moment are extracted as raw data.

[0094] The real-time difference between the longitudinal displacement of the platform and the amplitude of the longitudinal surge disturbance is calculated based on the original data, and the real-time difference between the longitudinal motion velocity of the platform and the changing trend of the longitudinal surge disturbance is calculated as a preliminary data feature.

[0095] The initial data features are extended to a time series, and combined with the initial data features from continuous historical moments, to construct enhanced data features containing information from multiple moments.

[0096] An adaptive segmented dynamic normalization method is adopted to perform real-time normalization processing on the enhanced data features. Specifically, the segmented interval threshold for normalization is determined in real time based on the dynamic deviation between the current longitudinal surge disturbance amplitude and the historical longitudinal surge disturbance amplitude. The local mean and local standard deviation of the interval are calculated for the data points in the enhanced data features located in different intervals, and the segmented dynamic normalization is performed on the data features using the local mean and local standard deviation of the interval to obtain stable data features that can adapt to the dynamic changes in the longitudinal surge disturbance state in real time.

[0097] The stable data features are arranged and combined dimensionally to form a high-dimensional feature array that meets the input requirements of the soft actor critic algorithm;

[0098] The high-dimensional feature array is subjected to feature filtering, and high-dimensional feature arrays with feature importance higher than a preset threshold are selected to construct the final real-time status feature vector of the sea-based rocket launch platform.

[0099] In this embodiment, obtaining the control action suggestion value of the sea-based rocket launch platform at the current moment specifically involves:

[0100] An experience replay dataset for training the soft actor commentator algorithm is established based on multiple sets of pre-collected sea state environmental data and platform longitudinal surge disturbance data. The sea state environmental data includes differential GPS measurement data, inertial measurement unit measurement data, wind speed data, and wave data. The platform longitudinal surge disturbance data includes longitudinal surge disturbance amplitude, longitudinal surge disturbance trend, platform longitudinal displacement, and platform longitudinal motion velocity.

[0101] A policy network for a soft actor critic algorithm is constructed, with the input layer defined as the real-time state feature vector of a marine rocket launch platform and the output layer as the current control actions corresponding to the platform's thrusters, ballast system, and mooring device.

[0102] The value network of the soft actor critic algorithm is constructed. The input layer is defined as the real-time state feature vector and current control action of the sea rocket launch platform. The output layer adopts a dual-branch structure. The first branch outputs the real-time evaluation value corresponding to the longitudinal surge disturbance suppression effect, and the second branch outputs the uncertainty variance estimate corresponding to the real-time evaluation value.

[0103] When training the soft actor critic algorithm controller, batches of sample data are randomly selected from the experience replay dataset and input into the policy network and value network. The probability distribution of the output action of the policy network and the real-time evaluation value and uncertainty variance estimate of the output of the value network are calculated for each batch of sample data.

[0104] Based on the calculation results of the batch sample data in the aforementioned experience replay dataset, the gradient descent method is used to optimize the policy network parameters, and the objective function jointly defined by maximizing the policy entropy of the policy network output action probability distribution and the real-time evaluation value output by the value network is obtained. This yields a soft actor critic algorithm controller that has been fully trained and can make robust decisions for different longitudinal surge disturbance states.

[0105] The real-time state feature vector of the sea-based rocket launch platform is input into the soft actor critic algorithm controller in real time. Based on the real-time evaluation value output by the value network and the corresponding uncertainty variance estimate, the entropy weight of the action output of the strategy network is dynamically adjusted. When the uncertainty variance estimate is higher than a preset threshold, the entropy weight of the action output of the strategy network is increased to enhance the exploration capability of the control action. The real-time control action probability distribution of each control device of the platform is calculated using the strategy network.

[0106] Based on the real-time control action probability distribution, the combination of control actions with the highest probability is selected using the maximum expected probability method to obtain the suggested control action value for the sea-based rocket launch platform at the current moment.

[0107] In this embodiment, the policy network and value network are specifically as follows:

[0108] The strategy network is a deep neural network structure, including an input layer for receiving the real-time state feature vector of the marine rocket launch platform, three fully connected hidden layers using rectified linear unit activation functions, and an output layer for outputting the probability distribution of the current control actions of the platform's thrusters, ballast system, and mooring device. The number of neurons in the input layer is equal to the number of feature dimensions of the real-time state feature vector. The number of neurons in the first hidden layer is twice the number of neurons in the input layer, the number of neurons in the second hidden layer is half the number of neurons in the first hidden layer, and the number of neurons in the third hidden layer is half the number of neurons in the second hidden layer. The feature data output by each hidden layer is batch normalized and passed to the next layer. The number of neurons in the output layer is the same as the number of control actions corresponding to the platform's thrusters, ballast system, and mooring device. The neurons in the output layer obtain the probability value corresponding to each control action through the Softmax function, thereby outputting the probability distribution of each control action at the current moment.

[0109] The value network is a deep neural network structure, including an input layer for receiving the real-time state feature vector and current control action of the sea-based rocket launch platform, three fully connected hidden layers using rectified linear unit activation functions, and an output layer with a dual-branch structure. The number of neurons in the input layer is the sum of the number of feature dimensions of the real-time state feature vector and the number of control action dimensions. The number of neurons in the first hidden layer is twice the number of neurons in the input layer, the number of neurons in the second hidden layer is half the number of neurons in the first hidden layer, and the number of neurons in the third hidden layer is half the number of neurons in the second hidden layer. The feature data output by each hidden layer is batch normalized and passed to the next layer. The first branch of the output layer includes a fully connected linear layer, outputting the real-time evaluation value corresponding to the longitudinal surge disturbance suppression effect. The second branch of the output layer includes a fully connected linear layer, outputting the uncertainty variance estimate corresponding to the real-time evaluation value.

[0110] In this embodiment, the step of performing a launch safety constraint check on the recommended control action values ​​to obtain the final control action that meets the requirements of rocket launch safety rules specifically involves:

[0111] Obtain recommended values ​​for control actions of the sea-based rocket launch platform, and establish safety constraints required by rocket launch safety rules based on the performance parameters and working limits of the platform's thrusters, ballast system, and mooring devices.

[0112] Based on the control action recommendation values, calculate the thrust output of the platform thruster, the load adjustment of the ballast system, and the mooring force adjustment of the mooring device, and determine whether the thrust output, load adjustment, and mooring force adjustment are within the safe range allowed by the performance parameters of the corresponding devices.

[0113] If the thrust output, load adjustment, and mooring force adjustment all meet the safety range, then the control action recommendation value is determined to comply with the rocket launch safety rules, and the control action recommendation value is directly output as the final control action.

[0114] If any of the thrust output, load adjustment, or mooring force adjustment exceeds the safe range allowed by the corresponding device performance parameters, the recommended control action value is determined to be inconsistent with the rocket launch safety rules, and the difference in the value exceeding the safe range is calculated.

[0115] Based on the differences in the values ​​and the safe range allowed by the corresponding device performance parameters, the thrust output of the platform thruster, the load adjustment of the ballast system and the mooring force adjustment of the mooring device are proportionally limited and adjusted in accordance with the rocket launch safety rules until the adjusted parameters of each device meet the safe range allowed by the performance parameters.

[0116] The control actions corresponding to the thrust output of the platform thruster, the load adjustment of the ballast system, and the mooring force adjustment of the mooring device after real-time proportional amplitude limiting adjustment are taken as the final control actions to meet the requirements of rocket launch safety rules.

[0117] In this embodiment, the distribution of execution commands to the thrusters, ballast system, and mooring device of the sea-based rocket launch platform through the thrust distribution strategy specifically includes:

[0118] Based on the final control action, the real-time thrust requirement of the platform thruster, the real-time load adjustment requirement of the ballast system, and the real-time anchoring force requirement of the anchoring device are determined respectively.

[0119] Based on the real-time thrust demand, the thrust distribution coefficient of the platform thrusters is determined according to the rated power, maximum response rate and current load status of the platform thrusters, and the thrust distribution of each thruster is calculated.

[0120] Based on the real-time load adjustment demand, the ballast water distribution ratio is determined according to the capacity limit of each ballast compartment of the ballast system and the current load status, and the ballast water injection or discharge volume of each ballast compartment is calculated.

[0121] Based on the real-time anchoring force demand, the anchoring force distribution ratio of each anchoring device is determined according to the rated anchoring force, anchor chain tension limit and current anchoring status, and the anchor chain tension adjustment amount of each anchoring device is calculated.

[0122] Based on the thrust distribution of the thruster, the ballast water injection or discharge rate of the ballast chamber, and the anchor chain tension adjustment of the mooring device, corresponding execution commands are generated and sent to the corresponding thruster, ballast system, and mooring device. The execution commands are obtained by matching the calculated thrust distribution of the thruster, the ballast water injection or discharge rate of the ballast chamber, and the anchor chain tension adjustment of the mooring device with preset mapping relationships for thruster thrust control commands, ballast system pump valve on / off status control commands, and mooring device anchor winch working status control commands. The control commands are then converted into power adjustment signals for the thruster motor, on / off control signals for each pump valve in the ballast system, and anchor winch release speed and braking torque adjustment signals for the mooring device, ultimately generating execution commands that can be accurately identified and executed by each device.

[0123] The thruster, ballast system, and mooring device receive and execute the execution commands in real time, generating a longitudinal torque acting on the sea-based rocket launch platform to suppress longitudinal surge interference on the platform in real time and ensure that the position and attitude of the sea-based rocket launch platform are always kept within a preset stable range.

[0124] Example 1:

[0125] To verify the feasibility of this invention in practice, it was applied to a longitudinal surge disturbance suppression mission conducted by a space agency on a sea-based rocket launch platform. This verified the platform's longitudinal stability and rocket launch safety control effectiveness under complex sea conditions. The platform, composed of a semi-submersible floating structure, frequently experiences significant longitudinal surge disturbances due to wind and waves in real-world applications. This severely impacts the safety and accuracy of rocket launches, posing a significant challenge to the execution of sea-based rocket launch missions.

[0126] In previous practical applications, the classical feedback control method relied on mathematical models of platform dynamics. The parameters of these models required frequent adjustments and calibrations for specific sea conditions, resulting in low control accuracy. Especially when longitudinal surge disturbances changed rapidly, the system struggled to respond accurately in real time, and the platform's longitudinal offset and tilt angle frequently exceeded permissible safety limits, leading to a high launch cancellation rate. While modern optimized control methods can improve control performance to some extent, their real-time response to disturbance changes under complex sea conditions remains insufficient, resulting in sluggish control actions and suppression effects that still fall short of practical application requirements.

[0127] To address the above issues, this embodiment employs the reinforcement learning-based anti-surge interference method for marine rocket launch platforms proposed in this invention. First, environmental sensors installed on the platform collect real-time measurement data from the Differential Global Positioning System (GPS), Inertial Measurement Unit (INS), wind speed, and wave data. Then, the collected multi-source data is fused to obtain a fused observation result of the platform's position, attitude, and motion state. Based on the fused observation result, the system uses the Empirical Mode Decomposition (EMD) method to estimate the current longitudinal surge disturbance amplitude and its changing trend, and utilizes a weighted least squares dynamic estimation method to update the parameters of the disturbance state estimation equation in real time, significantly improving the accuracy of longitudinal surge disturbance estimation.

[0128] Subsequently, this embodiment calculates the real-time state feature vector based on the longitudinal surge disturbance state estimation results, and inputs the state feature vector into a pre-trained reinforcement learning controller. The policy network employs a three-layer deep neural network structure, with each hidden layer using rectified linear unit activation functions. The number of nodes decreases layer by layer and undergoes batch normalization processing, ultimately outputting a control action probability distribution. The value network also employs a three-layer deep neural network structure with a dual-branch output structure, obtaining the real-time disturbance suppression evaluation value and the corresponding uncertainty variance estimate, respectively. The entropy weights of the policy network are dynamically adjusted based on the uncertainty variance of the value network, effectively improving the robustness of control action decisions and adaptability to unknown sea conditions.

[0129] In implementation, the real-time control action recommendations output by the reinforcement learning controller underwent rigorous launch safety constraint testing. The real-time requirements of the platform thrusters, ballast system, and mooring device were determined, and the control actions were adjusted in real-time based on the performance parameters of the corresponding devices to ensure that the final output control actions fully met rocket launch safety rules. Subsequently, accurate execution commands were obtained through pre-determined mapping relationships for thrust control commands, ballast system pump and valve on / off status control commands, and mooring device anchor winch operating status control commands. These commands were then sent to the corresponding actuators for real-time execution, significantly suppressing longitudinal surge interference on the platform.

[0130] To verify the actual control effect of the present invention, this embodiment selected sample time periods under five different sea state conditions in a real application scenario, and compared and analyzed the predictive control effect and measured effect of the method proposed in this invention. The results are summarized in Table 1 below:

[0131] Table 1 Comparison of the predictive control effect and the measured effect of the method of the present invention.

[0132] Sample Number | Measured Longitudinal Displacement Amplitude (m) | Predicted Longitudinal Displacement Amplitude (m) | Measured Longitudinal Tilt Angle (°) | Predicted Longitudinal Tilt Angle (°) | Control Action Response Time (s) | A1 | 0.35 | 0.32 | 0.95 | 0.92 | 0.46 | A2 | 0.41 | 0.38 | 1.12 | 1.09 | 0.48 | A3 | 0.29 | 0.27 | 0.82 | 0.78 | 0.42 | A4 | 0.53 | 0.49 | 1.35 | 1.31 | 0.51 | A5 | 0.38 | 0.36 | 1.03 | 1.01 | 0.45 surface

[0133] As can be clearly observed from the data in Table 1, the difference between the predicted longitudinal displacement amplitude and the measured value of the method of the present invention is less than 0.05m, and the difference between the predicted longitudinal tilt angle and the measured value is less than 0.05°, demonstrating extremely high prediction accuracy. Furthermore, the control action response time of the method of the present invention is less than 0.52s, significantly better than the response time of more than 0.9s of traditional feedback control methods. This fully demonstrates the outstanding advantages of the method of the present invention in real-time responsiveness and control accuracy, effectively ensuring the position and attitude stability of the platform under complex sea conditions, and successfully improving the execution rate and safety of actual launch missions.

[0134] The results of the above embodiments clearly demonstrate the technical advantages of this invention in real-time accurate estimation of longitudinal surge disturbance states, robust decision-making by reinforcement learning controllers, and adjustment of safety constraint control actions. The method and control system proposed in this invention have a rigorous structure, clear logic, and strong feasibility, and can significantly improve the control accuracy, adaptability, and stability of marine rocket launch platforms in complex sea conditions, possessing strong practical application and promotion value.

[0135] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and inventive concept of the present invention, should be covered within the scope of protection of the present invention.

Claims

1. A method for resisting surge interference on a marine rocket launch platform based on reinforcement learning, characterized in that, include: Multi-source data on the environment of the sea-based rocket launch platform are collected in real time by environmental sensors to obtain the current real-time motion status and environmental interference information; the current real-time motion status and environmental interference information are fused to obtain fused observation results; Based on the fusion observation results, the amplitude and trend of the current longitudinal surge disturbance of the marine rocket launch platform are estimated to obtain the longitudinal surge disturbance state estimation results; the longitudinal surge disturbance state estimation results are processed to obtain the real-time state feature vector. The real-time state feature vector is input into the pre-trained soft actor critic algorithm controller, and the control action suggestion value of the sea rocket launch platform at the current moment is obtained based on the output real-time longitudinal surge disturbance control strategy. The recommended values ​​of the control actions are tested for launch safety constraints to obtain the final control actions that meet the requirements of rocket launch safety rules. Based on the final control actions, the execution commands are distributed to the thrusters, ballast system and mooring devices of the sea-based rocket launch platform through a thrust distribution strategy to suppress longitudinal surge interference of the sea-based rocket launch platform.

2. The method for resisting surge interference on a marine rocket launch platform based on reinforcement learning according to claim 1, characterized in that, The multi-source data includes differential GPS measurement data, inertial measurement unit measurement data, wind speed data, and wave data.

3. The method for resisting surge interference on a marine rocket launch platform based on reinforcement learning according to claim 1, characterized in that, The process of fusing current real-time motion state and environmental interference information to obtain fused observation results specifically involves: Based on real-time received differential GPS measurement data and inertial measurement unit (IMU) measurement data, dynamic joint correction of clock drift and random noise is performed using extended Kalman filtering to obtain fused position and velocity state observations; initial integration is performed on the angular velocity and linear acceleration data in the IMU measurement data, and a quaternion attitude calculation method is used to compensate for the IMU drift error caused by platform motion to obtain initial attitude estimates; based on the initial attitude estimates, a nonlinear coupling relationship is established between wind speed data, wave data, and platform attitude, and a nonlinear state observer is constructed to correct the attitude angle and angular velocity to obtain platform attitude correction observations; adaptive weighted fusion is performed on the fused position and velocity state observations and the platform attitude correction observations to obtain fused observation results.

4. The method for resisting surge interference on a marine rocket launch platform based on reinforcement learning according to claim 1, characterized in that, The process involves estimating the current longitudinal surge disturbance amplitude and trend of the marine rocket launch platform based on the fused observation results, thereby obtaining the longitudinal surge disturbance state estimation result. Specifically, this includes: determining the real-time displacement and velocity data of the longitudinal motion direction of the marine rocket launch platform based on the fused observation results; constructing a data sequence containing historical longitudinal motion displacement and velocity using the real-time displacement and velocity data of the longitudinal motion direction through a dynamic sliding time window method; performing empirical mode decomposition on the historical longitudinal motion displacement and velocity data sequence to extract intrinsic mode components representing characteristics at different time scales, and using an autoregressive moving average model to predict the longitudinal surge disturbance amplitude at the next moment for the extracted low-frequency trend term; calculating the instantaneous frequency and amplitude changes of the high-frequency fluctuation components in the intrinsic mode components based on the Hilbert-Huang transform to determine the real-time trend of the current longitudinal surge disturbance; and fusing the predicted longitudinal surge disturbance amplitude and real-time trend using a weighted least squares dynamic estimation method to construct the longitudinal surge disturbance state estimation equation. Adaptive parameter updates are performed on the longitudinal surge disturbance state estimation equation to obtain the longitudinal surge disturbance state estimation result.

5. The method for resisting surge interference on a marine rocket launch platform based on reinforcement learning according to claim 1, characterized in that, The process of processing the longitudinal surge disturbance state estimation results to obtain a real-time state feature vector involves: extracting the longitudinal surge disturbance amplitude, trend, platform longitudinal displacement, and platform longitudinal velocity at the current moment as raw data based on the real-time longitudinal surge disturbance state estimation results; calculating the real-time difference between the platform longitudinal displacement and the longitudinal surge disturbance amplitude based on the raw data, and calculating the real-time difference between the platform longitudinal velocity and the trend of longitudinal surge disturbance as preliminary data features. The initial data features are extended to a time series format and combined with the initial data features from continuous historical moments to construct enhanced data features containing information from multiple moments. An adaptive segmented dynamic normalization method is used to perform real-time normalization on the enhanced data features to obtain stable data features. The stable data features are arranged and combined in dimensions to form a high-dimensional feature array. The high-dimensional feature array is then filtered to select high-dimensional feature arrays whose feature importance is higher than a preset threshold to construct a real-time state feature vector.

6. The method for resisting surge interference on a marine rocket launch platform based on reinforcement learning according to claim 1, characterized in that, The specific steps for obtaining the control action suggestion value of the marine rocket launch platform at the current moment are as follows: An experience replay dataset for training the soft actor / critic algorithm is established based on multiple pre-collected sets of sea state environmental data and longitudinal surge disturbance data of the platform; the policy network and value network of the soft actor / critic algorithm are constructed; during the training of the soft actor / critic algorithm controller, batch sample data is randomly extracted from the experience replay dataset and input into the policy network and value network, and the probability distribution of the policy network output action and the real-time evaluation value and uncertainty variance estimate of the value network output corresponding to each batch of sample data are calculated; based on the calculation results of the batch sample data in the experience replay dataset, the gradient descent method is used to optimize the policy network parameters, maximizing the objective function jointly defined by the policy entropy of the policy network output action probability distribution and the real-time evaluation value of the value network output, to obtain the soft actor / critic algorithm controller; the real-time state feature vector of the marine rocket launch platform is input into the soft actor / critic algorithm controller in real time to obtain the real-time control action probability distribution; based on the real-time control action probability distribution, the maximum expected probability method is used to select the control action combination with the highest probability to obtain the control action suggestion value of the marine rocket launch platform at the current moment.

7. The method for resisting surge interference on a marine rocket launch platform based on reinforcement learning according to claim 6, characterized in that, The policy network and value network are specifically defined as follows: The policy network is a deep neural network structure, including an input layer for receiving the real-time state feature vector of the marine rocket launch platform, three fully connected hidden layers using rectified linear unit activation functions, and an output layer for outputting the probability distribution of the current control actions of the platform's thrusters, ballast system, and mooring device. The number of neurons in the input layer is equal to the number of feature dimensions of the real-time state feature vector. The number of neurons in the first hidden layer is twice the number of neurons in the input layer, the number of neurons in the second hidden layer is half the number of neurons in the first hidden layer, and the number of neurons in the third hidden layer is half the number of neurons in the second hidden layer. The feature data output by each hidden layer is batch normalized and passed to the next layer. The number of neurons in the output layer is the same as the number of control actions corresponding to the platform's thrusters, ballast system, and mooring device. The neurons in the output layer obtain the probability values ​​corresponding to each control action through a Softmax function, thereby outputting the probability distribution of the current control actions. The probability distribution of each control action at the current moment is output. The value network is a deep neural network structure, including an input layer for receiving the real-time state feature vector of the sea-based rocket launch platform and the current control action, three fully connected hidden layers using rectified linear unit activation functions, and an output layer with a dual-branch structure. The number of neurons in the input layer is the sum of the number of feature dimensions of the real-time state feature vector and the number of control action dimensions. The number of neurons in the first hidden layer is twice the number of neurons in the input layer, the number of neurons in the second hidden layer is half the number of neurons in the first hidden layer, and the number of neurons in the third hidden layer is half the number of neurons in the second hidden layer. The feature data output by each hidden layer is batch normalized and passed to the next layer. The first branch of the output layer includes a linear fully connected layer, which outputs the real-time evaluation value corresponding to the longitudinal surge disturbance suppression effect. The second branch of the output layer includes a linear fully connected layer, which outputs the uncertainty variance estimate corresponding to the real-time evaluation value.

8. The method for resisting surge interference on a marine rocket launch platform based on reinforcement learning according to claim 1, characterized in that, The process of verifying the control action recommendations against launch safety constraints to obtain the final control actions that meet the requirements of rocket launch safety rules involves: acquiring the control action recommendations for the sea-based rocket launch platform; establishing safety constraints required by rocket launch safety rules based on the performance parameters and operational limits of the platform's thrusters, ballast system, and mooring devices; calculating the thrust output of the platform's thrusters, the load adjustment of the ballast system, and the mooring force adjustment of the mooring devices based on the control action recommendations; and determining whether each of these parameters falls within the safe range allowed by the corresponding device's performance parameters. If all three parameters meet the safety range, the control action recommendations are deemed to comply with rocket launch safety rules, and the control actions are then executed. The suggested values ​​are directly output as the final control actions. If any of the thrust output, ballast adjustment, or mooring force adjustment exceeds the safe range allowed by the corresponding device performance parameters, the suggested control action value is determined to be inconsistent with rocket launch safety rules, and the difference in value exceeding the safe range is calculated. Based on the difference in value and the safe range allowed by the corresponding device performance parameters, the thrust output of the platform thruster, the ballast system ballast adjustment, and the mooring force adjustment of the mooring device are proportionally limited and adjusted according to rocket launch safety rules until all device parameters after adjustment meet the safe range allowed by the performance parameters. The control actions corresponding to the thrust output of the platform thruster, the ballast system ballast adjustment, and the mooring force adjustment of the mooring device after real-time proportional limiting adjustment are taken as the final control actions.

9. The method for resisting surge interference on a marine rocket launch platform based on reinforcement learning according to claim 1, characterized in that, The thrust distribution strategy for allocating execution commands to the thrusters, ballast system, and mooring devices of the sea-based rocket launch platform specifically involves: determining the real-time thrust requirements of the platform's thrusters, the real-time load adjustment requirements of the ballast system, and the real-time mooring force requirements of the mooring devices based on the final control actions; determining the thrust distribution coefficient of the thrusters based on the real-time thrust requirements, according to the rated power, maximum response rate, and current load status of the platform's thrusters, and calculating the thrust distribution amount for each thruster; and determining the ballast water distribution ratio based on the real-time load adjustment requirements, according to the capacity limitations of each ballast compartment in the ballast system and the current load status. For example, the system calculates the ballast water injection or discharge volume for each ballast compartment; based on the real-time mooring force demand, it determines the mooring force distribution ratio for each mooring device according to the rated mooring force, the anchor chain tension limit, and the current mooring state, and calculates the anchor chain tension adjustment for each mooring device; based on the thrust distribution of the thruster, the ballast water injection or discharge volume of the ballast compartment, and the anchor chain tension adjustment for the mooring device, it generates corresponding execution commands; the thruster, ballast system, and mooring device receive and execute the execution commands in real time, generating a longitudinal torque acting on the sea-based rocket launch platform to suppress longitudinal surge interference experienced by the platform.