A reinforcement learning-based dynamic control method and system for ship composite propulsion

By using a reinforcement learning-based approach and combining environmental perception and propulsion system state data, the control strategy of the ship's composite propulsion system is dynamically adjusted, solving the problem of the disconnect between environmental perception and decision-making in existing technologies, and achieving stable and precise control of the ship in complex sea conditions.

CN122085692APending Publication Date: 2026-05-26GUANGDONG OCEAN UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
GUANGDONG OCEAN UNIVERSITY
Filing Date
2026-03-02
Publication Date
2026-05-26

AI Technical Summary

Technical Problem

Existing control methods for ship composite propulsion systems fail to effectively integrate environmental perception and decision control, making it difficult to achieve dynamic adaptive control in complex sea conditions and affecting the stability and accuracy of ship operation.

Method used

By employing a reinforcement learning-based approach, environmental perception data, navigation attitude data, and propulsion system status data of the ship's location are acquired. Dynamic decision-making is then carried out using a master control agent and an execution agent to generate a multi-degree-of-freedom thrust vector. Combined with risk analysis, dynamic adjustments are made to achieve adaptive control of the ship's composite propulsion system.

Benefits of technology

It improves the stability and control precision of ships in complex sea conditions, enhances their adaptability in complex environments, and ensures navigation safety and energy efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122085692A_ABST
    Figure CN122085692A_ABST
Patent Text Reader

Abstract

This invention discloses a method and system for dynamic control of ship composite propulsion based on reinforcement learning. The method includes acquiring environmental perception data of the area where the ship is located; acquiring the ship's navigation attitude data, navigation mission data, and operational status data of each propulsion system to obtain a composite state space data sequence of the ship in combination with the environmental perception data; outputting the ship's multi-degree-of-freedom thrust vector based on the composite state space data sequence and the master control agent; obtaining the target operational status data of each propulsion system based on the multi-degree-of-freedom thrust vector, the composite state data, and the execution agent of each propulsion system; and, when controlling the propulsion system based on the target operational status data, acquiring the risk analysis results of each propulsion system to determine the data adjustment amount of each propulsion system, so as to realize the dynamic adaptive control of the ship's composite propulsion system under complex sea conditions and improve the stability of ship operation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of ship control technology, specifically to a dynamic control method and system for ship composite propulsion based on reinforcement learning. Background Technology

[0002] As the primary carriers of maritime transportation and operations, the control performance of a ship's propulsion system directly affects navigation safety, operational efficiency, and energy consumption. With the development of larger, faster, and more intelligent ships, modern ships generally adopt composite propulsion systems, which are distributed propulsion devices composed of multiple azimuth thrusters, side thrusters, etc., achieving stable navigation of the ship through the coordinated work of multiple thrusters.

[0003] Existing technologies mostly employ control methods such as PID control, model predictive control, or rule-based energy management strategies to control ship composite propulsion systems. These control algorithms often distribute the total thrust demand evenly among the propellers through a preset thrust distribution algorithm, or adjust the propeller operating mode based on speed and sea state using a lookup table method.

[0004] However, the aforementioned control methods lack dynamic integration of environmental perception and decision-making control, leading to a severe disconnect between the ship's propulsion control and the actual environmental conditions, thus affecting the ship's control accuracy in complex sea states. Furthermore, these methods often employ fixed-weight thrust allocation algorithms, failing to capture the dynamic changes in coupling between propulsion systems as the environment alters. This results in discrepancies between the actual and expected outputs of each thruster, causing not only thrust loss but also hindering precise ship control. Additionally, these methods typically separate risk monitoring and control decision-making systems, preventing timely adjustments to control strategies to maintain stable ship operation when risks occur. In summary, existing control methods make it difficult for ship composite propulsion systems to achieve dynamic adaptive control in complex sea states, thereby impacting ship operational stability. Summary of the Invention

[0005] This invention discloses a reinforcement learning-based dynamic control method and system for ship composite propulsion, which is used to realize dynamic adaptive control of ship composite propulsion system under complex sea conditions and improve the stability of ship operation.

[0006] To achieve the above objectives, this invention discloses a reinforcement learning-based dynamic control method for ship composite propulsion, comprising: Acquire sea surface images of the area where the ship is located, and identify wave height and current velocity data in the sea surface images to obtain environmental perception data of the area where the ship is located; The ship's navigation attitude data, navigation mission data, and operational status data of each propulsion system in the ship are acquired, and combined with the environmental perception data to obtain a composite state space data sequence of the ship. The composite state space data sequence is input into the master control agent pre-trained by the reinforcement learning algorithm, so that the master control agent can output the multi-degree-of-freedom thrust vector of the ship. The multi-degree-of-freedom thrust vector and the composite state space data sequence are input into the execution agent that is pre-trained by the reinforcement learning algorithm corresponding to each of the propulsion systems, so that the execution agent generates the target running state data corresponding to each of the propulsion systems according to the preset parameterized action space. When controlling the propulsion system based on the target operating status data, the operating fluctuation data of each propulsion system is acquired, and the operating fluctuation data is analyzed for risk based on a pre-built risk model library to obtain the risk analysis results of each propulsion system. Based on the risk analysis results, the data adjustment amount for each of the propulsion systems is determined, so as to realize the dynamic control of each of the propulsion systems according to the data adjustment amount and the target operating status data.

[0007] This invention discloses a reinforcement learning-based dynamic control method for ship composite propulsion. It obtains environmental perception data representing sea state characteristics from sea surface image recognition of the ship's location, and combines this with navigation attitude data, navigation mission data, and operational status data of each propulsion system to obtain a composite state space data sequence representing the ship's multi-dimensional operational state. This sequence is then used by a master control agent to perform global thrust decoupling decisions, resulting in a multi-degree-of-freedom thrust vector representing the overall control objective of the ship. Based on the multi-degree-of-freedom thrust vector and the composite state data, local collaborative allocation is performed by the execution agent corresponding to each propulsion system to obtain target operational state data for each system. During dynamic control, risk pattern matching analysis is performed based on the operational fluctuation data of each propulsion system to obtain risk analysis results representing abnormal operational states. The data adjustment amount used for dynamic control is adjusted based on the risk analysis results and ship operational constraint data. Finally, based on the adaptive and precise data adjustment amount and target operational state data, dynamic adaptive control of the ship's composite propulsion system under complex sea conditions is achieved, improving the stability of ship operation.

[0008] On the other hand, the present invention discloses a ship composite propulsion dynamic control system based on reinforcement learning, including an environmental perception module, a composite state module, a main control module, an execution module, a monitoring module, and a dynamic adjustment control module; The environmental perception module is used to acquire sea surface images of the area where the ship is located, and identify wave height data and current velocity data in the sea surface images to obtain environmental perception data of the area where the ship is located. The composite state module is used to acquire the ship's navigation attitude data, navigation mission data, and operating status data of each propulsion system in the ship, so as to obtain the composite state space data sequence of the ship in combination with the environmental perception data. The main control module is used to input the composite state space data sequence into the main control agent obtained by pre-training the reinforcement learning algorithm, so as to output the multi-degree-of-freedom thrust vector of the ship through the main control agent; The execution module is used to input the multi-degree-of-freedom thrust vector and the composite state space data sequence into the execution agent that is pre-trained by the reinforcement learning algorithm corresponding to each of the propulsion systems, so that the execution agent generates target running state data corresponding to each of the propulsion systems according to the preset parameterized action space. The monitoring module is used to acquire operational fluctuation data of each propulsion system when controlling the propulsion system based on the target operational status data, and to perform risk analysis on the operational fluctuation data based on a pre-built risk model library to obtain the risk analysis result of each propulsion system. The dynamic adjustment control module is used to determine the data adjustment amount of each of the propulsion systems based on the risk analysis results, so as to realize the dynamic control of each of the propulsion systems based on the data adjustment amount and the target operating status data.

[0009] This invention discloses a reinforcement learning-based dynamic control system for ship composite propulsion. It obtains environmental perception data representing sea state characteristics from sea surface image recognition of the ship's location, and combines this with navigation attitude data, navigation mission data, and operational status data of each propulsion system to obtain a composite state space data sequence representing the ship's multi-dimensional operational state. This allows for subsequent global thrust decoupling decision-making through a master control agent based on the composite state space data sequence to obtain a multi-degree-of-freedom thrust vector representing the overall control objective of the ship. Furthermore, based on the multi-degree-of-freedom thrust vector and the composite state data, local collaborative allocation is performed through the execution agent corresponding to each propulsion system to obtain target operational state data for each system. During dynamic control, risk pattern matching analysis is performed based on the operational fluctuation data of each propulsion system to obtain risk analysis results representing abnormal operational states. The data adjustment amount used for dynamic control is adjusted based on the risk analysis results and ship operational constraint data. Finally, based on the adaptive and precise adjustment of the data adjustment amount and the target operational state data, dynamic adaptive control of the ship's composite propulsion system under complex sea conditions is achieved, improving the stability of ship operation. Attached Figure Description

[0010] Figure 1 This is a flowchart illustrating a reinforcement learning-based dynamic control method for ship composite propulsion disclosed in an embodiment of the present invention. Figure 2 This is a schematic diagram of a ship composite propulsion dynamic control system based on reinforcement learning disclosed in an embodiment of the present invention. Detailed Implementation

[0011] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0012] See Figure 1 To achieve dynamic adaptive control of a ship's composite propulsion system under complex sea conditions and improve the stability of ship operation, this invention provides a reinforcement learning-based dynamic control method for ship composite propulsion, which mainly includes: Step 101: Acquire a sea surface image of the area where the ship is located, and identify the wave height data and current velocity data in the sea surface image to obtain environmental perception data of the area where the ship is located; Step 102: Obtain the ship's navigation attitude data, navigation mission data, and operating status data of each propulsion system in the ship, so as to obtain the ship's composite state space data sequence in combination with the environmental perception data; Step 103: Input the composite state space data sequence into the master control agent pre-trained by the reinforcement learning algorithm, so as to output the multi-degree-of-freedom thrust vector of the ship through the master control agent; Step 104: Input the multi-degree-of-freedom thrust vector and the composite state space data sequence into the execution agent that has been pre-trained by the reinforcement learning algorithm corresponding to each of the propulsion systems, so that the execution agent generates the target running state data corresponding to each of the propulsion systems according to the preset parameterized action space; Step 105: When controlling the propulsion system based on the target operating status data, obtain the operating fluctuation data of each propulsion system, and perform risk analysis on the operating fluctuation data based on the pre-built risk model library to obtain the risk analysis result of each propulsion system; Step 106: Determine the data adjustment amount for each of the propulsion systems based on the risk analysis results, so as to realize the dynamic control of each propulsion system based on the data adjustment amount and the target operating status data.

[0013] In this embodiment, reinforcement learning is used to train the master control agent and the execution agent, enabling them to dynamically adjust control strategies according to the environment and ship state. The composite propulsion system consists of multiple thrusters of different or the same type, such as a ship propulsion device composed of a main propeller and various waterjet propulsion units, used to provide multi-degree-of-freedom thrust and torque to the ship through collaborative work, thereby achieving heading, speed, and attitude control. The master control agent is a pre-trained intelligent decision-making model used to receive the ship's composite state-space data sequence and output the overall multi-degree-of-freedom thrust vector of the ship according to the current environment and mission requirements, responsible for global thrust distribution and coordination. The execution agent is an intelligent decision-making model corresponding to each propulsion system, used to receive the multi-degree-of-freedom thrust vector output by the master control agent and the ship's composite state data, and to adjust the thrust according to the preset parameterized dynamic... The system provides a working space for generating specific target operating state data (such as power distribution, target angle, target speed, etc.) for the corresponding propulsion system. A risk pattern library is a pre-built database used to store standard feature templates and corresponding risk levels corresponding to various potential risk patterns in ship propulsion systems. This allows for matching and analysis of propulsion system operation fluctuation data to identify potential risks and assess their severity. The multi-degree-of-freedom thrust vector is a vector representation of the thrust or torque required by the ship in multiple degrees of freedom, such as longitudinal thrust components, lateral thrust components, and yaw moment components. It serves as the global target for local control by each executing agent. The parameterized action space is the adjustable parameter range used by the executing agent to generate target operating state data for the propulsion system. It defines the operable range and fine adjustment capability of the propulsion unit under different operating modes.

[0014] In this embodiment, step 101 includes: acquiring a sequence of sea surface images of the area where the ship is located in real time, and performing grayscale processing and motion blur correction on the sea surface image sequence to obtain a preprocessed sea surface image sequence; performing frame-by-frame recognition on the preprocessed sea surface image sequence using a pre-trained sea state visual analysis model to obtain a wave surface texture feature sequence, a wave crest distribution feature sequence, and a surface flow field stripe feature sequence; obtaining the spatial frequency of the wave surface texture feature sequence and the dominant wave period of the wave crest distribution feature sequence to invert the wave spectrum of the area where the ship is located based on the spatial frequency; obtaining wave height data of the area where the ship is located based on the dominant wave period, the wave spectrum, and a pre-calibrated pixel texture scale mapping relationship; obtaining the morphological offset between any two adjacent surface flow field stripe features in the surface flow field stripe feature sequence to obtain the pixel displacement between any two adjacent surface flow field stripe features based on the morphological offset and a preset optical flow analysis method; obtaining flow velocity data of the area where the ship is located based on the pixel texture scale mapping relationship, the sampling frequency of the preprocessed sea surface image sequence, and all pixel displacements corresponding to the surface flow field stripe feature sequence; and fusing the wave height data and flow velocity data to obtain environmental perception data of the area where the ship is located.

[0015] This embodiment continuously acquires a sequence of sea surface images of the area where the ship is located using visual devices such as high-definition cameras, infrared cameras, or multispectral sensors installed on the ship. Then, the color images are converted to grayscale images using a weighted average or simple averaging method to perform grayscale processing on the sea surface image sequence. Simultaneously, considering the impact of ship movement on image acquisition, motion blur correction is applied to the obtained grayscale images using deconvolution-based or deep learning-based methods to obtain a preprocessed sea surface image sequence. The aforementioned grayscale processing simplifies image information, reduces computational complexity, and preserves brightness information, which is helpful for texture and edge detection; while motion blur correction eliminates image blur caused by ship movement or camera shake, improving image clarity and ensuring the quality and accuracy of subsequent image analysis. After obtaining the preprocessed sea surface image sequence, a sea state visual analysis model trained using a deep learning-based convolutional neural network is employed to process the images and identify their features. Specifically, after constructing the initial sea state visual analysis model based on the convolutional neural network, it is trained using a large dataset of labeled sea surface images. This training enables the model to identify and segment wave surface textures, wave crest lines, and current field stripe regions. The sea state visual analysis model can automatically and efficiently extract key visual features related to sea states from the preprocessed sea surface images, providing foundational data for subsequent wave height and current velocity calculations.

[0016] After obtaining wave surface texture feature sequences, wave crest line distribution feature sequences, and surface flow field stripe feature sequences through a sea state visual analysis model, a two-dimensional Fourier transform is performed on the wave surface texture feature sequences to analyze their spectrum and identify the dominant frequency components with concentrated energy as spatial frequencies. Time series analysis, such as autocorrelation analysis or power spectral density estimation, is performed on the wave crest line distribution feature sequences to determine their periodicity, thereby obtaining the dominant wave period. Then, using empirical models or numerical inversion methods, combined with spatial frequencies and dominant wave periods, the wave spectrum is constructed or corrected through iterative optimization or lookup tables. The above steps of obtaining spatial frequencies and wave spectra transform the visual features extracted from the image into physically meaningful wave parameters, improving the accuracy of environmental perception data acquisition. To quantify the abstract wave spectrum and periodic information into specific wave height values, this embodiment utilizes pixel-texture scale mapping relationships to calculate wave height and current velocity. The pixel-texture scale mapping relationship serves as a bridge connecting image pixels with actual physical dimensions, ensuring the accuracy of wave height calculation. Specifically, the significant wave height can be calculated using the zero-order moment of the wave spectrum, where the calculation depends on the dominant wave period and spatial frequency. The pixel-texture scale mapping relationship can be obtained by calibration on a reference object of known size, for example, by measuring the actual size of a specific area of ​​the sea surface using a laser rangefinder or GPS-assisted measurement and mapping it to the pixel size in the image. When calculating the flow velocity, the morphological offset between any two adjacent surface flow field stripe features in the surface flow field stripe feature sequence is obtained, capturing the positional changes of the flow field stripes between consecutive frames. This offset is then combined with optical flow analysis to calculate the flow velocity. Optical flow analysis is a mature image motion estimation technique capable of accurately calculating pixel-level displacements, providing a foundation for subsequent flow velocity calculations. Image registration techniques (such as SIFT and SURF algorithms based on feature point matching, or cross-correlation algorithms based on region matching) are used to calculate the overall or local offset of the flow field stripes between adjacent frames to obtain the morphological offset.

[0017] Finally, to convert the pixel displacement information in the image into actual physical flow velocity, this embodiment combines the temporal and spatial scales of image acquisition to calculate the physical flow velocity. The sampling frequency provides temporal dimension information, and the pixel texture scale mapping relationship provides spatial dimension information. Combining these two with pixel displacement allows for accurate calculation of the flow velocity. The flow velocity can be calculated as: (pixel displacement) The pixel texture scale mapping is calculated as (1 / sampling frequency). Here, pixel displacement is the pixel movement distance obtained from optical flow analysis, the pixel texture scale mapping is the scaling factor that converts pixel distance to actual physical distance, and the sampling frequency is the frame rate of the image sequence. By fusing wave height and current velocity data, environmental perception data for the ship's location is obtained.

[0018] The above steps involve real-time acquisition and grayscale conversion and motion blur correction, ensuring the clarity and consistency of the original image data and laying the foundation for subsequent accurate analysis. A pre-trained sea state visual analysis model identifies wave surface texture, wave crest lines, and surface flow field stripe features frame-by-frame, utilizing advanced visual recognition capabilities to avoid the limitations of traditional methods and improve the accuracy of feature extraction. Furthermore, wave spectra are retrieved by acquiring spatial frequency and dominant wave period, and wave height data is calculated by combining pixel texture scale mapping. Flow velocity data is obtained by calculating pixel displacement through morphological offset and optical flow analysis, achieving precise quantification from image features to physical parameters. This multi-step, refined processing workflow fully considers the complexity of the sea surface environment and the challenges of image acquisition, ensuring the accuracy and real-time performance of wave height and flow velocity data. Ultimately, the environmental perception data formed by integrating these high-precision data provides high-quality and high-reliability environmental input for the subsequent dynamic control of ship propulsion by the main control agent. This enables the main control agent to more accurately assess the current sea state and output a multi-degree-of-freedom thrust vector that is more adaptable to environmental changes, thereby significantly improving the ship's operational stability in complex sea conditions.

[0019] In this embodiment, step 102 includes: acquiring the ship's multi-sensor time synchronization protocol to determine a clock source for a unified time reference based on the multi-sensor time synchronization protocol; acquiring the ship's navigation attitude data and the operating status data of each propulsion system in the ship to determine the acquisition time corresponding to each data item in the environmental perception data, navigation attitude data, and operating status data according to the clock source; performing time interpolation processing on the environmental perception data, navigation attitude data, and operating status data according to the acquisition time to use the environmental perception data, navigation attitude data, and operating status data at the same time as the composite state data of the ship at each time; and acquiring the ship's navigation mission data. The system performs semantic encoding on the navigation mission data to map it into a mission semantic embedding vector of a preset length. It then retrieves the ship operation constraint data corresponding to the navigation mission data from a pre-configured mission target mapping table. For any given moment, it normalizes the composite state data at that moment to convert it into a composite state data vector of a preset dimension. The system concatenates the mission semantic embedding vector, the ship operation constraint data, and the composite state data vector to obtain the ship's composite state space data at that moment. Finally, based on the composite state space data corresponding to each moment within a preset time window, it obtains the ship's composite state space data sequence.

[0020] In this embodiment, to ensure that data collected by different sensors on the ship have consistent timestamps, thereby enabling precise ship control using multi-source data from the same moment, clock synchronization of multiple sensors can be achieved by acquiring the ship's multi-sensor time synchronization protocol. Specifically, when using the multi-sensor time synchronization protocol for clock synchronization, the clocks of each sensor can be synchronized via a network based on a network time protocol or a precise time protocol, ensuring that all sensors remain synchronized with a high-precision master clock; alternatively, the timing function of the Global Positioning System (GPS) can be utilized, using the positioning receiver as a unified clock source to provide high-precision pulse-second signals and time information to all sensors, achieving multi-sensor time synchronization. This embodiment lays the foundation for accurate alignment of all subsequent data by determining a unified time reference clock source. Under the guidance of the unified clock source, each piece of data acquired from different sensors or data interfaces, such as environmental perception data output by the aforementioned environmental perception module, navigation attitude data including speed and heading acquired by the attitude sensor, and operational status data acquired by the propulsion system controller, is precisely timestamped, ensuring that each data point is associated with its actual acquisition time, thus providing an accurate basis for subsequent data alignment.

[0021] When collecting data to enable adaptive control of a ship, the different data acquisition frequencies of various sensors may differ, or data transmission delays may exist. Directly using raw data may not provide all the necessary information at the same time. Therefore, this embodiment employs temporal interpolation techniques. For example, linear interpolation can be used to estimate the data value at the target time based on data values ​​and timestamps from adjacent time points. Alternatively, more complex algorithms such as spline interpolation or nearest neighbor interpolation can be used to more accurately fill in missing data or align data with different sampling rates. This temporal interpolation technique aligns the originally inconsistent environmental perception data, navigation attitude data, and operational status data to a unified point in time, forming complete and time-synchronized composite status data. This comprehensively reflects the ship's internal and external environment and its own operational status at a specific moment. After acquiring the clock-aligned composite status data corresponding to any given time, a pre-trained language model based on the Transformer architecture is used to encode the text-based navigation mission instructions, converting them into fixed-dimensional embedding vectors. This allows the machine to obtain an understandable and processable numerical representation while ensuring that its semantic information is preserved. Next, to efficiently and accurately obtain the constraints related to a specific navigation mission, the ship operation constraint data corresponding to the navigation mission data is retrieved from a pre-configured mission target mapping table, thereby ensuring that the control decisions meet the actual operational requirements of the ship. The mission target mapping table is a key-value pair structure stored in a database or configuration file, where the key is the identifier of the navigation mission and the value is the corresponding set of ship operation constraint data. Finally, to eliminate the influence of different dimensions of the composite state data and scale the data to a uniform range, a normalization method can be used to linearly scale each feature value in the composite state data to the range of [0, 1] or [-1, 1]. Alternatively, a Z-score normalization method can be used to convert each feature value in the composite state data into a distribution with a mean of 0 and a standard deviation of 1, so as to concatenate the mission semantic embedding vector, the ship operation constraint data, and the composite state data vector to obtain the composite state space data of the ship at that moment. Finally, to capture the dynamic evolution of the ship's state, this embodiment obtains time-series data composed of composite state space data from multiple consecutive moments within a preset time window. For example, a fixed-length time window can be set, such as data from the past N time steps, and the sequence can be continuously updated using a sliding window. This sequence can provide time-series information such as ship motion trends, environmental change trends, and propulsion system operation history, providing rich contextual information for the subsequent reinforcement learning-based decision-making by the master agent, enabling it to better understand and predict the ship's dynamic behavior.

[0022] The above steps, by introducing a multi-sensor time synchronization protocol and a unified clock source, provide a precise time reference for all data acquisition, fundamentally eliminating time drift and data misalignment. Based on this, by precisely determining the acquisition time of each data point and performing time interpolation, seamless alignment of data from different sources and with different sampling frequencies is achieved, generating time-synchronized and comprehensive composite state data. Furthermore, the navigation mission data is encoded and fused with the composite state data, so that the state representation not only includes the ship's physical state and environmental information but also incorporates mission objectives, providing the agent with a more comprehensive decision-making basis. Finally, by constructing a continuous composite state space data sequence, the dynamic evolution of the ship's state can be captured, providing the master control agent with rich temporal context information. This significantly improves the reinforcement learning model's ability to perceive the ship's dynamic environment and its own state, laying a solid data foundation for achieving precise and adaptive dynamic control of the ship's composite propulsion system.

[0023] In this embodiment, step 103 includes: performing time-series modeling on the composite state space data sequence using a time-series encoder pre-set in the master control agent to extract the ship's motion state evolution features and obtain a time-series hidden state vector; querying the task attention association matrix pre-set in the master control agent based on the composite state space data sequence to obtain the attention weight vector corresponding to the time-series hidden state vector; multiplying the attention weight vector and the time-series hidden state vector element-wise by the master control agent to obtain the ship's comprehensive state representation vector; outputting the longitudinal thrust feature vector, lateral thrust feature vector, and yaw moment feature vector corresponding to the comprehensive state representation vector through a multi-feature extraction network pre-set in the master control agent; and concatenating the longitudinal thrust feature vector, lateral thrust feature vector, and yaw moment feature vector by the master control agent to obtain a joint feature vector and obtaining the longitudinal and lateral thrust feature vectors corresponding to the joint feature vector. The system calculates the following parameters: longitudinal thrust coupling coefficient, longitudinal yaw coupling coefficient, and lateral yaw coupling coefficient; obtains the longitudinal thrust reference value, lateral thrust reference value, and yaw torque reference value corresponding to the longitudinal thrust feature vector, and the yaw torque reference value corresponding to the yaw torque feature vector, all through a linear activation function preset in the main control agent; obtains the lateral thrust coupling compensation amount based on the longitudinal and lateral coupling coefficients and the longitudinal thrust reference value, and obtains the first yaw torque coupling compensation amount based on the longitudinal yaw coupling coefficient and the longitudinal thrust reference value, and calculates the second yaw torque coupling compensation amount based on the lateral yaw coupling coefficient and the lateral thrust reference value; and outputs a multi-degree-of-freedom thrust vector including longitudinal thrust components, lateral thrust components, and yaw torque components through the main control agent based on the longitudinal thrust reference value, lateral thrust reference value, yaw torque reference value, lateral thrust coupling compensation amount, first yaw torque coupling compensation amount, and second yaw torque coupling compensation amount.

[0024] This embodiment incorporates a temporal encoder within the master control agent to process sequential data and capture temporal dependencies, thereby understanding the changing patterns of the ship's state over time. Specifically, the temporal encoder employs a recurrent neural network structure, which processes sequential data in parallel through a self-attention mechanism, demonstrating excellent performance in capturing long-range dependencies to output the ship's temporal hidden state vector. Simultaneously, to dynamically allocate attention to different parts of the temporal hidden state vector based on the current composite state space data sequence, highlighting the state information most relevant to the current task objective, this embodiment sets up a learnable task attention association matrix within the master control agent. This matrix is ​​optimized alongside the master control agent through a reinforcement learning training process. During the query process, similarity calculations can be performed based on the current composite state space data sequence and preset key vectors in the task attention association matrix to obtain an attention score. This score is then normalized using a probability distribution function to obtain the attention weight. When attention weights are obtained through the task attention correlation matrix, each element of the attention weight vector can be multiplied with the corresponding element of the temporal hidden state vector by the element-wise multiplication method. This can highlight or suppress certain information in the temporal hidden state vector, forming a more comprehensive state representation that is more focused on the current task and environment. This operation can be achieved through the Hadamard product of vectors.

[0025] In this embodiment, a multi-feature extraction network is set up in the main control agent. To enable the multi-feature extraction network to extract different features, it can be configured as a fully connected neural network with multiple independent branches. This decouples and extracts features related to the ship's longitudinal thrust, lateral thrust, and yaw moment requirements from the comprehensive state representation vector. Each branch is responsible for extracting thrust or moment features in a specific direction. After obtaining multiple feature vectors through the multi-feature extraction network, these feature vectors are concatenated to form a longer joint feature vector for subsequent processing. The coupling coefficients are obtained to quantify the mutual influence between thrust or moment of different degrees of freedom. These coupling coefficients can be predicted by connecting one or more fully connected layers after the joint feature vector. These layers, after training, can output coefficients representing different coupling relationships. These coefficients can be scalars or vectors, used for subsequent compensation calculations. To map abstract feature vectors to specific thrust or torque reference values, this embodiment sets a linear activation function in the master control agent. The linear activation function maps the feature vectors directly or after a simple linear transformation to the ideal thrust or torque requirement without coupling effects. Preferably, each feature vector can be followed by a fully connected layer, whose activation function is a linear activation function, directly outputting the corresponding reference value. Since different propulsion systems have coupling effects, this embodiment calculates the amounts that need to be compensated for other degrees of freedom based on the previously obtained coupling coefficients and reference thrust or torque values, i.e., obtaining the lateral thrust coupling compensation amount, the first yaw torque coupling compensation amount, and the second yaw torque coupling compensation amount, respectively. Preferably, the calculation of each compensation amount is a linear combination; for example, the lateral thrust coupling compensation amount can be equal to the longitudinal-lateral coupling coefficient multiplied by the longitudinal thrust reference value. Finally, the reference value and the calculated compensation amounts are integrated to obtain the coupled-compensated multi-degree-of-freedom thrust vector. The final thrust and torque components are the algebraic sum of the reference values ​​and the corresponding compensation amounts. For example, the final lateral thrust component can be equal to the lateral thrust baseline value plus the lateral thrust coupling compensation.

[0026] The above steps utilize a time-series encoder to perform time-series modeling on a composite state-space data sequence that has been time-synchronized, interpolated, and fused with environmental perception data, navigation attitude data, operational status data, and navigation mission data. This comprehensively captures the evolutionary characteristics of the ship's motion state under complex sea conditions, providing rich and dynamic contextual information for subsequent thrust decision-making. Next, the introduction of a task attention correlation matrix enables the controlling agent to dynamically focus on the key information most relevant to thrust decision-making within the time-series hidden state vector, based on the current navigation mission and environmental conditions, thereby enhancing adaptability to environmental dynamics and task orientation. The comprehensive state representation vector obtained through element-wise multiplication accurately reflects the comprehensive needs currently faced by the ship. Based on this, a multi-feature extraction network decouples the comprehensive state representation vector into independent feature vectors for longitudinal and lateral thrust and yaw moment, laying the foundation for subsequent coupling analysis. Crucially, by concatenating these feature vectors and obtaining dynamically changing longitudinal-lateral coupling coefficients, longitudinal-yaw coupling coefficients, and lateral-yaw coupling coefficients, this embodiment can accurately quantify the mutual influence between thrust or moment forces of different degrees of freedom. Subsequently, based on the reference values ​​of thrust or torque in each direction obtained from the linear activation function, and combined with the dynamically calculated lateral thrust coupling compensation, first yaw moment coupling compensation, and second yaw moment coupling compensation, precise dynamic compensation for the thrust coupling effect was achieved. In summary, through effective modeling and compensation of the dynamic coupling between thrust components, the accuracy of thrust distribution was significantly improved, thrust loss was reduced, thereby ensuring the accuracy and stability of ship control under complex sea conditions, and enhancing navigation safety and energy efficiency.

[0027] In this embodiment, step 104 includes: for any execution agent corresponding to a propulsion system, obtaining local operating state data of the propulsion system and adjacent local operating state data of each adjacent propulsion system within a preset domain from the composite state data; semantically encoding the local operating state data, adjacent local operating state data, and multi-degree-of-freedom thrust vectors using a preset local feature encoder in the execution agent to obtain the local state representation vector of the propulsion system and the adjacent local state representation vector of each adjacent propulsion system; obtaining the attention weight between the propulsion system and each adjacent propulsion system using a preset collaborative attention algorithm in the execution agent based on the local state representation vector and all adjacent local state representation vectors; weightedly fusing all adjacent local state representation vectors according to the attention weights to obtain the neighborhood state representation vector of the propulsion system, and weightedly summing the neighborhood state representation vector and the local state representation vector to obtain the enhanced state representation vector of the propulsion system; inputting the enhanced state representation vector into the parameterized action network of the execution agent to output the target operating state data of the propulsion system through the parameterized action network.

[0028] To quantify the coupling effect between different propulsion systems, this embodiment obtains the local operating state data of the propulsion system and the adjacent local operating state data of each adjacent propulsion system within a preset domain from the composite state data for any given propulsion system's corresponding agent. Specifically, after outputting the composite state data to the agent, the agent filters out the local operating state data of the current propulsion system and its adjacent propulsion systems based on predefined neighborhood relationships. Then, the agent uses a preset local feature encoder to process the local operating state data, adjacent local operating state data, and multi-degree-of-freedom thrust vectors, converting the original, heterogeneous operating data and global thrust vector into a compact vector representation. This captures the key features and intrinsic relationships of the data, facilitating subsequent processing by machine learning algorithms. Specifically, the local feature encoder is constructed based on a feedforward neural network, which receives the input raw numerical or categorical data and outputs a fixed-dimensional vector.

[0029] After obtaining the vector output by the local feature encoder, in order to evaluate the relevance or influence of each neighboring propulsion system on the current propulsion system and enable the propulsion system to adaptively focus on the most critical neighbors, this embodiment employs a dot product attention mechanism—a collaborative attention algorithm—to calculate the similarity score between the local state representation vector of the current propulsion system and the local state representation vectors of each neighboring propulsion system. Then, attention weights are obtained through a normalization function. Based on these attention weights, all neighboring local state representation vectors are weighted and fused to obtain the neighborhood state representation vector of the propulsion system. Finally, the neighborhood state representation vector and the local state representation vector are weighted and summed to obtain the enhanced state representation vector of the propulsion system. Specifically, the neighborhood state representation vector of each propulsion system is obtained by multiplying each neighboring local state representation vector with its corresponding attention weight and summing the results. The enhanced state representation vector of each propulsion system is then obtained by weighting and summing this neighborhood state representation vector with the local state representation vector. This integrates the weighted information from neighbors and combines it with the current system's own operating state to form a comprehensive and context-rich representation, improving control accuracy. Finally, the enhanced state representation vector is input into the parameterized action network of the executing agent to output the target operational state data of the propulsion system. This transforms the high-level, context-aware state representation into specific, executable control instructions, thereby directly managing the operation of the propulsion system. Specifically, the parameterized action network is typically a neural network (e.g., a policy network in reinforcement learning) that receives the enhanced state representation vector as input and outputs parameters that define the probability distribution of possible actions (e.g., the mean and standard deviation of continuous actions, or the log odds of discrete actions). The target operational state data can be sampled from this distribution or derived directly from its parameters.

[0030] The above steps, by acquiring local operational state data of the current propulsion system and its neighboring propulsion systems and semantically encoding it using multi-degree-of-freedom thrust vectors, transform complex raw data into a unified and processable vector representation. Based on this, the collaborative attention algorithm dynamically calculates the attention weights between the current propulsion system and each neighboring propulsion system, thus adaptively capturing the dynamic coupling relationships between propulsion systems as the environment changes, avoiding the limitations of traditional fixed weight allocation. By weighted fusion of neighboring local state representation vectors and weighted summation with its own local state representation vector, an enhanced state representation vector containing information about itself and its neighborhood is generated, providing a more comprehensive and accurate decision-making basis for the parameterized action network executing the agent. Ultimately, based on this enhanced state representation, the parameterized action network can output more accurate target operational state data, thereby achieving refined and adaptive dynamic control of the ship's composite propulsion system, significantly improving the ship's navigation stability and efficiency in complex sea conditions.

[0031] In this embodiment, step 1045 includes: inputting the enhanced state representation vector into the discrete action branch of the parameterized action network, and processing the enhanced state representation vector through the fully connected layer and multi-class activation function of the discrete action branch to obtain the discrete action probability distribution of the propulsion system; inputting the enhanced state representation vector into the continuous action branch of the parameterized action network, and processing the enhanced state representation vector through the fully connected layer and nonlinear activation function of the continuous action branch to obtain the continuous action initial parameter vector of the propulsion system; processing the discrete action probability distribution through a preset selection strategy to obtain the current operating mode of the propulsion system; performing linear mapping processing on the continuous action initial parameter vector to obtain the mapped initial power allocation value, initial target angle, and initial target rotation speed; calling the preset adaptation function corresponding to the current operating mode to perform mode adaptation processing on the mapped initial power allocation value, initial target angle, and initial target rotation speed to obtain the mode-adapted target power allocation value, target angle, and target rotation speed; and combining the target power allocation value, target angle, and target rotation speed to obtain the target operating state data of the propulsion system.

[0032] In this embodiment, the discrete action branch serves as a branch of the parameterized action network, handling discrete decisions and matching different operational modes of the ship's propulsion system, such as high-speed navigation, low-speed scientific research, and dynamic positioning. Specifically, the discrete action branch incorporates fully connected layers and multi-class activation functions to convert the input augmented state representation vector into probability distributions representing different operational modes. Specifically, a discrete action branch using multi-layered fully connected neural networks is employed, with each layer employing nonlinear activation functions such as ReLU, and the final layer using the Softmax activation function. The output is a probability vector equal to the preset number of discrete actions, where each element represents the probability of the corresponding discrete action. In this embodiment, the continuous action branch, also a branch of the parameterized action network, receives the augmented state representation vector and processes it through preset fully connected layers and nonlinear activation functions to obtain the initial parameter vector for the continuous actions of the propulsion system, such as continuous variables like power allocation, target angle, and target rotational speed. Specifically, the fully connected layers and nonlinear activation functions in the continuous action branch convert the input augmented state representation vector into an initial parameter vector for continuous actions, which contains a preliminary estimate of the continuous control variables of the propulsion system. Preferably, the continuous action branch consists of one or more fully connected layers, with ReLU activation functions used between each layer. The last layer either does not use an activation function or uses activation functions such as Tanh / Sigmoid to limit the output to a specific range, thereby generating the initial parameter vector for the continuous action. After the discrete action branch outputs the probability distribution of different operating modes, the discrete action probability distribution is processed by a preset selection strategy to obtain the current operating mode of the propulsion system. The discrete action probability distribution represents the probability of each operating mode being selected under the current enhancement state. The preset selection strategy is used to determine the final current operating mode based on these probability distributions, ensuring that the mode that best meets the current state and task requirements is selected from multiple possible modes. Specifically, the discrete action with the highest probability is selected as the current operating mode using a preset greedy strategy in the executing agent. After the continuous action branch outputs the initial parameter vector for the continuous action, since the initial parameter vector for the continuous action is an abstract vector output by the neural network, it needs to be converted into actual operable physical quantities, namely, power allocation value, target angle, and target rotational speed. This embodiment uses a linear mapping processing method to convert the abstract vector into control parameters with physical meaning through linear transformation. Specifically, the executing agent directly maps the initial parameter vector of continuous action to the initial power allocation value, the initial target angle, and the initial target rotation speed through a preset linear transformation matrix and bias vector.

[0033] Finally, different operating modes impose different requirements and constraints on the propulsion system's power, angle, and rotational speed. This embodiment utilizes a preset adaptation function to finely adjust the initial control parameters based on the current operating mode, ensuring that the control commands conform to the characteristics and optimization objectives of that mode. Specifically, the preset adaptation function is a pre-trained small neural network or lookup table. These networks or lookup tables output the target control parameters after mode adaptation based on the current operating mode and initial control parameters. Finally, all the continuous control parameters after mode adaptation are integrated into a complete set of control commands, i.e., the target operating state data. This data can be directly sent to the propulsion system's underlying controller to guide it in executing corresponding actions.

[0034] In the above steps, the discrete action branch is processed through a fully connected layer and a multi-class activation function to generate a discrete action probability distribution, thereby intelligently capturing the possibility of the optimal operating mode in the current state and avoiding the subjectivity and rigidity of mode selection. Simultaneously, the continuous action branch is processed through a fully connected layer and a nonlinear activation function to generate a continuous action initial parameter vector, providing dynamic and state-related initial parameters for subsequent fine-grained control. The current operating mode is determined from the discrete action probability distribution using a preset selection strategy, ensuring optimized mode decision-making. Subsequently, the continuous action initial parameter vector is linearly mapped to convert it into operable initial power allocation values, initial target angles, and initial target speeds. Crucially, this embodiment further introduces mode adaptation processing, that is, calling a preset adaptation function corresponding to the current operating mode to adjust these initial values, thereby ensuring that the final target power allocation values, target angles, and target speeds accurately meet the specific requirements and constraints of the current operating mode. This step-by-step processing and mode adaptation mechanism effectively solves the problem of insufficient joint optimization of discrete and continuous actions when parametric motion networks output target operating state data. It avoids control parameter mismatch due to mode switching or environmental changes, significantly improving the control accuracy, adaptability, and stability of the propulsion system in complex sea conditions. By combining the target power allocation value, target angle, and target speed, the final target operating state data can provide the propulsion system with complete and highly optimized control commands, thereby achieving dynamic and precise control of the ship's composite propulsion system.

[0035] In this first embodiment, the training process of the executive agent includes: constructing an initial executive agent model according to a preset deep deterministic policy gradient algorithm; the initial executive agent model includes a policy network and a value network; acquiring multiple historical navigation data of the ship, and for any historical navigation data, acquiring historical local operating state data, historical target operating state data, and several historical adjacent local operating state data corresponding to each propulsion system in the historical navigation data; for any propulsion system, forming a state space corresponding to the propulsion system based on all historical local operating state data and all historical adjacent local operating state data, and forming an action space corresponding to the propulsion system based on all historical target operating state data; inputting the state space into the policy network, outputting action commands through the policy network, and inputting the state space and action commands into the value network, outputting the first value estimate of the current state-action pair through the value network; inputting the next moment state of the state space into the target policy network, outputting the next moment action command through the target policy network, and inputting the next moment state and the next moment action command into the target policy network. The system inputs into the target value network and outputs a second value estimate of the state-action pair for the next time step. It calculates the reward value of the action command based on a preset composite reward function to obtain the immediate reward value corresponding to the action command. Then, it performs Bellman equation calculations based on the immediate reward value and the second value estimate to obtain the target value estimate of the current state-action pair. The composite reward function includes propulsion efficiency reward value, fuel consumption reward value, maneuverability reward value, and navigation stability reward value. A loss function for the value network is constructed based on the mean square error between the first value estimate and the target value estimate, and the network parameters of the value network are updated using a backpropagation algorithm. The gradient of the action command output by the policy network is calculated based on the first value estimate output by the value network to obtain the policy gradient of the policy network, and the network parameters of the policy network are updated using the policy gradient. The network parameters of the policy network are synchronized to the target policy network, and the network parameters of the value network are synchronized to the target value network. When the loss function of the value network converges to a preset threshold, the executing agent corresponding to the propulsion system is obtained based on the converged network parameters and the initial executing agent model.

[0036] This embodiment constructs an initial execution agent model based on a pre-defined deep deterministic policy gradient algorithm. The deep deterministic policy gradient algorithm is a reinforcement learning algorithm based on an Actor (policy network)-Critic (value network) architecture. In the initial execution agent model, the policy network is responsible for outputting deterministic action instructions based on the current state, while the value network is responsible for evaluating the long-term value of these actions. Specifically, the deep deterministic policy gradient algorithm can be implemented based on a deep learning framework, where both the policy network and the value network are composed of multi-layer fully connected neural networks or convolutional neural networks, and an optimizer is used for parameter updates.

[0037] In the specific training process, multiple historical navigation data points of the ship are acquired. These historical navigation data are various sensor data and control commands recorded by the ship during actual operation, forming the basis for the intelligent agent to learn complex navigation patterns in the real world. For any given historical navigation data point, it is necessary to acquire the historical local operating state data, historical target operating state data, and several historical adjacent local operating state data corresponding to each propulsion system. The historical local operating state data may include the propeller speed, pitch, thrust, etc.; the historical target operating state data are the expected operating parameters given by the control system at that time; and the historical adjacent local operating state data reflect the operating status of other propellers, demonstrating the coupling between systems. This data can be extracted from the ship's historical navigation log through the ship's navigation data recorder, integrated automation system, or dedicated data acquisition system, and subjected to necessary preprocessing. Alternatively, it can be generated by running a ship model in a virtual simulation environment to provide a large amount of simulated navigation data, providing diverse training samples. Based on this, for any propulsion system, the state space corresponding to the propulsion system is formed according to all historical local operating state data and all historical adjacent local operating state data, and the action space corresponding to the propulsion system is formed according to all historical target operating state data. The state space is the complete set of information about the environment perceived by the agent, while the action space is the set of all possible actions the agent can perform. For example, the state space can be constructed by concatenating the historical local operating state data of the propulsion system (such as current speed, pitch, thrust, energy consumption, etc.) with the historical local operating state data of adjacent propulsion systems (such as the speed, pitch, thrust, etc. of adjacent thrusters); the action space is constructed by the historical target operating state data of the propulsion system (such as target speed, target angle, target power allocation, etc.). The state space is then input into the policy network, which outputs action commands. The policy network is a multilayer perceptron, whose input layer receives the state space vector, and whose output layer outputs continuous action commands through an activation function. Simultaneously, the state space and action commands are input into the value network, which outputs the first value estimate of the current state-action pair. The value network is also a multilayer perceptron, whose input layer receives the state space vector and the action command vector, and whose output layer outputs a scalar representing the value estimate. To stabilize the training process, the next state of the state space is input into the target policy network, which outputs the next action command. The target policy network is a copy of the main policy network, but its parameters are updated less frequently. The next-time state and the next-time action instruction are input into the target value network, which outputs a second value estimate of the next-time state-action pair. The target value network is also a copy of the main value network, used to provide a relatively stable learning objective. Furthermore, the reward value for the action instruction is calculated according to a preset composite reward function to obtain the immediate reward value corresponding to the action instruction.The composite reward function aims to comprehensively evaluate the merits of an agent's actions. It can include rewards for propulsion efficiency, fuel consumption, maneuverability, and navigation stability; for example, it can be designed as a weighted sum of these rewards. Then, the Bellman equation is used to calculate the target value estimate for the current state-action pair based on the immediate reward value and the second value estimate. The Bellman equation combines the present value of immediate rewards with the discounted sum of future rewards to evaluate long-term performance.

[0038] Next, a loss function for the value network is constructed based on the mean squared error (MSE) between the first value estimate and the target value estimate, and the network parameters of the value network are updated using the backpropagation algorithm. The MSE quantifies the prediction error of the value network, and the backpropagation algorithm uses this error signal to adjust the weights and biases of the value network, making its predictions more accurate. Simultaneously, gradient calculations are performed on the action instructions output by the policy network based on the first value estimate output by the value network, obtaining the policy gradient of the policy network, and the network parameters of the policy network are updated using the policy gradient. The goal of the policy network is to maximize the value assessment of the policy network's output actions by the value network; the policy gradient guides the policy network to adjust its parameters to output actions that yield higher values. To maintain training stability, the network parameters of the policy network are synchronized to the target policy network using a soft update method, and the network parameters of the value network are synchronized to the target value network. This soft update mechanism prevents oscillations caused by bootstrapping problems during training, ensuring the stability of the learning objective. Finally, when the loss function of the value network converges to a preset threshold, the executing agent corresponding to the propulsion system is obtained based on the converged network parameters and the initial executing agent model. The convergence of the loss function indicates that the value network can accurately evaluate the value of state-action pairs, and the policy network has learned a better control policy.

[0039] The training steps described above construct an initial agent model comprising a policy network and a value network. Utilizing historical navigation data, a state space and action space are defined for each propulsion system, enabling the agent to learn complex control strategies from real-world data. The policy network generates action commands, while the value network evaluates the value of those actions; both work together for optimization. Specifically, by introducing a target network to stabilize the learning objective and designing a composite reward function incorporating multi-dimensional indicators such as propulsion efficiency, fuel consumption, maneuverability, and navigation stability, the agent is guided to weigh multiple objectives and learn the optimal control strategy that balances performance and efficiency. The value network and policy network are iteratively updated using a mean squared error loss function and policy gradients, and a soft synchronization mechanism stabilizes the training process, ultimately resulting in a converged agent. This pre-training mechanism allows the agent, upon receiving multi-degree-of-freedom thrust vectors and composite state data from the master agent, to more accurately and intelligently generate target operational state data for each propulsion system based on its learned strategy. This not only improves the control precision of individual propulsion systems, but more importantly, because the training process considers the local operating state data of adjacent propulsion systems, the executing agent can perceive and adapt to the dynamic coupling relationship between propulsion systems, thereby achieving more refined and coordinated collaborative control of propulsion systems in complex sea conditions. In this way, the overall control precision of the ship and its adaptability to environmental changes are significantly improved, effectively reducing thrust loss and risks caused by improper control, and thus improving the ship's navigation stability, fuel economy, and maneuverability.

[0040] In this embodiment, step 105 includes: for any propulsion system, when controlling the propulsion system according to the target operating status data, acquiring the multi-source operating fluctuation data sequence of the propulsion system; extracting time-domain statistical features and frequency-domain features from the multi-source operating fluctuation data sequence to obtain the time-domain feature vector and frequency-domain feature vector of the propulsion system; inputting the time-domain feature vector and frequency-domain feature vector into a pre-constructed risk pattern library for multi-level matching analysis to obtain the risk pattern and risk level of the propulsion system; the risk pattern library pre-stores standard feature templates corresponding to various risk patterns and their corresponding risk levels.

[0041] In this embodiment, when controlling the propulsion system based on target operating status data, its operating status is continuously monitored. For any propulsion system, during the control process, a multi-source operating fluctuation data sequence (e.g., vibration sensor data, motor current fluctuation data, bearing temperature data) is continuously acquired. Time-domain statistical feature extraction and frequency-domain feature extraction are performed on the multi-source operating fluctuation data sequence to obtain the time-domain feature vector and frequency-domain feature vector of the propulsion system. These feature vectors are input into a pre-built risk pattern library for multi-level matching analysis. The risk pattern library pre-stores standard feature templates and corresponding risk levels for various risk patterns (such as cavitation, bearing wear, and motor overload). Through matching analysis, the risk patterns and risk levels of the propulsion system can be identified.

[0042] The above steps comprehensively capture multi-source operational fluctuation data of the propulsion system and utilize time-domain and frequency-domain feature extraction techniques to transform the raw data into highly discriminative feature vectors, thus overcoming the limitations of single data sources and single-dimensional analysis. Based on this, these feature vectors are input into a pre-constructed risk pattern library for multi-level matching analysis, enabling rapid and accurate identification of specific risk patterns and their severity within the propulsion system. This refined risk analysis mechanism allows the system to promptly and accurately diagnose potential problems and provide clear risk patterns and risk levels during the operation of the ship's composite propulsion system, providing a precise and reliable basis for subsequent dynamic adjustment of control strategies. This not only significantly improves the accuracy and timeliness of risk identification, avoiding control deviations caused by inaccurate or untimely risk identification, but also ensures stable operation of the ship in complex sea conditions, effectively solving the problem of inaccurate risk identification in existing technologies and guaranteeing the safety and reliability of ship navigation.

[0043] In this embodiment, step 106 includes: querying the basic adjustment parameters of each propulsion system from the pre-configured risk adjustment mapping table according to the risk mode; quantifying the basic adjustment parameters according to the risk level to obtain the initial data adjustment amount of each propulsion system and limiting the initial data adjustment amount according to the ship operation constraint data of each propulsion system to obtain the data adjustment amount of each propulsion system; and realizing the dynamic control of each propulsion system according to the data adjustment amount and the target operating status data.

[0044] In this embodiment, risk mode refers to an abstract concept that classifies and describes abnormal states or potential hazards that may occur during the operation of a ship's composite propulsion system. For example, risk modes can be classified into minor vibrations, power fluctuations, overload warnings, and accelerated component wear. Furthermore, risk modes can also be obtained through unsupervised or supervised learning of large amounts of operational fluctuation data using machine learning models (such as clustering algorithms or classifiers), thereby automatically identifying different risk categories. The pre-configured risk adjustment mapping table is a database or lookup table storing the correspondence between different risk modes and their corresponding basic adjustment parameters. The basic adjustment parameters are initial control adjustment suggestions or ranges preset for specific risk modes without fine-grained quantification. Their function is to provide a preliminary directional adjustment benchmark for different risk modes, which may include power adjustment coefficients, speed adjustment step sizes, thrust vector angle adjustment ranges, etc. The risk level is an indicator that quantifies the severity or urgency of the identified risk modes, usually expressed in numerical or hierarchical form. Its function is to refine the basic adjustment parameters, ensuring that the adjustment amount matches the actual severity of the risk and avoiding excessive or insufficient adjustments. Quantitative calculations can employ methods such as linear interpolation, nonlinear function mapping, or piecewise functions to map the risk level to the correction coefficients of the basic adjustment parameters. The initial data adjustment amount is a preliminary control adjustment value obtained after risk level quantification calculations, but before undergoing safety constraint checks. Its function is to serve as an intermediate result before amplitude limiting, reflecting the ideal adjustment range suggested by the system based on the risk model and level. Ship operation constraint data refers to data specified in ship design, construction, and operation procedures to limit the safe range and performance boundaries of propulsion system operating parameters. This can include the maximum / minimum power, maximum / minimum speed, maximum / minimum thrust vector angle, maximum acceleration, and maximum angular velocity of the propeller. Amplitude limiting refers to comparing the initial data adjustment amount with the ship operation constraint data and truncating or correcting the adjustment amount according to the constraint conditions to keep it within the safe allowable range. For example, if the initial adjustment amount causes a parameter to exceed the maximum allowable value, then it is set to the maximum allowable value. Alternatively, the limiting process can employ a more complex optimization algorithm; the data adjustment amount is the actual control adjustment value that is finally determined after the limiting process and used to correct the target operating state data.

[0045] Specifically, this embodiment determines the data adjustment amount for each propulsion system based on risk analysis results to achieve dynamic control. Based on the identified risk pattern (e.g., a propeller has a moderate cavitation risk), the basic adjustment parameters for that propulsion system are retrieved from a pre-configured risk adjustment mapping table. Based on the risk level (e.g., moderate risk), the basic adjustment parameters are quantified to obtain the initial data adjustment amount for that propulsion system. Simultaneously, based on the ship's operational constraints for that propulsion system (e.g., to ensure ship maneuverability, thrust adjustment cannot fall below a certain threshold), the initial data adjustment amount is limited to obtain the final data adjustment amount. For example, if cavitation risk is detected, the data adjustment amount might indicate a slight reduction in blade speed or an adjustment in blade angle. Based on the data adjustment amount and target operating state data, the system achieves dynamic control for each propulsion system. This closed-loop, risk-based dynamic adjustment mechanism enables the ship's composite propulsion system to achieve adaptive control in complex sea conditions, significantly improving the stability and safety of ship operation.

[0046] After identifying operational fluctuations in the propulsion system and conducting risk analysis as described above, the dynamic adjustment control module can quickly generate precise data adjustment amounts based on the risk pattern and level, combined with ship operational constraints. These adjustments directly affect the target operational status data generated by the execution module, enabling the propulsion system to respond promptly to risks and correct its operating parameters. This prevents further escalation of risks and ensures stable navigation of the ship even in complex sea conditions, significantly improving the adaptability and safety of the ship's composite propulsion system.

[0047] In this first embodiment, after obtaining the data adjustment amount for each propulsion system, the execution agent of each propulsion system is updated based on the data adjustment amount and the target operating state data. Specifically, the update process of the execution agent is as follows: First, during ship operation, multi-source operating data of each propulsion system is collected in real time, including target operating state data before dynamic control, operating fluctuation data generated during dynamic control, risk analysis results, actual data adjustment amounts, and propulsion system response data after dynamic control. The above data is associated and stored in an online experience data cache pool to form a complete experience sample including state, action, risk, adjustment, and effect. Second, the samples in the online experience data cache pool are screened, prioritizing the retention of samples with high learning value, including: samples with risk levels exceeding a preset threshold, samples with large deviations between risk analysis results and expectations, and samples where data adjustment amounts significantly change the operating state of the propulsion system. Through the above screening mechanism, it is ensured that the execution agent can learn efficiently from key risk events and avoid being overwhelmed by a large amount of redundant data from normal operation. Next, training data for online updating of the execution agent is extracted from the screened high-risk-value experience samples. For each experience sample, the local state observation before the risk occurs is taken as the current state, the risk analysis result and corresponding data adjustment amount are taken as the corrected action command, and the propulsion system response data after dynamic control is taken as the next moment's state. The immediate reward value corresponding to the sample is calculated according to the pre-set composite reward function in the executing agent, forming a complete four-tuple training sample of state, action, reward, and next state. The training samples from the online update sample set are input into the policy network and value network of the executing agent to fine-tune the network parameters. During the update process, a small learning rate and a small number of iterations are used to avoid excessive disturbance to the optimal policy already learned by the executing agent. Simultaneously, a copy of the network parameters before the update is retained for comparison and verification of the policy effect before and after the update. Finally, a safety boundary is set during the online update process of the executing agent. The difference in action commands output by the policy network before and after the update is calculated. If the difference exceeds a preset safety threshold, the update is paused and the network parameters before the update are restored. Simultaneously, the operating data generated by the executed agent in subsequent control processes after the update is compared in real time with historical baseline data. If an abnormal decrease in the stability of the propulsion system is detected, it is automatically rolled back to the stable version before the update. The optimization experience samples accumulated during online updates are periodically integrated into the offline training dataset, and the executive agent is periodically retrained on a large scale in a virtual simulation environment. By combining online fine-tuning with offline retraining, the executive agent can both respond to operational risks in real time to optimize policies and achieve systematic evolution from long-term accumulated experience.

[0048] Through the aforementioned online update process, the executing agent can continuously adapt to evolving factors such as equipment aging, seasonal changes in sea conditions, and the characteristics of the navigation area during the long-term operation of the ship, dynamically optimizing its collaborative control strategy. When the system adjusts the data based on the risk analysis results, the executing agent can learn from the effects of this adjustment, enabling the ship's composite propulsion system to possess continuous evolutionary adaptive capabilities, significantly improving the operational stability and control accuracy of the propulsion system throughout the ship's entire life cycle.

[0049] On the other hand, refer to Figure 2 This embodiment also discloses a ship composite propulsion dynamic control system based on reinforcement learning, including an environment perception module 201, a composite state module 202, a main control module 203, an execution module 204, a monitoring module 205, and a dynamic adjustment control module 206.

[0050] The environmental perception module 201 is used to acquire sea surface images of the area where the ship is located, and identify wave height data and current velocity data in the sea surface images to obtain environmental perception data of the area where the ship is located. The composite state module 202 is used to acquire the ship's navigation attitude data, navigation mission data, and operating status data of each propulsion system in the ship, so as to obtain the composite state space data sequence of the ship in combination with the environmental perception data. The main control module 203 is used to input the composite state space data sequence into the main control agent obtained by pre-training the reinforcement learning algorithm, so as to output the multi-degree-of-freedom thrust vector of the ship through the main control agent; The execution module 204 is used to input the multi-degree-of-freedom thrust vector and the composite state data into the execution agent that is pre-trained by the reinforcement learning algorithm corresponding to each of the propulsion systems, so that the execution agent generates target running state data corresponding to each of the propulsion systems according to the preset parameterized action space. The monitoring module 205 is used to acquire the operational fluctuation data of each propulsion system when controlling the propulsion system based on the target operational status data, and to perform risk analysis on the operational fluctuation data based on a pre-built risk pattern library to obtain the risk analysis result of each propulsion system. The dynamic adjustment control module 206 is used to determine the data adjustment amount of each of the propulsion systems based on the risk analysis results, so as to realize the dynamic control of each of the propulsion systems based on the data adjustment amount and the target operating status data.

[0051] This embodiment discloses a reinforcement learning-based dynamic control method and system for ship composite propulsion. It obtains environmental perception data representing sea state characteristics from sea surface image recognition of the ship's location, and combines this with navigation attitude data, navigation mission data, and operational status data of each propulsion system to obtain a composite state space data sequence representing the ship's multi-dimensional operational state. This allows for subsequent global thrust decoupling decision-making through a master control agent based on the composite state space data sequence to obtain a multi-degree-of-freedom thrust vector representing the overall control objective of the ship. Based on the multi-degree-of-freedom thrust vector and composite state data, local collaborative allocation is performed through the execution agent corresponding to each propulsion system to obtain target operational state data for each propulsion system. During dynamic control, risk pattern matching analysis is performed based on the operational fluctuation data of each propulsion system to obtain risk analysis results representing abnormal operational states. The data adjustment amount used for dynamic control is adjusted based on the risk analysis results and ship operational constraint data. Finally, based on the adaptive and precise adjusted data adjustment amount and target operational state data, dynamic adaptive control of the ship's composite propulsion system under complex sea conditions is achieved, improving the stability of ship operation.

[0052] The specific embodiments described above further illustrate the purpose, technical solution, and beneficial effects of the present invention. It should be understood that the above descriptions are merely specific embodiments of the present invention and are not intended to limit the scope of protection of the present invention. In particular, it should be noted that any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention for those skilled in the art.

Claims

1. A dynamic control method for ship composite propulsion based on reinforcement learning, characterized in that, include: The system acquires sea surface images of the area where the ship is located, and identifies wave height and current velocity data in the sea surface images to obtain environmental perception data of the area where the ship is located. The ship's navigation attitude data, navigation mission data, and operational status data of each propulsion system in the ship are acquired, and combined with the environmental perception data to obtain a composite state space data sequence of the ship. The composite state space data sequence is input into the master control agent pre-trained by the reinforcement learning algorithm, so that the master control agent can output the multi-degree-of-freedom thrust vector of the ship. The multi-degree-of-freedom thrust vector and the composite state space data sequence are input into the execution agent that is pre-trained by the reinforcement learning algorithm corresponding to each of the propulsion systems, so that the execution agent generates the target running state data corresponding to each of the propulsion systems according to the preset parameterized action space. When controlling the propulsion system based on the target operating status data, the operating fluctuation data of each propulsion system is acquired, and the operating fluctuation data is analyzed for risk based on a pre-built risk model library to obtain the risk analysis results of each propulsion system. Based on the risk analysis results, the data adjustment amount for each of the propulsion systems is determined, so as to realize the dynamic control of each of the propulsion systems according to the data adjustment amount and the target operating status data.

2. The method for dynamic control of ship composite propulsion based on reinforcement learning according to claim 1, characterized in that, The process of acquiring sea surface images of the area where the ship is located, and identifying wave height and current velocity data in the sea surface images to obtain environmental perception data of the area where the ship is located includes: Real-time acquisition of sea surface image sequences in the area where the ship is located, and grayscale processing and motion blur correction processing of the sea surface image sequences to obtain preprocessed sea surface image sequences; The pre-trained sea state visual analysis model is used to identify the pre-processed sea surface image sequence frame by frame to obtain wave surface texture feature sequence, wave crest line distribution feature sequence and surface flow field stripe feature sequence. The spatial frequency of the wave surface texture feature sequence and the dominant wave period of the wave crest distribution feature sequence are obtained, so as to invert the wave spectrum of the area where the ship is located based on the spatial frequency; The wave height data of the area where the ship is located is obtained based on the dominant wave period, the wave spectrum, and the pre-calibrated pixel texture scale mapping relationship; The morphological offset between any two adjacent surface flow field stripe features in the surface flow field stripe feature sequence is obtained, and the pixel displacement between any two adjacent surface flow field stripe features is obtained according to the morphological offset and the preset optical flow analysis method. The flow velocity data of the area where the ship is located is obtained based on the pixel texture scale mapping relationship, the sampling frequency of the preprocessed sea surface image sequence, and all the pixel displacements corresponding to the surface flow field stripe feature sequence. By integrating the wave height data and the current velocity data, environmental perception data of the area where the ship is located is obtained.

3. The reinforcement learning-based dynamic control method for ship composite propulsion according to claim 2, characterized in that, The acquisition of the ship's navigation attitude data, navigation mission data, and operational status data of each propulsion system in the ship, combined with the environmental perception data, to obtain a composite state-space data sequence of the ship, includes: The multi-sensor time synchronization protocol of the ship is obtained, and a clock source for a unified time reference is determined according to the multi-sensor time synchronization protocol; The ship's navigation attitude data and the operating status data of each propulsion system in the ship are acquired, so as to determine the acquisition time corresponding to each item of the environmental perception data, the navigation attitude data and the operating status data according to the clock source; Based on the acquisition time, the environmental perception data, the navigation attitude data, and the operational status data are subjected to time interpolation processing to obtain the environmental perception data, navigation attitude data, and operational status data at the same time as the composite status data of the ship at each time. The navigation mission data of the vessel is acquired, and the navigation mission data is semantically encoded to map the navigation mission data into a mission semantic embedding vector of a preset length. Obtain the ship operation constraint data corresponding to the navigation mission data from the pre-configured mission target mapping table; For any given moment, the composite state data at that moment is normalized to convert the composite state data into a composite state data vector of a preset dimension. The task semantic embedding vector, the ship operation constraint data, and the composite state data vector are concatenated to obtain the composite state space data of the ship at the specified moment. The composite state space data sequence of the ship is obtained based on the composite state space data corresponding to each of the multiple consecutive moments within a preset time window.

4. The reinforcement learning-based dynamic control method for ship composite propulsion according to claim 3, characterized in that, The step of inputting the composite state-space data sequence into a master control agent pre-trained by a reinforcement learning algorithm, so as to output the multi-degree-of-freedom thrust vector of the ship through the master control agent, includes: The composite state space data sequence is modeled temporally by a time encoder preset in the main control agent, and the motion state evolution characteristics of the ship are extracted to obtain a temporal hidden state vector. Based on the composite state space data sequence, query the task attention association matrix preset in the master control agent to obtain the attention weight vector corresponding to the temporal hidden state vector; The main control agent performs element-wise multiplication of the attention weight vector and the temporal hidden state vector to obtain the comprehensive state representation vector of the ship. The longitudinal thrust feature vector, lateral thrust feature vector, and yaw moment feature vector corresponding to the comprehensive state representation vector are output through the multi-feature extraction network preset in the main control agent. The main control agent splices the longitudinal thrust feature vector, the lateral thrust feature vector, and the yaw moment feature vector to obtain a joint feature vector, and obtains the longitudinal and lateral coupling coefficients, the longitudinal yaw coupling coefficient, and the lateral yaw coupling coefficient corresponding to the joint feature vector. The longitudinal thrust reference value corresponding to the longitudinal thrust feature vector, the lateral thrust reference value corresponding to the lateral thrust feature vector, and the yaw torque reference value corresponding to the bow torque feature vector are obtained by a linear activation function preset in the main control agent. The lateral thrust coupling compensation amount is obtained based on the longitudinal and lateral coupling coefficients and the longitudinal thrust reference value, and the first horn moment coupling compensation amount is obtained based on the longitudinal yaw coupling coefficient and the longitudinal thrust reference value, and the second horn moment coupling compensation amount is calculated based on the lateral yaw coupling coefficient and the lateral thrust reference value. Based on the longitudinal thrust reference value, the lateral thrust reference value, the yaw moment reference value, the lateral thrust coupling compensation amount, the first yaw moment coupling compensation amount, and the second yaw moment coupling compensation amount, the main control agent outputs a multi-degree-of-freedom thrust vector including the longitudinal thrust component, the lateral thrust component, and the yaw moment component.

5. The reinforcement learning-based dynamic control method for ship composite propulsion according to claim 4, characterized in that, The step of inputting the multi-degree-of-freedom thrust vector and the composite state space data sequence into the execution agent pre-trained by the reinforcement learning algorithm corresponding to each of the propulsion systems, so that the execution agent generates target operating state data corresponding to each of the propulsion systems according to the preset parameterized action space, includes: For any execution agent corresponding to a propulsion system, the local operating state data of the propulsion system and the adjacent local operating state data of each adjacent propulsion system within a preset domain are obtained from the composite state data. By semantically encoding the local operating state data, the adjacent local operating state data and the multi-degree-of-freedom thrust vector by a local feature encoder preset in the executing agent, the local state representation vector of the propulsion system and the adjacent local state representation vector of each adjacent propulsion system are obtained. Based on the local state representation vector and all the adjacent local state representation vectors, the attention weight between the propulsion system and each of the adjacent propulsion systems is obtained by a pre-set collaborative attention algorithm in the executing agent. The neighboring local state representation vectors are weighted and fused according to the attention weights to obtain the neighborhood state representation vector of the propulsion system, and the neighborhood state representation vector and the local state representation vector are weighted and summed to obtain the enhanced state representation vector of the propulsion system. The enhanced state representation vector is input into the parameterized action network of the executing agent to output the target operating state data of the propulsion system through the parameterized action network.

6. The method for dynamic control of ship composite propulsion based on reinforcement learning according to claim 5, characterized in that, The step of inputting the enhanced state representation vector into the parameterized action network of the executing agent, so as to output the target operating state data of the propulsion system through the parameterized action network, includes: The enhanced state representation vector is input into the discrete action branch of the parameterized action network, and the enhanced state representation vector is processed by the fully connected layer and multi-class activation function of the discrete action branch to obtain the discrete action probability distribution of the propulsion system. The enhanced state representation vector is input into the continuous action branch of the parameterized action network, and the enhanced state representation vector is processed by the fully connected layer and nonlinear activation function of the continuous action branch to obtain the continuous action initial parameter vector of the propulsion system. The current operating mode of the propulsion system is obtained by processing the discrete action probability distribution through a preset selection strategy. The initial parameter vector of the continuous action is linearly mapped to obtain the mapped initial power allocation value, initial target angle and initial target rotation speed. The preset adaptation function corresponding to the current operating mode is called to perform mode adaptation processing on the mapped initial power allocation value, initial target angle and initial target speed, so as to obtain the mode-adapted target power allocation value, target angle and target speed; The target power allocation value, the target angle, and the target rotational speed are combined to obtain the target operating status data of the propulsion system.

7. The reinforcement learning-based dynamic control method for ship composite propulsion according to claim 6, characterized in that, Before inputting the multi-degree-of-freedom thrust vector and the composite state data into the execution agent pre-trained by the reinforcement learning algorithm corresponding to each of the propulsion systems, the process includes: An initial execution agent model is constructed based on a preset deep deterministic policy gradient algorithm; the initial execution agent model includes a policy network and a value network; Acquire multiple historical navigation data of the ship. For any one of the historical navigation data, acquire the historical local operating status data, historical target operating status data, and several historical adjacent local operating status data corresponding to each propulsion system in the historical navigation data. For any propulsion system, a state space corresponding to the propulsion system is formed based on all the historical local operating state data and all the historical adjacent local operating state data corresponding to the propulsion system, and an action space corresponding to the propulsion system is formed based on all the historical target operating state data. The state space is input into the policy network, and action instructions are output through the policy network. The state space and the action instructions are input into the value network, and the first value estimate of the current state-action pair is output through the value network. The next time-space state is input into the target policy network, and the next time-space action instruction is output through the target policy network. The next time-space state and the next time-space action instruction are input into the target value network, and the second value estimate of the next time-space state-action pair is output through the target value network. The reward value of the action command is calculated according to a preset composite reward function to obtain the instant reward value corresponding to the action command. Then, the Bellman equation is used to calculate the target value estimate of the current state action pair based on the instant reward value and the second value estimate. The composite reward function includes propulsion efficiency reward value, fuel consumption reward value, maneuverability reward value and navigation stability reward value. The loss function of the value network is constructed based on the mean square error between the first value estimate and the target value estimate, and the network parameters of the value network are updated through the backpropagation algorithm. The gradient of the action instruction output by the policy network is calculated based on the first value estimate output by the value network to obtain the policy gradient of the policy network, and the network parameters of the policy network are updated through the policy gradient. The network parameters of the policy network are synchronized to the target policy network, and the network parameters of the value network are synchronized to the target value network. When the loss function of the value network converges to a preset threshold, the execution agent corresponding to the propulsion system is obtained based on the converged network parameters and the initial execution agent model.

8. The reinforcement learning-based dynamic control method for ship composite propulsion according to claim 7, characterized in that, When controlling the propulsion system based on the target operating state data, the operational fluctuation data of each propulsion system is acquired, and risk analysis is performed on the operational fluctuation data according to a pre-built risk model library to obtain the risk analysis results of each propulsion system, including: For any of the propulsion systems, when controlling the propulsion system according to the target operating state data, a multi-source operating fluctuation data sequence of the propulsion system is obtained; Time-domain statistical feature extraction and frequency-domain feature extraction are performed on the multi-source operational fluctuation data sequence to obtain the time-domain feature vector and frequency-domain feature vector of the propulsion system. The time-domain feature vector and the frequency-domain feature vector are input into a pre-constructed risk pattern library for multi-level matching analysis to obtain the risk pattern and risk level of the propulsion system; the risk pattern library contains pre-stored standard feature templates and corresponding risk levels for various risk patterns.

9. The reinforcement learning-based dynamic control method for ship composite propulsion according to claim 8, characterized in that, The step of determining the data adjustment amount for each propulsion system based on the risk analysis results, and then implementing dynamic control of each propulsion system based on the data adjustment amount and the target operating status data, includes: Based on the risk model, the basic adjustment parameters of each of the propulsion systems are queried from the pre-configured risk adjustment mapping table; The basic adjustment parameters are quantified and calculated according to the risk level to obtain the initial data adjustment amount for each propulsion system. The initial data adjustment amount is then limited according to the ship operation constraint data of each propulsion system to obtain the data adjustment amount for each propulsion system. Dynamic control of each propulsion system is achieved based on the data adjustment amount and the target operating status data.

10. A dynamic control system for ship composite propulsion based on reinforcement learning, characterized in that, It includes an environmental perception module, a composite status module, a main control module, an execution module, a monitoring module, and a dynamic adjustment control module; The environmental perception module is used to acquire sea surface images of the area where the ship is located, and identify wave height data and current velocity data in the sea surface images to obtain environmental perception data of the area where the ship is located. The composite state module is used to acquire the ship's navigation attitude data, navigation mission data, and operating status data of each propulsion system in the ship, so as to obtain the composite state space data sequence of the ship in combination with the environmental perception data. The main control module is used to input the composite state space data sequence into the main control agent obtained by pre-training the reinforcement learning algorithm, so as to output the multi-degree-of-freedom thrust vector of the ship through the main control agent; The execution module is used to input the multi-degree-of-freedom thrust vector and the composite state space data sequence into the execution agent that is pre-trained by the reinforcement learning algorithm corresponding to each of the propulsion systems, so that the execution agent generates target running state data corresponding to each of the propulsion systems according to the preset parameterized action space. The monitoring module is used to acquire operational fluctuation data of each propulsion system when controlling the propulsion system based on the target operational status data, and to perform risk analysis on the operational fluctuation data based on a pre-built risk model library to obtain the risk analysis result of each propulsion system. The dynamic adjustment control module is used to determine the data adjustment amount of each of the propulsion systems based on the risk analysis results, so as to realize the dynamic control of each of the propulsion systems based on the data adjustment amount and the target operating status data.