A reinforcement learning control method for ship cooling systems
By using reinforcement learning control methods to optimize the operation of spray valves and condensate circulation pumps in real time, the problem of traditional ship cooling systems being unable to adapt to changes in heat transfer performance caused by marine organism attachment was solved. This achieved precise control and adaptive adjustment, improving system efficiency and reliability.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- CHINA STATE SHIPBUILDING CORP LTD RESEARCH INSTITUTE 719
- Filing Date
- 2025-06-16
- Publication Date
- 2026-07-17
AI Technical Summary
Traditional ship cooling system control strategies cannot be updated in real time, making it difficult to adapt to changes in heat transfer performance caused by factors such as marine organism attachment, resulting in reduced heat transfer efficiency and affecting the performance and energy consumption of the ship's power system.
By employing reinforcement learning control methods, and through the construction of reward functions and network training, the opening degree of spray valves and the speed of condensate circulation pumps are optimized in real time to achieve precise control of vacuum/subcooling.
It improves the control precision and adaptability of the cooling system, reduces energy consumption and operating costs, and ensures the stable operation of the ship's power system.
Smart Images

Figure CN120779711B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of marine power system control technology, and more specifically, to a reinforcement learning control method for a marine cooling system. Background Technology
[0002] The efficient and stable operation of a ship's propulsion system is crucial for its navigation safety and performance, and the ship's cooling system, as a core component, plays a key role in maintaining the normal operating condition of the propulsion equipment. In the ship's cooling system, heat exchange is achieved through key components such as seawater heat exchangers to ensure that the cooling requirements of the ship's propulsion unit are met. With the continuous development of ship technology, increasingly higher demands are being placed on the control precision and adaptability of cooling systems. Especially in the complex marine environment, the surface of seawater heat exchangers is easily affected by issues such as the attachment of marine organisms, leading to gradual changes in their heat transfer performance, which presents new challenges to the control strategies of the cooling system.
[0003] Traditional ship cooling systems typically employ strategies based on fixed control parameters, such as pre-set vacuum / subcooling control parameters. While this method can maintain stable system operation to some extent, it has significant limitations. Firstly, once the parameters of traditional control strategies are determined, they cannot be updated in real-time according to actual heat transfer changes in the seawater heat exchanger during system operation, making it difficult to adapt to the gradual deterioration of heat transfer performance caused by factors such as marine organism attachment. Secondly, when a fouling layer forms on the surface of the seawater heat exchanger, the system's heat transfer efficiency decreases, and traditional control methods struggle to ensure that the vacuum and subcooling of the cooling system remain within precise control ranges. This negatively impacts the overall performance and operating efficiency of the ship's propulsion system, increasing energy consumption and equipment maintenance costs. Summary of the Invention
[0004] In view of at least one defect or improvement requirement of the prior art, this application provides a reinforcement learning control method for a ship cooling system, which can solve at least one of the problems existing in the background art.
[0005] To achieve the above objectives, according to the first aspect of this application, a reinforcement learning control method for a ship cooling system is provided, the method comprising the following steps:
[0006] Determine the input measured values and output action values of the ship's cooling system. The input measured values are state parameters that are strongly correlated with the vacuum / subcooling control target. The output action values include the spray valve opening degree and the condensate circulation pump speed.
[0007] A reinforcement learning reward function is constructed, which comprehensively considers both penalty and incentive factors, and reflects the degree of influence of the seawater heat exchanger evolution process on the vacuum degree / subcooling degree through the reward function;
[0008] Offline training of network parameters is performed, including offline training of policy network and offline training of value network. The policy network is pre-trained using data under traditional vacuum / supercooling control. The reward value is calculated based on the reward function and the measured values and action values of the policy network. The value network is pre-trained based on the measured values, action values and reward values.
[0009] The offline-trained policy network and value network are deployed to the cooling system control device and enter the online training phase. Through real-time interaction between the policy network, value network and cooling system, the parameters of the policy network and value network are optimized online using a reinforcement learning algorithm based on deep deterministic policy gradients until the vacuum / supercooling control effect meets the application requirements.
[0010] Furthermore, the aforementioned reinforcement learning control method for ship cooling systems, which comprehensively considers both penalty and incentive factors, specifically includes:
[0011] The main penalty factor is given based on the deviation from the vacuum / subcooling target; the larger the deviation, the larger the penalty value.
[0012] Secondary penalty factors are given based on the changes in the spray valve opening and the condensate circulation pump speed. The greater the change in the speed, the greater the penalty.
[0013] The main incentive factors are given based on the degree of proximity to the vacuum / supercooling target. When the actual output is close to the target value, a positive reward is given.
[0014] Absolute penalties are based on training time, and penalties are imposed when training is terminated prematurely.
[0015] Furthermore, the above-mentioned reinforcement learning control method for ship cooling systems is characterized in that the reward function is:
[0016]
[0017] in, As the main factor of punishment, As a secondary punitive factor, As the main motivating factor, It is an absolute punitive factor.
[0018] Furthermore, in the above-mentioned reinforcement learning control method for ship cooling systems, the calculation formulas for the primary penalty factor, secondary penalty factor, primary incentive factor, and absolute penalty factor are as follows:
[0019]
[0020] in, Set the vacuum level for the condenser. This is the measured value of the condenser vacuum. This is the measured value of the condensate temperature in the condenser; This is the saturation temperature of the condensate in the condenser. This refers to the opening degree of the spray valve. This is a program termination signal. When system parameters deviate significantly from normal conditions, the training session ends prematurely. The value is either 0 or 1. When premature termination occurs... A reward is given when the condenser vacuum fluctuation is within ±0.5 kPa; otherwise, a penalty is imposed.
[0021] Furthermore, in the above-mentioned reinforcement learning control method for ship cooling systems, if the training results do not meet the requirements during online training, the network hyperparameters are adjusted and training is repeated; if satisfactory results still cannot be obtained after repeated training, the influence of measured values and action values on the evolution of the cooling system heat exchanger is reassessed, and the reward function is adjusted.
[0022] Furthermore, in the above-mentioned reinforcement learning control method for ship cooling systems, the policy network is selected as a BP neural network with a total of 4 fully connected layers. The first three fully connected layers are set with 20, 40, and 80 neurons respectively.
[0023] The value network is a recurrent neural network with temporal features. It is a multi-input neural network with two feature input layers, namely state variables and action values. The output layer is a fully connected layer. There are three fully connected layers before the summation layer, each containing 128 neurons. After the summation layer, there are five more fully connected layers besides the output layer, each containing 128 neurons.
[0024] Furthermore, in the aforementioned ship cooling system reinforcement learning control method, the input measured values include spray water specific enthalpy, turbine exhaust steam specific enthalpy, condenser vacuum degree, deviation of condenser vacuum degree from set value, condenser subcooling degree, and condensate circulation pump speed.
[0025] According to a second aspect of this application, a reinforcement learning control device for a ship cooling system is also provided, comprising:
[0026] The parameter determination module is used to determine the input measured values and output action values of the ship's cooling system. The input measured values are state parameters that are strongly correlated with the vacuum / subcooling control target. The output action values include the spray valve opening degree and the condensate circulation pump speed.
[0027] The reward function construction module is used to construct a reinforcement learning reward function that comprehensively considers both penalty and incentive factors. The reward function reflects the degree of influence of the seawater heat exchanger evolution process on the vacuum degree / subcooling degree.
[0028] The offline training module is used for offline training of network parameters, including offline training of the policy network and offline training of the value network. It uses data under traditional vacuum / supercooling control to pre-train the policy network, calculates the reward value based on the reward function and the measured values and action values of the policy network, and pre-trains the value network based on the measured values, action values, and reward values.
[0029] The online training module is used to deploy the offline trained policy network and value network to the cooling system control device and enter the online training phase. Through real-time interaction between the policy network, value network and cooling system, the parameters of the policy network and value network are optimized online using a reinforcement learning algorithm based on deep deterministic policy gradients until the vacuum / supercooling control effect required by the application is met.
[0030] According to a third aspect of this application, a reinforcement learning control device for a ship cooling system is also provided, comprising at least one processing unit and at least one storage unit, wherein the storage unit stores a computer program that, when executed by the processing unit, causes the processing unit to perform the steps of any of the methods described above.
[0031] According to a fourth aspect of this application, a storage medium is also provided, which stores a computer program executable by a ship cooling system reinforcement learning control device, which, when run on the ship cooling system reinforcement learning control device, causes the ship cooling system reinforcement learning control device to perform the steps of any of the methods described above.
[0032] In summary, compared with the prior art, the above-described technical solutions conceived in this application can achieve the following beneficial effects:
[0033] The reinforcement learning control method for ship cooling systems provided in this application achieves precise control and adaptive adjustment of the ship cooling system by determining the measured input values and output action values of the ship cooling system, constructing a reinforcement learning reward function, conducting offline training of network parameters, and optimizing network parameters online. It can automatically adjust the network weight parameters, effectively address system evolution problems, improve the control accuracy, reliability, and operating efficiency of the ship cooling system, reduce energy consumption and operating costs, and ensure the stable operation of the ship's power system. Attached Figure Description
[0034] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0035] Figure 1 A flowchart illustrating a reinforcement learning control method for a ship cooling system provided in this application embodiment;
[0036] Figure 2 This is a schematic diagram of the cooling system structure provided in an embodiment of this application;
[0037] Figure 3 This is a schematic diagram of the multivariable optimal coordination control logic flow of the cooling system provided in an embodiment of this application;
[0038] Figure 4 This is a schematic diagram of the online optimization process for network parameters provided in an embodiment of this application;
[0039] Figure 5 This is a schematic diagram of the strategy network structure provided in an embodiment of this application;
[0040] Figure 6 A schematic diagram of the value network structure provided for an embodiment of this application. Detailed Implementation
[0041] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application. Furthermore, the technical features involved in the various embodiments described below can be combined with each other as long as they do not conflict with each other.
[0042] The terms "first," "second," "third," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish different objects, not to describe a specific order. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or apparatus that includes a series of steps or units is not limited to the listed steps or units, but may optionally include steps or units not listed, or may optionally include other steps or units inherent to these processes, methods, products, or apparatuses.
[0043] Figure 1 A flowchart illustrating a reinforcement learning control method for a ship cooling system provided in this application embodiment is shown below. Figure 1 As shown in the figure, an embodiment of this application provides a reinforcement learning control method for a ship cooling system, which includes the following steps:
[0044] Determine the input measured values and output action values of the ship's cooling system. The input measured values are state parameters that are strongly correlated with the vacuum / subcooling control target. The output action values include the spray valve opening degree and the condensate circulation pump speed.
[0045] A reinforcement learning reward function is constructed, which comprehensively considers both penalty and incentive factors, and reflects the degree of influence of the seawater heat exchanger evolution process on the vacuum degree / subcooling degree through the reward function;
[0046] Offline training of network parameters is performed, including offline training of policy network and offline training of value network. The policy network is pre-trained using data under traditional vacuum / supercooling control. The reward value is calculated based on the reward function and the measured values and action values of the policy network. The value network is pre-trained based on the measured values, action values and reward values.
[0047] The offline-trained policy network and value network are deployed to the cooling system control device and enter the online training phase. Through real-time interaction between the policy network, value network and cooling system, the parameters of the policy network and value network are optimized online using a reinforcement learning algorithm based on deep deterministic policy gradients until the vacuum / supercooling control effect meets the application requirements.
[0048] Specifically, the ship's cooling system mainly consists of a mixing condenser, a seawater heat exchanger, a condensate circulation pump, a reflux regulating valve, a nozzle regulating valve assembly, connecting pipes and accessories, such as... Figure 2 As shown, the turbine exhaust steam entering the mixing condenser comes into direct contact with the cooling condensate sprayed from the nozzle regulating valve group. After thorough mixing and cooling, the condensate is drawn out by the variable frequency condensate circulation pump. The pressurized condensate from the circulation pump is divided into two streams: one stream returns directly to the mixing condenser via the return regulating valve without participating in heat exchange; the other stream passes through the seawater heat exchanger, exchanges heat with seawater, and then enters the nozzle regulating valve group. After flow regulation by the nozzle regulating valve group, the condensate is sprayed into the mixing condenser. Seawater, through the inlet and outlet pipes and flow guides on the ship's side, converts the dynamic pressure head of the oncoming water flow during ship navigation into the internal static pressure of the cooling system, driving the cooling seawater to flow through the seawater heat exchanger, cooling the condensate flowing through the seawater heat exchanger, thus meeting the cooling requirements of the ship's power plant.
[0049] Seawater heat exchangers in cooling systems are constantly immersed in the marine environment, inevitably leading to phenomena such as marine organisms adhering to the heat exchanger surface. This results in the formation of biofouling, sludge, and fouling layers, severely impacting the thermal-hydraulic performance and heat transfer efficiency of the seawater heat exchanger. In practical engineering applications, traditional vacuum / subcooling control strategies rely on trial-and-error tuning of control parameters to achieve stable control performance under operating conditions. Once determined, these parameters cannot be updated online based on system operating characteristics, making it difficult to adapt to the heat transfer evolution characteristics caused by marine organisms adhering to the seawater heat exchanger. This affects the precise control of vacuum / subcooling and the safe and reliable operation of the cooling system.
[0050] This application proposes a reinforcement learning control method for a ship cooling system, the control logic flow of which is as follows: Figure 3As shown, the first step is to determine the measured input values and output action values of the ship's cooling system. The measured input values encompass several key parameters closely related to the vacuum / subcooling control objectives, including the spray water specific enthalpy, the turbine exhaust steam specific enthalpy, the condenser vacuum, the deviation of the condenser vacuum from the setpoint, the condenser subcooling, and the condensate circulation pump speed. These parameters comprehensively reflect the system's operating status. The output action values include the spray valve opening and the condensate circulation pump speed; these are key action variables used by the control system to adjust the system's state.
[0051] A reinforcement learning reward function is constructed, which comprehensively considers multiple factors, including a primary penalty factor, secondary penalty factors, a primary incentive factor, and an absolute penalty factor. The primary penalty factor focuses on the deviation from the vacuum / subcooling target; the larger the deviation, the larger the penalty value, thus emphasizing the importance of precise control. Secondary penalty factors target changes in the spray valve opening and condensate circulation pump speed; the larger the change in these values, the larger the penalty value, used to avoid frequent and large-scale adjustments to the system, reducing unnecessary energy consumption and equipment wear. The primary incentive factor provides a positive reward when the actual output approaches the target value, guiding the system towards the target state. The absolute penalty factor is related to the training time, ensuring the integrity and stability of the training process. This multi-factor comprehensive reward function effectively reflects the degree of influence of the seawater heat exchanger evolution process on the vacuum / subcooling.
[0052] The offline training process for network parameters includes two parts: offline training of the policy network and offline training of the value network. In the offline training of the policy network, data from traditional vacuum / supercooling control is used to pre-train the policy network. This data contains the relationship between the measured input values and the corresponding output action values under traditional control methods, enabling the policy network to initially learn the system's control laws. The offline training of the value network, on the other hand, uses the reward function and the reward values calculated from the measured and action values of the policy network. Based on these measured, action, and reward values, the value network is pre-trained, allowing it to make preliminary evaluations of the policy network's output actions. The policy network uses a backpropagation (BP) neural network with four fully connected layers. The first three fully connected layers have 20, 40, and 80 neurons respectively. This structural design ensures sufficient computational power while enabling rapid response, meeting the requirements of real-time control. The value network uses a recurrent neural network with temporal features. It is a multi-input neural network with two feature input layers: state and action values. The output layer is a fully connected layer. There are three fully connected layers before the summation layer, each containing 128 neurons. After the summation layer, there are five more fully connected layers besides the output layer, each containing 128 neurons. This structure enables the value network to better handle the relationship between system states and actions, improving the accuracy of evaluating policy networks.
[0053] After the offline-trained policy network and value network are deployed to the cooling system control device, the system enters the online training phase. In this phase, the policy network and value network interact with the cooling system in real time. The policy network outputs corresponding action values based on the current measured input values, while the value network evaluates the policy network's performance based on these action values and the measured values, and adjusts the network parameters using a reinforcement learning algorithm based on deep deterministic policy gradients. For example, when the system detects a deviation in the condenser vacuum level, the policy network outputs adjusted spray valve openings and condensate circulation pump speeds based on the current measured values. The value network calculates reward values based on the new measured values and action values, and updates the network parameters accordingly to reduce the vacuum deviation until satisfactory control is achieved. If the training results are found to be unsatisfactory during online training, such as insufficient vacuum control accuracy, the network hyperparameters, such as the learning rate and discount factor, can be adjusted, and training can be repeated. If satisfactory results are not obtained after repeated training, the impact of measured values and action values on the evolution of the cooling system heat exchanger should be reassessed, and relevant parameters in the reward function, such as the gain factor, should be adjusted to better adapt to the evolution characteristics of the system.
[0054] The reinforcement learning control method for ship cooling systems provided in this application achieves precise control and adaptive adjustment of the ship cooling system by determining the measured input values and output action values of the ship cooling system, constructing a reinforcement learning reward function, conducting offline training of network parameters, and optimizing network parameters online. It can automatically adjust the network weight parameters, effectively address system evolution problems, improve the control accuracy, reliability, and operating efficiency of the ship cooling system, reduce energy consumption and operating costs, and ensure the stable operation of the ship's power system.
[0055] Optionally, the reinforcement learning control method for ship cooling systems provided in this application embodiment, which comprehensively considers penalty factors and incentive factors, specifically includes:
[0056] The main penalty factor is given based on the deviation from the vacuum / subcooling target; the larger the deviation, the larger the penalty value.
[0057] Secondary penalty factors are given based on the changes in the spray valve opening and the condensate circulation pump speed. The greater the change in the speed, the greater the penalty.
[0058] The main incentive factors are given based on the degree of proximity to the vacuum / supercooling target. When the actual output is close to the target value, a positive reward is given.
[0059] Absolute penalties are based on training time, and penalties are imposed when training is terminated prematurely.
[0060] Optionally, in the reinforcement learning control method for ship cooling systems provided in this application embodiment, the reward function is:
[0061]
[0062] in, As the main factor of punishment, As a secondary punitive factor, As the main motivating factor, It is an absolute punitive factor.
[0063] Optionally, in the reinforcement learning control method for ship cooling systems provided in this application embodiment, the calculation formulas for the primary penalty factor, secondary penalty factor, primary incentive factor, and absolute penalty factor are as follows:
[0064]
[0065] in, Set the vacuum level for the condenser. This is the measured value of the condenser vacuum. This is the measured value of the condensate temperature in the condenser; This is the saturation temperature of the condensate in the condenser. This refers to the opening degree of the spray valve. This is a program termination signal. When system parameters deviate significantly from normal conditions, the training session ends prematurely. The value is either 0 or 1. When premature termination occurs... A reward is given when the condenser vacuum fluctuation is within ±0.5 kPa; otherwise, a penalty is imposed.
[0066] Specifically, to measure the impact of the seawater heat exchanger evolution process on vacuum / subcooling, a reinforcement learning reward function needs to be designed. In the offline training of the policy network and value network parameters, the vacuum / subcooling control objective follows the effect of traditional control actions. In the online optimization module of the policy network and value network parameters, the vacuum / subcooling control objective is updated to output action values that maintain a constant condenser vacuum and meet the subcooling requirements. This application mainly considers the following factors:
[0067] Key penalty factors: First, it is necessary to ensure that the vacuum / subcooling control targets are achieved, i.e., the condenser vacuum is constant and the subcooling meets the requirements. The square of the deviation between the actual vacuum / subcooling output and the target value is used as the penalty value; the larger the deviation, the larger the penalty value. An appropriate gain factor is set to balance the weights among the targets, and this penalty factor is ensured to dominate.
[0068] Secondary penalty factors: The changes and fluctuations in control variables should be minimized, meaning the deviations between the output action values (such as the spray valve opening and condensate circulation pump speed) and the steady-state action values, as well as the deviations between two output steps, should be as small as possible. The penalty values are the squared deviations between the actual output action value and the initial or steady-state action value, and the squared deviations between two output steps; the larger the deviation, the larger the penalty value. An appropriate gain factor is set to balance the weights of each action value, ensuring that this penalty factor is weaker than the primary penalty factor.
[0069] Key incentive factors: A positive reward is designed to ensure the output continuously approaches the target vacuum / supercooling value; that is, a positive reward is given when the actual output approaches the target value. Appropriate rewards are given based on the degree of closeness between the actual output value and the target value. An appropriate gain factor is set to ensure that this reward factor is comparable to the main penalty factor.
[0070] Absolute penalty factor: This ensures the program can be trained completely, i.e., a penalty is imposed when training is prematurely terminated. It is assumed that after the training termination time, each penalty factor will give its maximum penalty value at each step. Penalties are imposed based on the estimated maximum penalty value at each step and the remaining untrained time, ensuring that the longer the training time, the greater the total reward value.
[0071] The values of each factor are calculated in detail. Taking the condenser vacuum setpoint as an example, let's assume it's a specific value. In actual operation, the measured condenser vacuum may vary due to factors such as the fouling layer on the seawater heat exchanger. Substituting the measured value and the setpoint into the calculation formula yields the value of the primary penalty factor, reflecting the degree of influence of vacuum deviation on the control effect. Similarly, based on parameters such as spray valve opening and program termination signal, the values of secondary penalty factors, primary excitation factors, and absolute penalty factors are calculated respectively. For example, when the spray valve opening changes significantly, the value of the secondary penalty factor will increase accordingly, indicating that the control system needs to adjust its strategy to reduce unnecessary changes in action.
[0072] Optionally, in the reinforcement learning control method for ship cooling systems provided in this application embodiment, if the training results do not meet the requirements during online training, the network hyperparameters are adjusted and training is performed again; if satisfactory results still cannot be obtained after repeated training, the influence of measured values and action values on the evolution of the cooling system heat exchanger is reassessed, and the reward function is adjusted.
[0073] Specifically, the offline-trained policy network and value network are deployed to the cooling system control device, enabling real-time interaction between the policy network, value network, and cooling system, thus entering the online training phase and achieving online optimization of the policy network and value network parameters. A reinforcement learning algorithm based on deep deterministic policy gradients is used to achieve online optimization of the policy network and value network parameters. The online training and optimization process is as follows: Figure 4As shown, the online network parameter optimization module mainly consists of a policy network, a value network, measured values, action values, and a reward function. It obtains the initial parameters of the policy network and value network through an offline training process, and then deploys them to the cooling system control device for online training and parameter updates.
[0074] Considering the evolution process of the seawater heat exchanger during the actual operation of the cooling system, if the vacuum / subcooling results obtained during online training meet the requirements, the training process is stopped and the weights of the policy network and value network are saved. If the vacuum / subcooling results obtained during online training do not meet the requirements, the network hyperparameters need to be adjusted and training is repeated. If satisfactory vacuum / subcooling results cannot be obtained even after repeated training, the influence of measured values and action values on the evolution of the cooling system heat exchanger needs to be re-evaluated, and the reward function needs to be adjusted until a satisfactory vacuum / subcooling control effect is achieved.
[0075] Optionally, in the reinforcement learning control method for ship cooling systems provided in this application embodiment, the policy network is selected as a BP neural network with a total of 4 fully connected layers. The first three fully connected layers are set with 20, 40, and 80 neurons respectively.
[0076] The value network is a recurrent neural network with temporal features. It is a multi-input neural network with two feature input layers, namely state variables and action values. The output layer is a fully connected layer. There are three fully connected layers before the summation layer, each containing 128 neurons. After the summation layer, there are five more fully connected layers besides the output layer, each containing 128 neurons.
[0077] Specifically, the strategy network, after training, obtains a control strategy for vacuum / subcooling, primarily functioning to output action values such as the spray valve opening and the condensate circulation pump speed. The strategy network outputs control action values based on measured values. To ensure the real-time performance of reinforcement learning control, a simple BP neural network is typically chosen, with a structure as follows: Figure 5 As shown, the strategy network mainly consists of fully connected layers and rule layers. The input layer of the strategy network is a feature layer, which takes into account the state variables of the cooling system, i.e., the measured values selected in this application. The output layer is a fully connected layer, which outputs control action values, i.e., the opening degree of the spray valve and the speed of the condensate circulation pump. The strategy network has a total of 4 fully connected layers. Except for the last fully connected layer, the other fully connected layers have 20, 40, and 80 neurons respectively.
[0078] The offline training data for the policy network includes measured values and action values of the cooling system under traditional vacuum / supercooling control. Based on the measured values, action values are output, training a policy network that achieves the same effect as traditional vacuum / supercooling control. The parameters of the policy network in the trained agent are then stored. This process can be viewed as using a neural network to simulate traditional control actions, also known as simulation learning.
[0079] The value network evaluates the performance of the policy network during training, but does not participate in the vacuum / supercooling control process of the cooling system after training. The value network needs to evaluate the policy network's performance based on the cooling system state and the policy network's output, resulting in high complexity. The value network structure is as follows: Figure 6 As shown, this is a multi-input neural network with two feature input layers: one for state variables and one for action values. The output layer is a fully connected layer. There are three fully connected layers before the summing layer, each containing 128 neurons. After the summing layer, there are five more fully connected layers besides the output layer, each containing 128 neurons. Due to the complexity of the cooling system, including time delays and strong nonlinearity, a recurrent neural network incorporating time features is chosen here to identify the system state and evaluate the network's effectiveness. This results in more accurate system state identification, but requires more training time and computational resources.
[0080] The value network simultaneously receives measured values and action values from the cooling system and provides reward function values. Therefore, after completing the offline training of the policy network, the value network is trained based on the measured values, action values, and reward values to obtain the initial offline parameters.
[0081] Optionally, the reinforcement learning control method for ship cooling systems provided in this application embodiment includes the input measured values such as spray water specific enthalpy, turbine exhaust steam specific enthalpy, condenser vacuum degree, deviation of condenser vacuum degree from set value, condenser subcooling degree, and condensate circulation pump speed.
[0082] Optionally, embodiments of this application also provide a reinforcement learning control device for a ship cooling system, comprising:
[0083] The parameter determination module is used to determine the input measured values and output action values of the ship's cooling system. The input measured values are state parameters that are strongly correlated with the vacuum / subcooling control target. The output action values include the spray valve opening degree and the condensate circulation pump speed.
[0084] The reward function construction module is used to construct a reinforcement learning reward function that comprehensively considers both penalty and incentive factors. The reward function reflects the degree of influence of the seawater heat exchanger evolution process on the vacuum degree / subcooling degree.
[0085] The offline training module is used for offline training of network parameters, including offline training of the policy network and offline training of the value network. It uses data under traditional vacuum / supercooling control to pre-train the policy network, calculates the reward value based on the reward function and the measured values and action values of the policy network, and pre-trains the value network based on the measured values, action values, and reward values.
[0086] The online training module is used to deploy the offline trained policy network and value network to the cooling system control device and enter the online training phase. Through real-time interaction between the policy network, value network and cooling system, the parameters of the policy network and value network are optimized online using a reinforcement learning algorithm based on deep deterministic policy gradients until the vacuum / supercooling control effect required by the application is met.
[0087] This application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the above-described method. The computer-readable storage medium may include, but is not limited to, any type of disk, including floppy disks, optical disks, DVDs, CD-ROMs, microdrives, as well as magneto-optical disks, ROMs, RAMs, EPROMs, EEPROMs, DRAMs, VRAMs, flash memory devices, magnetic cards or optical cards, nanosystems (including molecular memory ICs), or any type of medium or device suitable for storing instructions and / or data.
[0088] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that this application is not limited to the described order of actions, as some steps may be performed in other orders or simultaneously according to this application. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions and modules involved are not necessarily essential to this application.
[0089] In the above embodiments, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions in other embodiments.
[0090] In the several embodiments provided in this application, it should be understood that the disclosed apparatus can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some service interface; the indirect coupling or communication connection between devices or units may be electrical or other forms.
[0091] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the objectives of the embodiments of this application, depending on actual needs.
[0092] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0093] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage device (CMD). Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a memory and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned memory includes various media capable of storing program code, such as USB flash drives, read-only memory (ROM), random access memory (RAM), portable hard drives, magnetic disks, or optical disks.
[0094] Those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be implemented by a program instructing related hardware. The program can be stored in a computer-readable storage medium, which may include: a flash drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk, etc.
[0095] The foregoing description is merely an exemplary embodiment of this disclosure and should not be construed as limiting the scope of this disclosure. Any equivalent changes and modifications made in accordance with the teachings of this disclosure shall still fall within the scope of this disclosure. Those skilled in the art will readily conceive of embodiments of this disclosure upon considering the specification and practicing the disclosure herein. This application is intended to cover any variations, uses, or adaptations of this disclosure that follow the general principles of this disclosure and include common knowledge or customary techniques in the art not described herein. The specification and embodiments are to be considered exemplary only, and the scope and spirit of this disclosure are defined by the claims.
[0096] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0097] Those skilled in the art will readily understand that the above description is merely a preferred embodiment of this application and is not intended to limit this application. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this application should be included within the scope of protection of this application.
Claims
1. A reinforcement learning control method for a ship cooling system, characterized in that, Includes the following steps: Determine the input measured values and output action values of the ship's cooling system. The input measured values are state parameters that are strongly correlated with the vacuum / subcooling control target. The output action values include the spray valve opening degree and the condensate circulation pump speed. A reinforcement learning reward function is constructed, which comprehensively considers both penalty and incentive factors, and reflects the degree of influence of the seawater heat exchanger evolution process on the vacuum degree / subcooling degree through the reward function; Offline training of network parameters is performed, including offline training of policy network and offline training of value network. The policy network is pre-trained using data under traditional vacuum / supercooling control. The reward value is calculated based on the reward function and the measured values and action values of the policy network. The value network is pre-trained based on the measured values, action values and reward values. The offline trained policy network and value network are deployed to the cooling system control device and enter the online training stage. Through the real-time interaction between the policy network, value network and cooling system, the parameters of the policy network and value network are optimized online using a reinforcement learning algorithm based on deep deterministic policy gradient until the vacuum / supercooling control effect required by the application is met. The reward function is: in, As the main factor of punishment, As a secondary punitive factor, As the main motivating factor, As an absolute punitive factor; The calculation formulas for the primary punishment factor, secondary punishment factor, primary incentive factor, and absolute punishment factor are as follows: in, Set the vacuum level for the condenser. This is the measured value of the condenser vacuum. This is the measured value of the condensate temperature in the condenser; This is the saturation temperature of the condensate in the condenser. This refers to the opening degree of the spray valve. This is a program termination signal. When system parameters deviate significantly from normal conditions, the training session ends prematurely. The value is either 0 or 1. When premature termination occurs... A reward is given when the condenser vacuum fluctuation is within ±0.5 kPa; otherwise, a penalty is imposed.
2. The reinforcement learning control method for a ship cooling system as described in claim 1, characterized in that, The comprehensive consideration of punitive and incentive factors specifically includes: The main penalty factor is given based on the deviation from the vacuum / subcooling target; the larger the deviation, the larger the penalty value. Secondary penalty factors are given based on the changes in the spray valve opening and the condensate circulation pump speed. The greater the change in the speed, the greater the penalty. The main incentive factors are given based on the degree of proximity to the vacuum / supercooling target. When the actual output is close to the target value, a positive reward is given. Absolute penalties are based on training time, and penalties are imposed when training is terminated prematurely.
3. The reinforcement learning control method for a ship cooling system as described in claim 1, characterized in that, During online training, if the training results do not meet the requirements, the network hyperparameters are adjusted and training is repeated; if satisfactory results still cannot be obtained after repeated training, the influence of measured values and action values on the evolution of the cooling system heat exchanger is reassessed, and the reward function is adjusted.
4. The reinforcement learning control method for a ship cooling system as described in claim 1, characterized in that, The strategy network is a backpropagation neural network with four fully connected layers. The first three fully connected layers have 20, 40, and 80 neurons, respectively. The value network is a recurrent neural network with temporal features. It is a multi-input neural network with two feature input layers, namely state variables and action values. The output layer is a fully connected layer. There are three fully connected layers before the summation layer, each containing 128 neurons. After the summation layer, there are five more fully connected layers besides the output layer, each containing 128 neurons.
5. The reinforcement learning control method for a ship cooling system as described in claim 1, characterized in that, The input measured values include spray water specific enthalpy, turbine exhaust steam specific enthalpy, condenser vacuum degree, deviation of condenser vacuum degree from set value, condenser subcooling degree, and condensate circulation pump speed.
6. A reinforcement learning control device for a ship cooling system, characterized in that, include: The parameter determination module is used to determine the input measured values and output action values of the ship's cooling system. The input measured values are state parameters that are strongly correlated with the vacuum / subcooling control target. The output action values include the spray valve opening degree and the condensate circulation pump speed. The reward function construction module is used to construct a reinforcement learning reward function that comprehensively considers both penalty and incentive factors. The reward function reflects the degree of influence of the seawater heat exchanger evolution process on the vacuum degree / subcooling degree. The offline training module is used for offline training of network parameters, including offline training of the policy network and offline training of the value network. It uses data under traditional vacuum / supercooling control to pre-train the policy network, calculates the reward value based on the reward function and the measured values and action values of the policy network, and pre-trains the value network based on the measured values, action values, and reward values. The online training module is used to deploy the offline trained policy network and value network to the cooling system control device and enter the online training stage. Through the real-time interaction between the policy network, value network and cooling system, the parameters of the policy network and value network are optimized online using a reinforcement learning algorithm based on deep deterministic policy gradient until the vacuum / supercooling control effect required by the application is met. The reward function is: in, As the main factor of punishment, As a secondary punitive factor, As the main motivating factor, As an absolute punitive factor; The calculation formulas for the primary punishment factor, secondary punishment factor, primary incentive factor, and absolute punishment factor are as follows: in, Set the vacuum level for the condenser. This is the measured value of the condenser vacuum. This is the measured value of the condensate temperature in the condenser; This is the saturation temperature of the condensate in the condenser. This refers to the opening degree of the spray valve. This is a program termination signal. When system parameters deviate significantly from normal conditions, the training session ends prematurely. The value is either 0 or 1. When premature termination occurs... A reward is given when the condenser vacuum fluctuation is within ±0.5 kPa; otherwise, a penalty is imposed.
7. A reinforcement learning control device for a ship cooling system, characterized in that, It includes at least one processing unit and at least one storage unit, wherein the storage unit stores a computer program that, when executed by the processing unit, causes the processing unit to perform the steps of the method according to any one of claims 1 to 5.
8. A storage medium, characterized in that, It stores a computer program that can be executed by a ship cooling system reinforcement learning control device, which, when run on the ship cooling system reinforcement learning control device, causes the ship cooling system reinforcement learning control device to perform the steps of the method according to any one of claims 1 to 5.