Optimal control method for Rankine cycle waste heat recovery system based on deep reinforcement learning
By training the intelligent agent controller through deep reinforcement learning, the problem of efficient optimization control of the Rankine cycle waste heat recovery system under transient fluctuating heat source conditions was solved, and the safety and real-time performance of the system were improved.
Patent Information
- Application Number
- CN202211384624.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-11-07
- Publication Date
- 2025-09-09
- Estimated Expiration
- 2042-11-07
AI Technical Summary
The existing Rankine cycle waste heat recovery system is difficult to achieve efficient optimization control under transient fluctuating heat source conditions, and cannot take into account system safety.
A control method based on deep reinforcement learning is adopted. By establishing a dynamic simulation model and reward mechanism, the intelligent agent controller is trained to directly adjust the working fluid flow to optimize the system output power while ensuring system safety.
It achieves efficient optimization control and safety assurance of the system under transient fluctuation conditions, reduces the amount of computing tasks, and improves the real-time performance and output power of the system.
Smart Images

Figure CN115857330B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of intelligent control of energy power systems, and specifically relates to an optimization control method for a Rankine cycle waste heat recovery system based on deep reinforcement learning. Background Art
[0002] A large amount of waste heat is generated in modern industrial production and transportation, such as the waste heat from lime kiln flue gas in steel production, the exhaust gas and cylinder jacket water waste heat from internal combustion engines, etc. The effective utilization of this part of waste heat is of great significance to energy conservation, emission reduction and the realization of dual carbon goals. The Rankine cycle (including various working fluids) is considered to be the mainstream waste heat recovery technology at present due to its high thermal efficiency, strong adaptability of working fluids, and mature components. However, since the conditions of the waste heat source may continue to change in practice, the most typical example is the transient fluctuation of the waste heat of automotive internal combustion engines as the vehicle operating conditions change. Therefore, the Rankine cycle waste heat recovery system may often be in transient conditions. In order to ensure the safe and efficient operation of the waste heat recovery system, system control is crucial.
[0003] Existing Rankine cycle waste heat recovery systems use a supervisory control method based on steady-state calibration, but this fails under transient heat source conditions. Some scholars have proposed using model-based prediction optimization control to address this issue. While model-based online optimization can yield optimized control actions for each non-steady-state condition, the complexity and nonlinearity of the system model, coupled with strict safety constraints, creates a massive online computational workload, making it difficult to ensure real-time control. Furthermore, accurate predictions of future operating conditions significantly impact the effectiveness of model-based control, while precise predictions of truck operating conditions are difficult in reality. Furthermore, the system also has stringent safety requirements for overtemperature and overpressure. Therefore, efficient optimization control of Rankine cycle waste heat recovery systems while balancing safety requirements under fluctuating heat source conditions is currently a significant challenge.
[0004] Deep reinforcement learning (DRL) is a combination of deep neural networks and reinforcement learning (RL). It has powerful perception and decision-making capabilities. The general learning process can be described as follows:
[0005] (1) At each moment, the agent interacts with the environment to observe the state of the environment and uses a deep neural network to perceive the state to obtain a specific state feature representation;
[0006] (2) Evaluate the value function of each action based on the immediate reward fed back by the environment, select actions to maximize future rewards, and map the current state to the corresponding action through a deep neural network;
[0007] (3) The environment responds to this action and obtains the next observation. By continuously looping the above process and continuously updating the strategy, the optimal strategy to achieve the goal can be obtained.
[0008] To this end, the DRL algorithm is considered to be introduced to solve the optimization control problem of the Rankine cycle waste heat recovery system under transient fluctuating heat source conditions while taking into account safety. Summary of the Invention
[0009] The purpose of the present invention is to overcome the defects of the prior art and propose a Rankine cycle waste heat recovery system optimization control method based on deep reinforcement learning. The control method does not use the traditional PID control (proportional integral differential control) algorithm. Based on the deep reinforcement learning algorithm, it directly takes action by observing the state of the waste heat recovery system, thereby changing the working fluid flow rate to obtain the maximum system output power, thereby solving the problem of optimizing the control of the Rankine cycle waste heat recovery system under transient fluctuating heat source conditions while taking into account safety requirements.
[0010] The optimization control method of the Rankine cycle waste heat recovery system based on deep reinforcement learning includes:
[0011] Step 1: Establish a dynamic simulation model of the Rankine cycle waste heat recovery system to form a learning environment for the deep reinforcement learning algorithm; use an intelligent agent as the controller of the waste heat recovery system, and the dynamic simulation model of the waste heat recovery system is the environment for interaction with the intelligent agent; the dynamic simulation model includes a heat exchanger model, pump, expander, liquid storage tank, and various valves and pipelines;
[0012] Step 2: Design the rewards, observations, states, and actions of the deep reinforcement learning algorithm; specifically:
[0013] The reward of the deep reinforcement learning algorithm is set by the reward function, which is:
[0014]
[0015] Where r represents the reward, Wnet represents the net work of the system, k is a proportional coefficient, p represents pressure, T represents temperature, subscript t represents turbine, in represents inlet, and max represents the maximum allowable value. If the action causes the system state to exceed the preset safety limit, a negative reward, i.e., a penalty, is obtained, and the current training segment is stopped and a new training segment is restarted.
[0016] The temperature and pressure of the working fluid at the inlet of the expander in the waste heat recovery system and the temperature and flow rate of the heat source entering the waste heat recovery system are set as observation quantities of the intelligent agent; the intelligent agent makes a decision action by observing the system state, and the action is to control the working fluid flow rate by directly or indirectly adjusting the speed of the pump, so as to achieve the optimal temperature and pressure state of the system;
[0017] The agent evaluates the quality of its actions based on the rewards it receives, thereby continuously interacting with the environment and converging towards a direction with greater rewards. Ultimately, an agent controller is trained to obtain the maximum reward, that is, the maximum accumulated output work.
[0018] Step 3: After setting the training environment boundary conditions, the intelligent controller is trained according to the preset training algorithm. After training, several sets of untrained heat source boundary conditions that fluctuate over time and vary within the same amplitude range as the training heat source fluctuation conditions are input into the Rankine cycle waste heat recovery system for testing.
[0019] When the system cumulative output work of the test results is less than the system cumulative output work in the actual operation data according to the traditional PID constant temperature or constant pressure control method, return to step three, reset the boundary conditions of the training environment and the training algorithm parameters, continue training until the system cumulative output work is better, end the training, and use the obtained intelligent agent controller for the optimal control of the Rankine cycle waste heat recovery system.
[0020] Furthermore, the system is a high-precision dynamic simulation model of a Rankine cycle waste heat recovery system, or a real Rankine cycle waste heat recovery system; when the system is a real Rankine cycle waste heat recovery system, a system simulation model based on the real Rankine cycle waste heat recovery system is established before executing step one, and the system simulation model is used to pre-train the reinforcement learning agent for optimization control. The training process is the same as steps one to three above. Only after the pre-training, the agent can use the real Rankine cycle waste heat recovery system as the training environment for the next step.
[0021] Furthermore, the optimization control method is applicable to a subcritical organic Rankine cycle and a transcritical organic Rankine cycle.
[0022] Furthermore, the training algorithm of the intelligent agent controller includes a deep deterministic policy gradient algorithm and a deep Q network algorithm.
[0023] Furthermore, in step three, the selection of training environment boundary conditions includes the following steps: From the actual operating data of the waste heat recovery system, a set of randomly fluctuating heat source data within the range of the waste heat recovery system's heat source's regular fluctuations is selected as the training environment boundary conditions, including heat source flow rate and temperature data. A larger data sample size improves the control effect after training the agent, and also improves extrapolation and stability. However, this also results in higher training costs, which requires consideration.
[0024] Compared with the existing technology, the optimization control method of the Rankine cycle waste heat recovery system based on deep reinforcement learning described in the present invention has the following advantages and beneficial effects:
[0025] The optimization control method of the present invention has a very small computational task in actual use, does not require prediction of future interference, and does not require a large amount of online optimization calculations, thus effectively ensuring the real-time performance of the control.
[0026] The optimization control method can not only enable the system to obtain more recovery power, but also ensure the safety of the system. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] Figure 1 This is a schematic diagram of an embodiment of the present invention using the optimized control method for the Rankine cycle waste heat recovery system and using the working fluid pump speed as the control action;
[0028] Figure 2 This is a schematic diagram of an embodiment of the present invention using the optimized control method for the Rankine cycle waste heat recovery system and using the expander inlet state signal as a control action;
[0029] Figure 3 is the heat source boundary condition during the deep reinforcement learning algorithm training in the embodiment;
[0030] Figure 4 is the untrained heat source boundary condition in the embodiment. DETAILED DESCRIPTION
[0031] The present invention will be further described below with reference to the accompanying drawings and specific embodiments. The embodiments of the present invention are intended to enable those skilled in the art to better understand the present invention and are not intended to limit the present invention in any way. The workflow and working principle of the present invention will be further described below using a preferred embodiment of the present invention.
[0032] The Rankine cycle waste heat recovery system can be divided into a subcritical cycle and a transcritical cycle according to whether the parameters at the high-pressure end are in a supercritical state. The present invention is applicable to both subcritical and transcritical cycles and is not limited to working fluids and cycle configurations. That is, the control method is applicable to Rankine cycle systems with any working fluid and cycle configuration. The following embodiments are for a transcritical organic Rankine cycle (ORC).
[0033] The existing Rankine cycle waste heat recovery system uses a supervisory control method based on steady-state calibration. The present invention uses deep reinforcement learning (DRL) to solve the optimization control problem of the transcritical Rankine cycle waste heat recovery system while taking into account safety under transient fluctuating heat source conditions.
[0034] The key elements of a deep reinforcement learning algorithm are environment, reward, agent, state, and action.
[0035] In this embodiment, the agent observes the working fluid temperature and pressure at the expander inlet, as well as the flow rate and inlet temperature of the heat source. The agent makes decisions based on the state of the environment (i.e., the waste heat recovery system). In this embodiment, the action is a control signal for the pump speed, which changes the working fluid flow rate. The agent evaluates the quality of its actions based on the rewards it receives. This allows it to continuously interact with the environment and learn the actions that maximize reward under each environmental state. Ultimately, it trains an agent controller that maximizes cumulative reward, i.e., cumulative output work.
[0036] like Figure 1 As shown in FIG, a deep reinforcement learning-based optimization control method for a Rankine cycle waste heat recovery system using pump speed as a control action includes:
[0037] Step 1: Establish a learning environment for deep reinforcement learning algorithms
[0038] S101: Use Simulink to establish a dynamic simulation model of a transcritical organic Rankine cycle (ORC) waste heat recovery system, including establishing dynamic simulation models of each component in the system, and then connecting them to each other according to the relationship between the components to form a dynamic simulation model of the system; specifically including
[0039] (1) Heat exchanger model
[0040] In system simulation models, the heat exchanger is often simplified to a typical countercurrent heat exchanger. The heat exchanger is divided into multiple control volumes, and the hot and cold fluids in each control volume obey the same mass and energy conservation equations (ignoring momentum conservation), as shown in equations (1) and (2). The tube wall has no mass conservation equations, only the energy conservation equation (3). The entire set of equations (1-3) for all control volumes constitutes the simulation model of the entire heat exchanger. Boundary conditions and initial values can then be assigned to solve the equations simultaneously.
[0041]
[0042]
[0043] in, V, p, and They represent the mass flow rate of the fluid, the volume of the control body, the average density, the pressure, the average enthalpy and the average temperature respectively; the subscripts "in" and "out" represent the inlet and outlet of the control body; α and A represent the heat transfer coefficient and heat transfer area between the fluid and the tube wall; the subscript "w" represents the tube wall; the derivation result of the above formula is simplified based on the fact that the average density is a function of the average enthalpy and pressure.
[0044] (2) Pump and expander
[0045] Since the response speed of the pump and expander is much faster than that of the heat exchanger, a steady-state model is generally used for system simulation. The pump in this system is a positive displacement diaphragm pump, and its model is:
[0046]
[0047] h p_out =h p_in +(h sp_out -h p_in ) / η sp (5)
[0048] Among them, η v is the volumetric efficiency, ρ p is the working fluid density, V cyl is the pump capacity, and ω is the pump speed, which can be controlled by a frequency converter and is the key control variable of this system. The subscript p represents pump, and s represents isentropic.
[0049] The expander model is simplified to a valve, as shown in Equation (6-7). It's important to note that this model is consistent with the actual system used in the experimental verification below, where an expansion valve is used instead of an expander. The subscript t represents the expander.
[0050]
[0051] h t_out =h t_in -(h t_in -h st_out )η st (7)
[0052] (3) Liquid storage tank
[0053] The liquid storage tank can store excess liquid fluid to ensure liquid at the pump inlet and maintain the system's cold-end pressure. It is usually considered a control volume, with the fluid inside in a two-phase coexistence state. Ignoring the momentum conservation equation, the derived mass conservation equation and energy conservation equation are as follows:
[0054]
[0055] The subscripts “l”, “v” and “rec” represent saturated liquid, saturated gas and liquid storage tank, respectively.
[0056] S102: Using an intelligent agent as a controller of the waste heat recovery system, and an environment interacting with the intelligent agent is a dynamic simulation model of the waste heat recovery system.
[0057] Step 2: Set up a reward function. Rewards are calculated using a reward function. Since the goal of the waste heat recovery system is to maximize waste heat recovery and increase output power, the reward function specifies that the greater the system's output power, the greater the reward for that action.
[0058] The specific reward function is shown as follows:
[0059]
[0060] Where r represents the reward, Wnet represents the net work of the system, k is a proportional coefficient, p represents pressure, T represents temperature, the subscript t represents the turbine, in represents the inlet, and max represents the maximum allowable value. If the action causes the system state to exceed the safety limit (7 MPa in this embodiment, the specific value varies greatly depending on the actual waste heat recovery system), a negative reward, i.e., a penalty, is obtained, the current training segment is terminated, and a new training segment is restarted.
[0061] Step 3: Select a training algorithm: Use the Deep Deterministic Policy Gradient (DDPG) algorithm to train the DRL agent controller to achieve safe and optimal control of the transcritical ORC system. Depending on the actual needs, various general deep reinforcement learning algorithms can be used, such as the Deep Q-Network (DQN) algorithm.
[0062] Step 4: Select the observed state. The agent's observations of the environment should be a set of variables that uniquely represent the current state of the system. The observed variables are the temperature and pressure of the working fluid at the expander inlet of the waste heat recovery system, and the temperature and flow rate of the internal combustion engine flue gas entering the waste heat recovery system.
[0063] Step 5: Select an agent action. The Rankine cycle waste heat recovery system described herein primarily adapts to changes in the heat source by directly or indirectly controlling the flow rate of the working fluid, thereby achieving safe and efficient control. In this step of the present embodiment, the flow rate signal or speed signal of the working fluid pump in the system is used as the agent action to directly control the system's working fluid flow rate.
[0064] Step 6: Train the DRL agent controller
[0065] S601: From the actual operation data of the waste heat recovery system, a set of randomly fluctuating heat source data within the range of the heat source of the waste heat recovery system that often fluctuates is selected as the boundary conditions of the training environment, including the flow rate and temperature data of the heat source, such as Figure 3As shown in the figure, the policy gradient algorithm is used to train the DRL agent controller based on the selected depth. A larger data sample size results in better control performance after the agent is trained, and better extrapolation and stability are achieved. However, this also results in higher training costs, which requires consideration.
[0066] S602: In order to test the effect of the trained DRL agent controller, several sets of heat source boundary conditions that have not been trained and have the same amplitude range as the heat source fluctuation conditions trained in S601 and fluctuate over time are set. Figure 4 The input shown is given to the Rankine cycle waste heat recovery system (i.e., the environment of the deep reinforcement learning algorithm) to test the safety optimization control effect of the intelligent agent on the system after training.
[0067] If the system cumulative output work of the test result is less than the system cumulative output work of the traditional PID constant temperature or constant pressure control method (controlling the expander inlet working fluid temperature or pressure to a constant value), return to step 3, reset the boundary conditions of the training environment and the training algorithm parameters, and continue training until the system cumulative output work is better. End the training and use the obtained intelligent agent controller for the optimization control of the Rankine cycle waste heat recovery system;
[0068] Only after testing can the intelligent agent perform efficient and optimized control of the Rankine cycle waste heat recovery system while taking into account safety under transient fluctuating heat source conditions.
[0069] By continuously training the DRL agent, it learned a strategy for achieving maximum power output. This differs from existing PID thermostats, which control the working fluid temperature at the expander inlet at a constant value. PID control only tracks the target temperature and fails to recognize the rapid increase in flow rate to increase power output.
[0070] Example 2
[0071] like Figure 2 As shown, this embodiment is directed to a waste heat recovery system of a transcritical organic Rankine cycle (ORC). Its control method is similar to that of Example 1, and only its distinguishing features are described below:
[0072] In step 5, the reference control signal of the temperature or pressure at the expander inlet in the waste heat recovery system is used as the agent action. The working fluid flow controller tracks the temperature or pressure state of the expander inlet specified by the agent action by controlling the working fluid flow rate. Therefore, the agent action is also the reference signal of the system working fluid flow controller. If the action setting scheme of Example 1 is adopted, then due to the inexplicability of the deep neural network, the agent may make some unreasonable actions that are difficult to understand, affecting the safety of the system; if the action setting scheme of Example 2 is adopted, as long as the range of the action signal is limited to below the system safety state value, such as below the safe temperature and pressure (this embodiment uses the temperature reference control signal as the agent action, with a safety range of 100-200°C), then even if the deep neural network is inexplicable, the agent cannot give the working fluid flow controller a reference tracking signal that exceeds the safety value. At most, an unreasonable action affects the efficiency of the system, thereby ensuring that the system always operates in a safe state. However, due to the interference of the flow controller, the optimization effect of the system waste heat recovery performance of Example 2 is likely to be inferior to that of Example 1. In short, the action setting scheme of Example 1 has better optimization effect, but poor safety; the action setting scheme of Example 2 has worse optimization effect, but high safety. .
[0073] The embodiments described above are only used to illustrate the technical ideas and features of the present invention. Their purpose is to enable those skilled in the art to understand the contents of the present invention and implement them accordingly. The scope of the patent of the present invention cannot be limited by these embodiments alone. That is, any equivalent changes or modifications made to the spirit disclosed by the present invention still fall within the scope of the patent of the present invention.
Claims
1. Optimal control method for Rankine cycle waste heat recovery system based on deep reinforcement learning, including: Step 1: Establish a dynamic simulation model of the Rankine cycle waste heat recovery system to form a learning environment for the deep reinforcement learning algorithm; An intelligent agent is used as a controller of the waste heat recovery system, and a dynamic simulation model of the waste heat recovery system is an environment that interacts with the intelligent agent; the dynamic simulation model includes a heat exchanger model, a pump, an expander, a liquid storage tank, and various valves and pipelines; Step 2: Design the rewards, observations, states, and actions of the deep reinforcement learning algorithm; specifically: The reward of the deep reinforcement learning algorithm is set by the reward function, which is: Where r represents the reward, Wnet represents the net work of the system, k is a proportionality coefficient, p represents the pressure, T represents the temperature, the subscript t represents the turbine, in represents the inlet, and max represents the maximum allowable value; If the action causes the system state to exceed the preset safety limit, a negative reward, i.e., a penalty, is obtained, and the current training segment is stopped and a new training segment is restarted; The temperature and pressure of the working fluid at the inlet of the expander in the waste heat recovery system and the temperature and flow rate of the heat source entering the waste heat recovery system are set as observation quantities of the intelligent agent; the intelligent agent makes a decision action by observing the system state, and the action is to control the working fluid flow rate by directly or indirectly adjusting the speed of the pump, so as to achieve the optimal temperature and pressure state of the system; The agent evaluates the quality of its actions based on the rewards it receives, thereby continuously interacting with the environment and converging towards a direction with greater rewards. Ultimately, an agent controller is trained to obtain the maximum reward, that is, the maximum accumulated output work. Step 3: After setting the boundary conditions of the training environment, train the intelligent agent controller according to the preset training algorithm; After training, several sets of untrained heat source boundary conditions that fluctuate over time and change within the same amplitude range as the heat source fluctuation conditions during training are input into the Rankine cycle waste heat recovery system for testing; When the system cumulative output work of the test results is less than the system cumulative output work in the actual operation data according to the traditional PID constant temperature or constant pressure control method, return to step three, reset the boundary conditions of the training environment and the training algorithm parameters, and continue training until the system cumulative output work is greater than the system cumulative output work in the actual operation data according to the traditional PID constant temperature or constant pressure control method. The training is terminated and the obtained intelligent body controller is used for the optimal control of the Rankine cycle waste heat recovery system.
2. The method for optimizing and controlling a Rankine cycle waste heat recovery system based on deep reinforcement learning according to claim 1, characterized in that: The system is a high-precision dynamic simulation model of a Rankine cycle waste heat recovery system, or a real Rankine cycle waste heat recovery system. When the system is a real Rankine cycle waste heat recovery system, a system simulation model based on the real Rankine cycle waste heat recovery system is established before executing step one, and the system simulation model is used to pre-train the reinforcement learning agent for optimization control. The training process is the same as steps one to three above. Only after the pre-training, the agent can use the real Rankine cycle waste heat recovery system as the training environment for the next step.
3. The optimization control method for Rankine cycle waste heat recovery system based on deep reinforcement learning according to claim 1 is characterized in that: The optimization control method is applicable to a subcritical organic Rankine cycle and a transcritical organic Rankine cycle.
4. The method for optimizing and controlling a Rankine cycle waste heat recovery system based on deep reinforcement learning according to claim 1, characterized in that: The training algorithm of the intelligent agent controller includes a deep deterministic policy gradient algorithm and a deep Q network algorithm.
5. The method for optimizing and controlling a Rankine cycle waste heat recovery system based on deep reinforcement learning according to claim 1, characterized in that: The selection of the boundary conditions of the training environment in step three includes the following steps: from the actual operation data of the waste heat recovery system, a set of randomly fluctuating heat source data within the range of frequent fluctuations of the heat source of the waste heat recovery system is selected as the boundary conditions of the training environment, including the flow and temperature data of the heat source.