A Modeling Method for a Dynamic Simulation Model of a Shell-and-Tube Heat Exchanger Based on Reinforcement Learning
By introducing reinforcement learning algorithms into the dynamic simulation model of the partition wall heat exchanger, the accurate values of the characterization parameters are obtained, and the problems of large model errors and slow calculations in the prior art are solved, and a dynamic simulation model with high precision and fast calculation is realized.
Patent Information
- Application Number
- CN202210709700.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-22
- Publication Date
- 2025-06-03
- Estimated Expiration
- 2042-06-22
AI Technical Summary
When establishing a dynamic simulation model of a partition wall heat exchanger in the prior art, there are problems such as large error, slow calculation speed, unexplained model identification, and poor applicability and stability.
Using a reinforcement learning-based method, a dynamic simulation model with high precision and rapid calculation is constructed by establishing a mechanism framework model including multiple characterization parameters and using reinforcement learning algorithms to obtain the accurate values of these parameters.
The accuracy and calculation speed of the dynamic simulation model of the wall heat exchanger are improved, the stability and interpretability of the model are ensured, and the high-precision needs are met in all operating conditions.
Smart Images

Figure CN115081327B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of energy utilization, and particularly relates to a method for modeling a dynamic simulation model of a shell-and-tube heat exchanger based on reinforcement learning. Background Art
[0002] Shell-and-tube heat exchangers are widely used in industrial systems such as energy, power, petroleum, metallurgy, chemical industry, and pharmaceuticals, and play a crucial role in production. Among them, shell-and-tube heat exchangers are the most widely used heat exchangers at present. Heat exchangers have a crucial impact on the steady-state and dynamic performance of the entire system. Therefore, effectively controlling important parameters such as temperature and pressure inside the heat exchanger during the system operation is a necessary condition to ensure the safe and efficient operation of the system. The development of an efficient and accurate heat transfer process control system is often based on a high-precision dynamic simulation model.
[0003] Currently, the methods for establishing the dynamic simulation model of heat exchangers are mainly divided into two types: one is the modeling method based on heat transfer mechanism, and the other is various model identification methods based on experimental data. In the modeling method based on heat transfer mechanism, the general dynamic simulation model of heat exchangers will be simplified to one-dimensional, at most two-dimensional, otherwise it is very difficult to complete the dynamic simulation calculation for control. The model identification method is difficult to ensure that the model always has a high accuracy in an untrained dataset, and sometimes it may even output results with very large deviations. Currently, both of the two main methods for modeling the dynamic simulation model of heat exchangers have obvious disadvantages: the mechanism modeling has large errors and slow calculation speed; the model identification is not interpretable, and its applicability and stability are poor. Therefore, this paper will propose a new modeling method that combines the advantages of the two methods to overcome the existing problems, so as to obtain a model with high precision, high calculation speed, and high stability under all working conditions. Summary of the Invention
[0004] The purpose of the present invention is to overcome the deficiencies in the prior art and propose a method for modeling a dynamic simulation model of a shell-and-tube heat exchanger based on reinforcement learning. This method uses a set of non-repeating and most basic heat transfer mechanism models to establish a mechanism skeleton model including multiple characterization parameters, minimizing the number of model formulas, and at the same time using the reinforcement learning algorithm to obtain the accurate values of these parameters, solving the problem of poor accuracy in the transient process of the dynamic simulation model of the shell-and-tube heat exchanger and improving the calculation speed.
[0005] A method for modeling a dynamic simulation model of a shell-and-tube heat exchanger based on reinforcement learning includes:
[0006] The first step is to establish the skeleton of the dynamic simulation model of the heat exchanger:
[0007] Set the inlet and outlet flow rates and inlet temperature of the heat source, and the inlet and outlet flow rates and inlet temperature of the working fluid as the boundary conditions of the model for the actual physical process of the heat exchanger to be modeled. Derive the heat transfer mechanism model of the heat exchanger to be modeled to obtain a mechanism skeleton model including multiple characteristic parameters, and the characteristic parameters change in real time with the state of the heat exchanger. Among them, the boundary conditions are dynamic data, that is, the boundary conditions involve the time series of parameters. The mechanism skeleton model equation of any side fluid is specifically as follows:
[0008] Where V, ρ, h, p, α, A, T, C respectively represent the volume, mass flow rate, density, temperature, pressure, heat transfer coefficient, heat transfer area, temperature and specific heat of the fluid; the subscripts ave, in, out, w respectively represent the calculated average value, inlet, outlet, wall and fluid;
[0009] The calculated average value refers to the various fluid parameters calculated using the calculated average temperature (i.e., half of the sum of the inlet and outlet temperatures of the fluid) and the fluid pressure; f 1 , f 2 respectively represent the two sides of the fluid; β 1 , β 2 , β 3 , β 4 , β 5 are characteristic parameters;
[0010] Second step, use the algorithm of reinforcement learning to obtain the accurate values of the characteristic parameters:
[0011] Set the outlet temperature and pressure of the working fluid collected as learning samples. Regard the value of the derived characteristic parameter at every moment as a decision variable, that is, the action of the agent in reinforcement learning. The derived mechanism skeleton model of the heat exchanger is regarded as the environment. The key state parameters of the working fluid calculated by the model, that is, the outlet temperature, pressure, average wall temperature and their change rates are regarded as the observation quantities. Use the error between the outlet temperature of the working fluid output by the model and the outlet temperature of the working fluid of the actual heat exchanger collected to construct the reward function. The smaller the error, the larger the value of the reward function. Among them, the learning samples are dynamic experimental data, that is, the time series of temperature or pressure;
[0012] By training the agent with reinforcement learning, make the agent output the optimal action strategy at every moment, that is, the characteristic parameter that minimizes the error between the model output and the actual value, and obtain the corresponding relationship between the input value and the output value. The input value is the key state parameter and error of the heat exchanger, and the output value is the characteristic parameter;
[0013] Step 3: Fit the functional relationship between the key state parameters of the heat exchanger and the corresponding characterization parameters, so that accurate characterization parameters can be output only based on the current key state parameters of the heat exchanger model (without the need to input errors), thereby constructing a high-precision heat exchanger simulation model.
[0014] Further, the characterization parameters include: β 1 represents the derivative of the corrected calculated average density with respect to time, β 2 represents the derivative of the corrected calculated average internal energy with respect to time, β 3 represents the corrected convective heat transfer amount between the cold fluid and the tube wall, β 4 represents the corrected convective heat transfer amount between the hot fluid and the tube wall, β 5 represents the derivative of the corrected average wall temperature with respect to time.
[0015] Further, when the characterization parameters are trained by the method of reinforcement learning, its reward function is composed of the error between the fluid pressure or outlet temperature output by the skeleton model and the fluid pressure or outlet temperature output by the actual heat exchanger. The characteristic of the reward function is that the smaller the error, the larger the value of the reward function.
[0016] Further, the reinforcement learning algorithm adopts a value-based reinforcement learning algorithm or a policy-based reinforcement learning algorithm; the fitting method in Step 3 adopts a polynomial fitting or a neural network fitting method.
[0017] Compared with the prior art, the beneficial effects brought by the technical solution of the present invention are:
[0018] The present invention uses a group of non-repeating and most basic heat transfer mechanism models as the mechanism skeleton model of the new model established by the modeling method, ensuring that the modeling method has the most basic interpretability and stability, while minimizing the number of model formulas and improving the calculation speed; on the other hand, due to the simplification of the mechanism model, some parameters that are difficult to accurately calculate are generated, which are called characterization parameters in the method, and a reinforcement learning algorithm is used to obtain the accurate values of these parameters to meet the requirements of high model precision. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] Figure 1 is the schematic diagram of the modeling method of the present invention;
[0020] Figure 2 shows the schematic diagram of the mechanism model of a typical countercurrent double-pipe heat exchanger;
[0021] Figure 3 shows the schematic diagram of the boundary conditions of the modeling object in the embodiment;
[0022] Figure 4It is a comparison between the calculation results of the modeling method adopted in the embodiment and the experimental data. Detailed implementation manners
[0023] The technical solution of the present invention will be further described in detail below in conjunction with the accompanying drawings and specific embodiments. The specific embodiments described are only used to explain the present invention and are not intended to limit the present invention.
[0024] As Figure 3 shown, in this embodiment, a dynamic simulation model is established for a spiral tube heat exchanger by using a modeling method of a partition wall type heat exchanger dynamic simulation model based on reinforcement learning. The spiral tube heat exchanger is an evaporator in a basic Rankine cycle, and the heat source is hot air, which is introduced into the hot side of the heat exchanger by a fan and discharged into the environment; the cold side is the R245fa working medium, and the liquid working medium is pressurized by a diaphragm pump and then enters the heat exchanger. After being heated, the working medium becomes gaseous and flows out of the heat exchanger and expands in an expansion valve, and then enters a condenser to be cooled into a liquid. In this embodiment, a set of the most basic heat transfer mechanism models that do not repeat, that is, the skeleton model equations (Equation 17) of the modeling method, are used to establish a dynamic simulation model of the heat exchanger, ensuring that the model has the most basic interpretability and stability, and at the same time minimizing the number of model formulas and improving the calculation speed. Since the specific heat of air on the heat source side is much smaller than the specific heat of the working medium, a simple steady-state model is adopted.
[0025] The design of the heat exchanger is shown in Table 1.
[0026] Table 1: Main design parameters of the evaporator in the heat exchanger
[0027]
[0028] As Figure 1 shown, it specifically includes the following steps:
[0029] The first step is to establish the skeleton of the dynamic simulation model of the heat exchanger: Based on a code writing platform such as matlab or python, a mechanism skeleton model is established. The established mechanism skeleton model is a set of non-repeating heat transfer mechanism models derived from basic heat transfer mechanisms, including multiple characterization parameters. Among them, the characterization parameters change in real time with the state of the heat exchanger, including the derivative of the corrected calculated average density with respect to time, the derivative of the corrected calculated average internal energy with respect to time, the corrected convective heat transfer amount between the cold fluid and the tube wall, the corrected convective heat transfer amount between the hot and cold fluids and the tube wall, and the derivative of the corrected average wall temperature with respect to time.
[0030] Figure 2The following is a schematic diagram of the mechanism model of a typical countercurrent shell-and-tube heat exchanger. The heat exchange section is divided into many control volumes. The mass and energy conservation equations of the cold and hot fluids in each control volume are the same, as shown in equations (1) and (2). There is no mass conservation equation on the wall, only the energy conservation equation (3). In fact, if all types of inter-wall heat exchangers are divided into N control volumes along the flow direction, each control volume also follows the following equation.
[0031]
[0032]
[0033]
[0034] The set of equations (1)-(3) in each control volume is the core mechanism of the traditional mechanism modeling of the interlayer heat exchanger. The traditional mechanism modeling method is to establish N sets of such equations in N control volumes, and then assign boundary conditions to solve them simultaneously.
[0035] Since the cold fluid and the hot fluid (f 1 and f 2 ) has the same energy and mass conservation equations. 1 As an example, the model skeleton is derived. If the mass conservation equations of all the fluids in the control body are summed, we can get:
[0036]
[0037] Further simplifying it, we can get:
[0038]
[0039]
[0040] is the actual average density, V f1 is the total volume of the fluid. However, the actual average density is very difficult to calculate, so the density corresponding to the arithmetic mean temperature (the sum of the inlet and outlet temperatures divided by 2) and the fluid pressure is used as the average density, and all the parameters obtained by the arithmetic mean temperature and pressure are called calculated average parameters, represented by the subscript ave, such as the calculated average temperature T f1_ave and calculate the average density ρ f1_ave Because there is an error between the calculated average density and the actual average density, the derivative of the average density with respect to time in equation (4) will naturally have an error with the derivative of the actual average density. Here we use β 1 Correct this error, that is:
[0041]
[0042] Similarly, summing up the energy conservation equation of the fluid gives:
[0043]
[0044] Further simplifying it gives:
[0045]
[0046] Where:
[0047]
[0048] u f1 is the true average internal energy per unit mass, is the true total internal energy, whether or m f1 is very difficult to calculate accurately. And the rate of change of the total internal energy of the fluid calculated using the calculated average parameters with respect to time, and the error between the rate of change of the true total internal energy with respect to time can be represented by a coefficient β 2 as:
[0049]
[0050] If we let:
[0051]
[0052] Then we get
[0053]
[0054] β 2 The physical meaning of can be understood as the comprehensive correction of the convective heat transfer coefficient between the cold fluid f 1 and the tube wall, and the arithmetic average temperature calculated using the calculated average temperature, that is, the correction of the convective heat transfer amount of the cold fluid f 1 Similarly, summing up the energy conservation equation of the wall gives:
[0055]
[0056] The specific heat and density of the tube wall change very little with temperature and can be considered to be equal to the average specific heat and density, so
[0057]
[0058] Using the same method as above to further simplify the right side of the equation gives:
[0059]
[0060] where β 4 can be understood physically as a comprehensive correction of the convective heat transfer coefficient between the hot fluid and the tube wall and the arithmetic mean temperature calculated using the calculated average temperature, that is, a correction of the convective heat transfer amount of the hot fluid. β 5 can be understood physically as a correction of the time derivative of the average wall temperature.
[0061] In this way, we can use the system of equations (17) composed of equations (7), (13), and (16) to describe the skeletal model equation for a certain fluid side, specifically:
[0062]
[0063] where V, ρ, h, p, α, A, T, C respectively represent the volume, mass flow rate, density, temperature, pressure, heat transfer coefficient, heat transfer area, temperature, and specific heat of the fluid; the subscripts ave, in, out, and w respectively represent the calculated average value, inlet, outlet, and wall and fluid; β 1 , β 2 , β 3 , β 4 , β 5 are characteristic parameters that change in real time with the state of the heat exchanger, and are respectively the derivative of the corrected calculated average density with respect to time, the derivative of the corrected calculated average internal energy with respect to time, the corrected convective heat transfer amount between the cold fluid and the tube wall, the corrected convective heat transfer amount between the hot fluid and the tube wall, and the derivative of the corrected average wall temperature with respect to time;
[0064] The calculated average value refers to the various fluid parameters calculated using the calculated average temperature (i.e., half of the sum of the inlet and outlet temperatures of the fluid) and the fluid pressure; f 1 , f 2 respectively represent the fluids on both sides.
[0065] The accuracy of the model will mainly be determined by the 5 characteristic parameters β 1 , β 2 , β 3 , β 4 , β 5 that are difficult to accurately calculate by mechanism. The accurate values of the characteristic parameters are obtained through the following reinforcement learning algorithm.
[0066] Second step, use the reinforcement learning algorithm to obtain the accurate values of the characteristic parameters:
[0067] In order to learn these five characteristic parameters from the actual experimental data of the heat exchanger, the rotational speed of the pump is changed, and then the inlet and outlet flow rates and the inlet temperature of the heat source, as well as the inlet and outlet flow rates and the inlet temperature of the working fluid, are collected as the boundary conditions of the model. The boundary conditions are dynamic experimental data, that is, the time series of the above parameters. The outlet temperature and pressure of the working fluid are collected as the learning samples, and the learning samples are dynamic experimental data, that is, the time series of temperature or pressure. The hot fluid is high-temperature air, and the inlet temperature and flow rate are directly given; since the outlet is connected to the atmospheric environment, it is assumed that the inlet and outlet flow rates are basically the same.
[0068] The important parameters of the reinforcement learning training algorithm for the said characteristic parameters are shown in Table 2.
[0069] Table 2: Important Parameters of the Training Algorithm
[0070]
[0071]
[0072] The key elements of reinforcement learning are the agent, the environment, the state, the action, and the reward. Regarding the values of these five characteristic parameters β 1 , β 2 , β 3 , β 4 , β 5 at every moment as a decision variable, that is, regarding these five parameters as the actions of the agent; regarding the skeleton model of the heat exchanger as the environment; regarding the outlet temperature, pressure, average wall temperature of the fluids on both sides of the heat exchanger model and their change rates as the observation quantities; using the deviation between the outlet temperature of the working fluid output by the model and the outlet temperature of the working fluid of the actual heat exchanger collected to construct the reward function as shown in formula (18). The smaller the deviation, the larger the value of the reward function. In this way, the optimal action strategy of these five parameters at every moment can be obtained through reinforcement learning, and the corresponding relationship between the input value and the output value can be obtained, where the input value is the key state parameters and errors of the heat exchanger, and the output value is the characteristic parameter; in this implementation case, the deep reinforcement learning algorithm is used.
[0073]
[0074] The specific learning process is as follows:
[0075] (1) At each moment, the agent outputs a set of five characteristic parameters, and uses a deep neural network to observe and perceive the key state of the skeleton model, that is, the outlet temperature, pressure, average wall temperature of the working fluid calculated by the model and their change rates, so as to obtain the specific state feature representation.
[0076] (2) Calculate the immediate reward value according to the error between the calculated working fluid outlet pressure value based on the skeleton model and the actual working fluid outlet pressure value of the heat exchanger according to formula (2) to evaluate the value function of each action, select actions to maximize future rewards, and map the current state to corresponding actions;
[0077] (3) The skeleton model responds to this action and obtains the next set of observations. By continuously looping through the above process and continuously updating the mapping relationship between states and actions, the mapping relationship between the state and action that minimizes the error between the model and experimental data can be finally obtained, that is, the variation relationship of the 5 characteristic parameters with the model state.
[0078] Step 3: Fit the functional relationship between the key state parameters of the heat exchanger and the corresponding characteristic parameters:
[0079] When obtaining the characteristic parameters of the heat exchanger in each state through the deep reinforcement learning algorithm, not only the key state parameters of the heat exchanger need to be observed, but also the error needs to be observed. And it is impossible to know the error when actually using the model. Therefore, it is also necessary to fit the functional relationship between the key state parameters of the heat exchanger and the corresponding characteristic parameters. Here, various fitting methods such as polynomial fitting and neural network fitting can be used. In this way, when using the model, accurate characteristic parameters can be output only based on the current key state parameters of the heat exchanger model, making the model output reach a very high accuracy.
[0080] Finally, the comparison between the calculation results of the heat exchanger model using the modeling method of the present invention and the experimental data is as Figure 4 shown. The modeling method of the present invention only uses a set of equations (17) for calculation. Under the hardware configuration conditions: Matlab2019a version, processor Intel Ii7-9700CPU@3.00GHz, the model established by the modeling method of the present invention only takes 16.7s to complete the calculation results as Figure 4 shown, and its calculation results have high accuracy. While the model of the traditional finite volume method (divided into 20 control volumes, that is, 20 sets of equations) requires about 10 minutes of calculation time.
[0081] The above deep reinforcement learning algorithm, deep neural network, and polynomial fitting are well-known algorithms in the art, and the embodiments of the present invention will not elaborate on them. Optionally, the reinforcement learning algorithm can also be a value-based reinforcement learning algorithm or a policy-based reinforcement learning algorithm.
[0082] Although the preferred embodiments of the present invention have been described above in conjunction with the accompanying drawings, the present invention is not limited to the above specific embodiments. The above specific embodiments are merely illustrative and not restrictive. Under the inspiration of the present invention, those of ordinary skill in the art can also make many forms without departing from the spirit of the present invention and the scope protected by the claims. All of these fall within the protection scope of the present invention.
Claims
1. A method for modeling a dynamic simulation model of a shell-and-tube heat exchanger based on reinforcement learning, characterized in that, it includes: The first step is to establish the framework of the heat exchanger dynamic simulation model. Set the inlet and outlet flow rates and inlet temperature of the heat source, and the inlet and outlet flow rates and inlet temperature of the working fluid as the boundary conditions of the model for the actual physical process of the heat exchanger to be modeled. Derive the heat transfer mechanism model of the heat exchanger to be modeled to obtain a mechanism framework model including multiple characterization parameters, and the characterization parameters change in real time with the state of the heat exchanger; among them, the boundary conditions are dynamic data, that is, the boundary conditions involve the time series of parameters. The mechanism framework model equation of any side fluid is specifically: Among them, V, ρ, h, p, A, T, C represent the volume, mass flow rate, density, enthalpy, pressure, heat transfer area, temperature, and specific heat of the fluid, respectively; α represents the heat transfer coefficient between the fluid and the tube wall, and the subscripts ave, in, out, and w represent the calculated average value, inlet, outlet, and wall surface, respectively; The calculated average value refers to various fluid parameters calculated using the calculated average temperature and fluid pressure; f 1 , f 2 respectively represent the fluids on both sides; β 1 , β 2 , β 3 , β 4 , β 5 are characteristic parameters; Among them, β 1 Correction coefficient corresponding to the derivative of the average density with respect to time; β 2 is the correction coefficient for the rate of change of the total internal energy of the fluid corresponding to the average parameter calculation with respect to time; β 3 is the correction parameter for the convective heat transfer amount of the cold fluid f 1 β 4 is the correction parameter for the convective heat transfer amount of the hot fluid f 2 ; β 5 is the correction to the time derivative of the average wall temperature; The second step is to use the reinforcement learning algorithm to obtain the accurate values of the characterization parameters. Collect the outlet temperature and pressure of the working fluid as learning samples. Regard the value of the derived characterization parameter at each moment as a decision variable, that is, the action of the intelligent agent in reinforcement learning. Regard the derived mechanism framework model of the heat exchanger as the environment, and the key state parameters of the working fluid calculated by the model as the observation quantities, that is, the outlet temperature, pressure, average wall temperature and their change rates are used as the observation quantities; use the outlet temperature or pressure of the working fluid output by the model and the error between the outlet temperature or pressure of the working fluid of the actual heat exchanger collected to construct the reward function. The smaller the error, the larger the value of the reward function; among them, the learning samples are dynamic experimental data, that is, the time series of temperature or pressure. By training the intelligent agent with reinforcement learning, make the intelligent agent output the optimal action strategy at each moment, and obtain the corresponding relationship between the input value and the output value, where the input value is the key state parameter and error of the heat exchanger, and the output value is the characterization parameter. The third step: Fit the functional relationship between the key state parameters of the heat exchanger and the corresponding characterization parameters, so that the accurate characterization parameters can be output only according to the current key state parameters of the heat exchanger model, thereby constructing the heat exchanger simulation model.
2. The method for modeling a dynamic simulation model of a shell-and-tube heat exchanger based on reinforcement learning according to claim 1, characterized in that, When the characterization parameter is trained by the reinforcement learning method, its reward function is composed of the error between the enthalpy value or pressure of the fluid outlet output by the framework model and the enthalpy value or pressure of the fluid outlet output by the actual heat exchanger, and the smaller the error, the larger the value of the reward function.
3. The method for modeling a dynamic simulation model of a shell-and-tube heat exchanger based on reinforcement learning according to claim 1, characterized in that, The reinforcement learning algorithm adopts a value-based reinforcement learning algorithm or a policy-based reinforcement learning algorithm; the fitting method in the third step adopts a polynomial fitting or a neural network fitting method.