Model training method, control method and device based on energy storage heat management system
By applying deep reinforcement learning algorithms in energy storage thermal management systems, establishing training environment models and reinforcement learning agent networks, and optimizing control strategies, the problems of energy waste and increased system energy consumption in the existing technology are solved, and more efficient energy management and lower operating costs are achieved.
Patent Information
- Application Number
- CN202311461804.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-03
- Publication Date
- 2025-05-06
AI Technical Summary
The control strategy of existing energy storage thermal management systems cannot be optimized, resulting in waste of energy, increased system energy consumption and increased operating costs. At the same time, the training time of optimization algorithms is long and may lead to actual system dangers.
The algorithm based on deep reinforcement learning is adopted, and by establishing training environment models and reinforcement learning agent networks, optimizing the control strategy of the energy storage thermal management system, using offline data for environmental modeling, reducing the actual system sampling time, and improving the robust performance of the algorithm.
By optimizing control strategies, the energy efficiency of the energy storage thermal management system is improved, the operating costs are reduced, the training time of reinforcement learning agents is reduced, the robustness of the algorithm is enhanced, and the potential dangers of the actual system are avoided.
Smart Images

Figure CN119939237A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of deep reinforcement learning technology, and in particular to a model training method based on a training environment, a model training method based on a reinforcement learning agent network, a control method based on an energy storage thermal management system, an energy storage thermal management device, and a computer storage medium. Background Art
[0002] At present, the main operation strategy of the energy storage thermal management system is: when the ambient temperature and electricity price are low at night, the system is in cold storage mode, and the refrigerator is turned on at maximum power to store as much cold as possible in the cold storage tank; when the ambient temperature and electricity price are high during the day, the system is in cooling mode, and the cold storage tank releases cold. When the cold storage tank is insufficient, the system switches to cooling mode. The control of the system mainly uses traditional control methods such as RBC (Rule-based Control) and PID. The power of the refrigerator will change with the value of the battery temperature and the set point temperature. In the cooling stage, the cooling ratio of the refrigerator and the cold storage tank in the system is a fixed value.
[0003] Currently, traditional control methods cannot obtain the optimal control strategy, resulting in energy waste in the energy storage thermal management system, that is, increased system energy consumption and increased costs.
[0004] Currently, optimization algorithms need to be continuously trained and iterated, but the energy storage thermal management system cannot be sampled frequently due to structural characteristics and other reasons. Therefore, the optimization algorithm requires a lot of time to train, and due to the trial and error mechanism of the optimization algorithm, it may cause damage or danger to the actual system. Summary of the invention
[0005] In order to solve the above technical problems, the present application proposes a model training method based on a training environment, a model training method based on a reinforcement learning agent network, a control method based on an energy storage thermal management system, an energy storage thermal management device and a computer storage medium.
[0006] In order to solve the above technical problems, the present application proposes a model training method based on a training environment, and the model training method includes:
[0007] Acquire system operation data to be trained, wherein the system operation data includes the real system state at the next moment;
[0008] Inputting the control information and environment information in the system operation data into the training environment model to be trained;
[0009] Obtaining the next moment predicted system state output by the training environment model;
[0010] The training environment model is trained according to the actual system state at the next moment and the predicted system state at the next moment.
[0011] The system operation data to be trained is historical operation data of the system and / or simulation operation data.
[0012] Wherein, the control information includes: refrigerator power, water pump power, and / or valve opening;
[0013] The environmental information includes: ambient temperature, load, and / or electricity price;
[0014] The system operation data also includes the real system status at the current moment.
[0015] Wherein, the training of the training environment model according to the real system state at the next moment and the predicted system state at the next moment includes:
[0016] Acquire the loss value of the training environment model according to the real system state at the next moment and the predicted system state at the next moment;
[0017] The training environment model is trained according to the loss value using a support vector regression algorithm or a random forest algorithm.
[0018] In order to solve the above technical problems, the present application also proposes a model training method based on a reinforcement learning agent network, and the model training method includes:
[0019] Use the training environment model to obtain the current system status of the energy storage thermal management system;
[0020] Inputting the current system state into the reinforcement learning agent network to be trained to obtain an action performed according to the current system state;
[0021] Inputting the action into the training environment model to obtain a new system state and a reward for executing the action;
[0022] Training the reinforcement learning agent network according to the goal of maximizing the accumulated reward obtained;
[0023] Wherein, the training environment model is obtained by training using the above-mentioned model training method.
[0024] Wherein, the step of training the reinforcement learning agent network with the goal of maximizing the obtained reward includes:
[0025] Selecting an action to control the energy storage thermal management system according to the current system state through the reinforcement learning agent network, and obtaining a reward value fed back by the energy storage thermal management system after executing the action;
[0026] With the goal of maximizing the cumulative reward, the reinforcement learning agent network is continuously trained according to the reward value at each moment.
[0027] The reward value includes a first reward for guiding the action to obtain optimal system performance and a second reward for preventing the battery from being out of the optimal temperature for a long time;
[0028] The first reward includes: a reward for high system energy efficiency, a penalty for high operating costs, and a penalty for batteries not being at optimal operating temperature;
[0029] The second reward is a cumulative penalty for the battery not being at the optimal operating temperature for a plurality of consecutive moments.
[0030] In order to solve the above technical problems, the present application also proposes a control method based on an energy storage thermal management system, the control method comprising:
[0031] Acquiring the current system state of the energy storage thermal management system;
[0032] Inputting the current system state into a pre-trained reinforcement learning agent network to obtain a system control strategy;
[0033] Controlling the operation of the energy storage thermal management system according to the system control strategy;
[0034] Wherein, the reinforcement learning agent network is trained by the above-mentioned model training method.
[0035] In order to solve the above technical problems, the present application also proposes an energy storage thermal management device, which includes a memory and a processor coupled to the memory; wherein the memory is used to store program data, and the processor is used to execute the program data to implement the model training method and / or control method as described above.
[0036] In order to solve the above technical problems, the present application also proposes a computer storage medium, which is used to store program data. When the program data is executed by a computer, it is used to implement the above model training method and / or control method.
[0037] Compared with the prior art, the beneficial effects of the present application are: the energy storage thermal management device obtains the system operation data to be trained, wherein the system operation data includes the real system state at the next moment; the control information and environmental information in the system operation data are input into the training environment model to be trained; the predicted system state at the next moment output by the training environment model is obtained; the training environment model is trained according to the real system state at the next moment and the predicted system state at the next moment. Through the above-mentioned model training method, by establishing a training environment model based on actual data, it is possible to avoid using a thermodynamic model that cannot reflect all dynamic characteristics, while avoiding the problem of too long sampling time of the actual system, reducing the training time of the reinforcement learning agent, and improving the robustness of the algorithm. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0039] in:
[0040] Figure 1 It is a schematic diagram of the framework of an embodiment of an energy storage thermal management system provided by the present application;
[0041] Figure 2 It is a flowchart of an embodiment of a model training method based on a training environment provided by the present application;
[0042] Figure 3 is a schematic diagram of the training environment modeling process provided by this application;
[0043] Figure 4 is a schematic diagram of the sequential decision-making process of the reinforcement learning agent provided in this application;
[0044] Figure 5 It is a schematic diagram of the DQN neural network structure provided by this application;
[0045] Figure 6 It is a flowchart of an embodiment of a model training method based on a reinforcement learning agent network provided by the present application;
[0046] Figure 7 It is a flow chart of an embodiment of a control method based on an energy storage thermal management system provided by the present application;
[0047] Figure 8 It is a structural schematic diagram of an embodiment of an energy storage thermal management device provided by the present application;
[0048] Fig. 9 It is a structural diagram of an embodiment of a computer storage medium provided by the present application. DETAILED DESCRIPTION
[0049] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.
[0050] The terms "first", "second", "third", "fourth", etc. (if any) in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can, for example, be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units that are clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0051] The object of this application is a liquid-cooled energy storage thermal management system with a cold storage tank. A vertical cold storage tank is used, and the connection method is a two-stage pump in parallel. The specific structural diagram is as follows Figure 1 The main modules of the system include refrigerator, cold storage tank, battery, water pump and diverter valve.
[0052] The energy storage thermal management system targeted by this application has three working modes: 1) cold storage, that is, the refrigerator stores cold in the cold storage tank while supplying cold to the battery; 2) cold release: the cold storage tank releases cold energy to the battery; 3) cooling, that is, the refrigerator supplies cold to the battery alone.
[0053] At present, the operation strategy of the energy storage thermal management system cannot be optimized. The operation strategy of the existing technology is not optimal, mainly manifested in two aspects: ① The switching of the working mode is not optimal. For example, when the ambient temperature is high and the electricity price is high, the cold storage tank has been discharged in advance. At this time, only low-efficiency refrigerators (ambient temperature is inversely proportional to the efficiency of the refrigerator) can be used to cool the system, which will lead to a decrease in the overall energy efficiency of the system and an increase in operating costs; ② The cooling ratio of the refrigerator and the cold storage tank during the cooling stage is not optimal. For example, operating the system according to the existing ratio may result in residual cold storage tanks after an operating cycle, which leads to a waste of system energy, that is, increased system energy consumption and increased costs.
[0054] In addition, since existing optimization algorithms are all based on thermodynamic models, and thermodynamic models are difficult to reflect the full dynamic characteristics of the system, the optimization strategies calculated by thermodynamic models will have many deviations in actual applications. The best training environment for reinforcement learning optimization algorithms is the actual system, and the energy storage thermal management system cannot be sampled frequently due to structural characteristics and other reasons (the parameters of equipment such as refrigerators cannot change frequently, so the time step is usually 1 hour, while reinforcement learning training samples require tens of thousands). Therefore, if you want to achieve good optimization results for reinforcement learning, you need a lot of time for training, and due to the trial and error mechanism of reinforcement learning, it may cause the danger of excessive battery temperature in the actual system.
[0055] In this regard, in order to improve the energy efficiency of the energy storage thermal management system and reduce the system operating costs on the basis of meeting the battery thermal load, this application proposes an energy storage thermal management system strategy optimization method based on the Deep Reinforcement Learning algorithm; at the same time, in view of the problem of low sampling efficiency of reinforcement learning in the actual application of this system, an improved solution for using offline data for environmental modeling is proposed.
[0056] Please refer to Figure 2 , Figure 2 It is a flow chart of an embodiment of a model training method based on a training environment provided in the present application.
[0057] In order to improve the practical application ability of reinforcement learning methods in energy storage thermal management systems, this application needs to use offline data to model the training environment of the RL agent, that is, to build a state transition model of the energy storage thermal management system, that is, a training environment model. The main methods used are RF (random forest), SVR (Support Vector Regression) and other methods. These algorithms can handle multi-dimensional data samples well to achieve the purpose of predictive regression. The performance of different algorithms can be verified by indicators such as mean square error MSE, mean offset error MBE and determination coefficient R_Squared. After completing the modeling of the system, the current system state and the currently executed action are input into the model, and the state of the system at the next moment will be output. In other words, the model can give timely feedback to the outside world, so it can be used as a training environment for the RL agent.
[0058] like Figure 2 As shown, the specific steps are as follows:
[0059] Step S11: Acquire the system operation data to be trained, wherein the system operation data includes the real system state at the next moment.
[0060] In the embodiment of the present application, the present application needs to collect offline data of the energy storage thermal management system at each time step. Since the parameters in the modules such as the refrigerator in the system cannot change frequently, the time step is usually 1 hour, or other time lengths. The data that need to be collected include but are not limited to: ambient temperature, electricity price, load, refrigerator power, water pump power, valve opening, battery temperature, remaining cold storage tank and cold release of the refrigerator, etc.
[0061] The methods for obtaining offline data provided in this application mainly include the following two methods:
[0062] Actual system historical operation data. By adding flow meters and thermometers to the corresponding parts of the energy storage thermal management system and using the battery's BMS (battery management system), a set of required data can be obtained at each time step. The system environment, load, and control quantity will change over time, so as the system runs, a set of time series data can be collected.
[0063] Operation data output by the simulation model. When the energy storage thermal management system has been in operation for a short time or has no corresponding historical data, there will be a problem that the offline data is insufficient to meet the training environment modeling. At this time, the TRNSYS software can be used to build a simulation model to output sufficient offline data.
[0064] During the construction of TRNSYS, it is necessary to select TRNSYS modules according to the parameters of equipment such as refrigerators, cold storage tanks and water pumps in the actual system. It is necessary to ensure that the modules and structures in the simulation model are consistent with the actual system, and the simulation model must be able to switch back and forth between cold storage, cold release and refrigeration modes.
[0065] While building the simulation model, collect offline data of the actual system as much as possible. After the model is built, you can use these actual data to verify the simulation model. When the output results of the TRNSYS simulation model are basically consistent with the actual data, you can increase the simulation time to output sufficient offline data.
[0066] Furthermore, due to the different dimensions of temperature, flow rate, and power in offline data, in order to reduce the difficulty of modeling in the training environment, the collected offline data needs to be normalized:
[0067]
[0068] Among them, X min , X max Respectively represent the minimum and maximum values of the corresponding data, X i and x i Represent the data before and after normalization, respectively.
[0069] The normalized offline data can be divided into a training set and a test set in a ratio of 7:3. The modeling process corresponding to the model training method of the embodiment of the present application is as follows: Figure 3 As described above, after preprocessing the offline data, the offline data is continuously input into subsequent model regression until a final regression model, ie, a training environment model, is obtained.
[0070] Step S12: input the control information and environment information in the system operation data into the training environment model to be trained.
[0071] In the embodiments of the present application, Figure 3As shown, the modeling process of the entire training environment provided by this application mainly consists of three parts, namely, selection of model input and output, model regression and model verification.
[0072] Specifically, the purpose of model construction is to obtain real-time changes in various states within the system when the model inputs relevant control quantities, that is, control information and external environmental information. The control of the system targeted by this application is mainly achieved by adjusting the power of the refrigerator, the opening of each valve and the power of the water pump. At the same time, the ambient temperature, load and electricity price have a decisive influence on the operation strategy of the system, and the current moment values of the battery temperature, the remaining cold capacity of the cold storage tank and the cold release of the refrigerator have a strong influence on the next moment. Therefore, the input of the regression model is selected including but not limited to: ambient temperature, load, electricity price, refrigerator power, water pump power, opening of each valve, battery temperature at the current moment, remaining cold capacity of the cold storage tank at the current moment, and cold release of the refrigerator at the current moment.
[0073] Step S13: Obtain the predicted system state at the next moment output by the training environment model.
[0074] In the embodiment of the present application, the output of the selected regression model includes but is not limited to: the battery temperature at the next moment, the remaining cold capacity of the cold storage tank at the next moment, and the cold capacity released by the refrigeration machine at the next moment.
[0075] Step S14: training the environment model according to the actual system state at the next moment and the predicted system state at the next moment.
[0076] In an embodiment of the present application, in the model regression stage, since the offline data contains multiple dimensions, the present application selects model regression methods such as the RF algorithm and the SVR algorithm that can process data of multiple dimensions.
[0077] Specifically, the RF algorithm is improved from the Bagging ensemble learning method. When facing a regression problem, the RF algorithm averages the prediction results of many decision trees and finally integrates them into the final result. The data and features used by different decision trees are random. The framework of this algorithm is the RandomForestRegressor in the sklearn module library in Python. Since the regression model of this application is used as the training environment of the RL agent, in order to make the model error smaller, the number of decision trees in the RF parameter n_estimators is set to 200, and the maximum depth of the decision tree max_depth is set to 15.
[0078] The SVR algorithm requires that the predicted data be as close to the hyperplane as possible, so that the total deviation of the data from the hyperplane is minimized. The algorithm has a certain error margin. The framework of the algorithm is svm.SVR in the sklearn module library in Python. In order to reduce the impact of model error on this solution, the kernel function is selected as 'rbf', where the penalty factor C of the error term is set to 10 and the kernel function coefficient gamma is set to 0.001.
[0079] In order to make the regression model closer to the actual system, after obtaining the prediction models of various regression methods, it is necessary to compare and analyze their performance. This application uses mean square error MSE, mean offset error MBE and determination coefficient R_Squared to test the performance of the regression model:
[0080]
[0081]
[0082]
[0083] Among them, y i represents the model prediction value, represents the actual value, n represents the number of test set data, Represents the average value of the actual test set data.
[0084] It should be noted that the loss function used in the model training phase can be the same as the loss function used in the above-mentioned model testing phase, which will not be repeated here.
[0085] This application selects the regression model with the best performance based on three performance indicators, which will be used as the training environment for the next RL agent. Since the data is normalized in the data preprocessing stage, the predicted data also needs to be denormalized:
[0086] y i =X min +Y i (X max -X min )
[0087] Among them, y i and Y i Respectively represent the predicted data before and after denormalization.
[0088] In an embodiment of the present application, the energy storage thermal management device obtains the system operation data to be trained, wherein the system operation data includes the real system state at the next moment; inputs the control information and environmental information in the system operation data into the training environment model to be trained; obtains the predicted system state at the next moment output by the training environment model; and trains the training environment model according to the real system state at the next moment and the predicted system state at the next moment. Through the above-mentioned model training method, by establishing a training environment model based on actual data, it is possible to avoid using a thermodynamic model that cannot reflect all dynamic characteristics, while avoiding the problem of too long sampling time of the actual system, reducing the training time of the reinforcement learning agent, and improving the robustness of the algorithm.
[0089] This application aims to solve the training difficulties of reinforcement learning in actual energy storage thermal management systems, and proposes a method of establishing a training environment model using a regression algorithm. In order to prevent the waste of historical offline data of the energy storage thermal management system, this application proposes a way to use historical offline data for training environment modeling. The model established using historical data will be closer to the real system and more conducive to the implementation of the algorithm. In order to increase the amount of data used to establish the training environment model, this application proposes a method of building a TRNSYS simulation model based on actual data and actual parameters. By building a simulation model that is closer to the actual system, it can output sufficient system data under different working conditions.
[0090] This application models the energy storage thermal management system strategy optimization problem as an MDP (Markov Decision Process, including the definition of state, action, state transition and reward), and trains a reinforcement learning agent (RL agent) network of the energy storage thermal management system operation strategy based on the value-based reinforcement learning algorithm DQN (Deep Q Network). In each time step, the cooling ratio of the refrigeration machine and the cold storage tank in the system is output (including control quantities such as the refrigeration machine power, valve opening and water pump power in the system), so as to achieve the goals of maximum system energy efficiency, minimum operating cost and just using up the cold storage tank at the end of an operation cycle.
[0091] like Figure 4As shown in the figure, at each decision moment in the sequential decision-making process of each system operation cycle, RLagent outputs the control such as the refrigerator power and valve opening to be adjusted at that moment based on the observed state of the current energy storage thermal management system, and obtains the reward and the state of the next moment fed back by the system for the action taken under this state. If the cooling capacity output by the current refrigerator and cold storage tank of the system cannot meet the load demand of the battery end in the current time step, a cumulative penalty will be given to RLagent. After the cumulative penalty reaches a certain amount, the energy storage thermal management process will terminate and feedback a negative reward signal for punishing the failure to meet the battery load at the last moment; otherwise, the RL agent will complete the adjustment of various control quantities within the system within an operation cycle, and obtain a reward signal that comprehensively reflects the completion of various optimization goals of the energy storage thermal management system.
[0092] This application is based on the DQN algorithm. The key to DQN is the action value function Q π The calculation of (s,a) represents the average return of taking action a in state s under strategy π, where strategy π(s,a) represents the policy function of selecting different actions in a certain state. As the algorithm iterates, DQN uses the TD (Temporal Difference) algorithm to calculate Q π (s,a) is updated, and Q π (s,a) will also gradually converge to the optimal action value function Q * (s, a) means that the probability of the RL agent choosing the action with the greatest value in this state is 1, that is, the optimal strategy of the system is obtained.
[0093] To use the function Q in high dimensions π (s,a), consider using a neural network to represent it, such as Figure 5 As shown, its input layer is the system state, and the output layer obtained after each hidden layer is the action value function Q π (s,a).
[0094] Please refer to Figure 6 , Figure 6 This is a flow chart of an embodiment of a model training method based on a reinforcement learning agent network provided in this application. The key to the RL agent of this application, that is, the training of the reinforcement learning agent network, is to construct the state space, action space and reward function for the energy storage management system.
[0095] like Figure 6 As shown, the specific steps are as follows:
[0096] Step S21: using the training environment model to obtain the current system state of the energy storage thermal management system.
[0097] In the embodiment of the present application, the state s obtained by the RL agent at each moment in the decision-making process of the energy storage thermal management system t It consists of four parts: time (the system environment will change with time, and the time variable can be used to characterize the changes in three state quantities: load, ambient temperature and load), battery temperature (characterizing the battery state), remaining cold capacity of the cold storage tank (characterizing the cold storage tank state), and cold release capacity of the refrigeration machine (characterizing the refrigeration machine state).
[0098] Step S22: Input the current system state into the reinforcement learning agent network to be trained to obtain the action performed according to the current system state.
[0099] In the embodiment of the present application, the action consists of three control modes: refrigerator power, valve opening, and water pump power of the energy storage thermal management system. In order to improve the performance of the energy storage thermal management system and the training efficiency of the RL agent, it is necessary to discretize the action space: for refrigerator power, its minimum value is 0 and its maximum value is the refrigerator rated power, and map it to the n-divided action space C = {c 1 ,c 2 ,...,c n}; The pump power can be mapped to the action space W={w 1 ,w 2 ,...,w l}; The action value is 0 when the valve is closed, and 1 when the valve is fully open, which is mapped to the m-divided action space V = {v 1 ,v 2 ,...,v m}. So the entire action space is composed of A = {(c i ,w j ,v k )|c i ∈C,w j ∈W,v k ∈V}.
[0100] According to the above action space, the energy storage thermal management device inputs the current system state of the energy storage thermal management system into the reinforcement learning agent network to obtain the action performed according to the current system state.
[0101] Step S23: Input the action into the training environment model to obtain the new system state and the reward for performing the action.
[0102] In the embodiments of the present application, Figure 4 As shown in the figure, the training environment model generates a new system state according to the action decided by the reinforcement learning agent network, and feeds the new system state back to the reinforcement learning agent network. The reinforcement learning agent network can then calculate the reward of the action based on the new system state.
[0103] In the decision-making optimization process of the energy storage thermal management system, the reward r received by the RL agent for completing the current control at each moment is t Contains the reward r_a used to guide it to achieve the optimal performance of the system t and the bonus of preventing the battery from being at optimal temperature for a long time t .
[0104] Specifically, the reward r_a t (Feedback at every moment): The RL agent needs to maintain the battery at the optimal operating temperature (15-35°C) while making the system energy efficiency as high as possible and the operating cost as low as possible, so the reward is set to r_a t :
[0105]
[0106]
[0107] The first two items represent the reward for high system energy efficiency and the penalty for high operating costs, respectively. b Indicates system load, P c and P w represents the electricity price at that moment; F(T) represents the penalty for the battery not being at the optimal operating temperature, T represents the battery temperature, T l and T u They represent the upper and lower bounds of the optimal operating temperature of the battery respectively, and a is a constant.
[0108] In order to balance the mutually exclusive relationship among energy efficiency, cost and battery temperature, coefficients α and β are added to adjust their weights.
[0109] Reward r_b t (Feedback is performed at the termination time T): The weighted sum that reflects the degree of satisfaction of the battery operating temperature.
[0110] Step S24: Train the reinforcement learning agent network with the goal of maximizing the accumulated reward obtained.
[0111] In the embodiment of the present application, due to the characteristics of lithium batteries, if they are not at the optimal operating temperature for a long time, their battery life will be greatly reduced and there is a risk of causing a fire. In order to prevent danger in actual application, it is necessary to weight the sum of the battery temperature satisfaction at each moment and set a penalty margin. When the cumulative penalty exceeds the margin, the maximum cooling capacity should be used immediately and the AL agent training should be terminated.
[0112]
[0113] In summary, through the construction of the above reward function, the decision-making optimization process of the energy storage thermal management system is transformed into a reward value for evaluating different control actions. The optimal reward function is obtained by comparing and combining different control actions, that is, the action with the highest reward value is used as the decision-making optimization solution for the energy storage thermal management system.
[0114] In order to improve the energy efficiency of the energy storage thermal management system and reduce operating costs, this application creatively proposes an operation strategy optimization method based on reinforcement learning, which realizes the optimal switching of the system working mode and achieves the optimal cooling ratio of the refrigeration machine and the cold storage tank in the cooling mode.
[0115] Please continue reading Figure 7 , Figure 7 It is a flow chart of an embodiment of a control method based on an energy storage thermal management system provided in the present application.
[0116] like Figure 7 As shown, the specific steps are as follows:
[0117] Step S31: Acquire the current system state of the energy storage thermal management system.
[0118] Step S32: Input the current system state into the pre-trained reinforcement learning agent network to obtain the system control strategy.
[0119] In the embodiment of the present application, the reinforcement learning agent network is trained by the model training method in the above embodiment, and the process will not be repeated here.
[0120] Step S33: Control the operation of the energy storage thermal management system according to the system control strategy.
[0121] Those skilled in the art will appreciate that, in the above method of specific implementation, the order in which the steps are written does not imply a strict execution order and does not constitute any limitation on the implementation process. The specific execution order of the steps should be determined by their functions and possible internal logic.
[0122] In order to implement the above-mentioned model training method and / or control method, this application also proposes an energy storage thermal management device, please refer to Figure 8 , Figure 8 It is a structural schematic diagram of an embodiment of an energy storage thermal management device provided in the present application.
[0123] The energy storage thermal management device 400 of this embodiment includes a processor 41 , a memory 42 , an input / output device 43 , and a bus 44 .
[0124] The processor 41, memory 42, and input / output device 43 are respectively connected to a bus 44. The memory 42 stores program data, and the processor 41 is used to execute the program data to implement the model training method and / or control method described in the above embodiments.
[0125] In the embodiment of the present application, the processor 41 may also be referred to as a CPU (Central Processing Unit). The processor 41 may be an integrated circuit chip having the ability to process signals. The processor 41 may also be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA), or other programmable logic devices, discrete gates or transistor logic devices, or discrete hardware components. A general-purpose processor may be a microprocessor or the processor 41 may also be any conventional processor, etc.
[0126] This application also provides a computer storage medium, please continue to refer to Fig. 9 , Fig. 9 It is a structural diagram of an embodiment of a computer storage medium provided in the present application. The computer storage medium 600 stores a computer program 61. When the computer program 61 is executed by the processor, it is used to implement the model training method and / or control method of the above-mentioned embodiment.
[0127] When the embodiments of the present application are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor to perform all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a disk or an optical disk.
[0128] The above description is only an implementation method of the present application, and does not limit the patent scope of the present application. Any equivalent structure or equivalent process transformation made using the contents of the present application specification and drawings, or directly or indirectly used in other related technical fields, are also included in the patent protection scope of the present application.
Claims
1. A model training method based on a training environment, characterized in that: The model training method comprises: Acquire system operation data to be trained, wherein the system operation data includes the real system state at the next moment; Inputting the control information and environment information in the system operation data into the training environment model to be trained; Obtaining the next moment predicted system state output by the training environment model; The training environment model is trained according to the actual system state at the next moment and the predicted system state at the next moment.
2. The model training method according to claim 1, characterized in that: The system operation data to be trained is historical operation data of the system and / or simulation operation data.
3. The model training method according to claim 1 or 2, characterized in that: The control information includes: refrigerator power, water pump power, and / or valve opening; The environmental information includes: ambient temperature, load, and / or electricity price; The system operation data also includes the real system status at the current moment.
4. The model training method according to claim 1, characterized in that: The training of the training environment model according to the real system state at the next moment and the predicted system state at the next moment comprises: Acquire the loss value of the training environment model according to the real system state at the next moment and the predicted system state at the next moment; The training environment model is trained according to the loss value using a support vector regression algorithm or a random forest algorithm.
5. A model training method based on a reinforcement learning agent network, characterized in that: The model training method comprises: Use the training environment model to obtain the current system status of the energy storage thermal management system; Inputting the current system state into the reinforcement learning agent network to be trained to obtain an action performed according to the current system state; Inputting the action into the training environment model to obtain a new system state and a reward for executing the action; Training the reinforcement learning agent network according to the goal of maximizing the accumulated reward obtained; Wherein, the training environment model is obtained by training using the model training method described in any one of claims 1 to 4.
6. The model training method according to claim 5, characterized in that: The step of training the reinforcement learning agent network with the goal of maximizing the obtained reward includes: Selecting an action to control the energy storage thermal management system according to the current system state through the reinforcement learning agent network, and obtaining a reward value fed back by the energy storage thermal management system after executing the action; With the goal of maximizing the cumulative reward, the reinforcement learning agent network is continuously trained according to the reward value at each moment.
7. The model training method according to claim 6, characterized in that: The reward value includes a first reward for guiding the action to obtain optimal system performance and a second reward for preventing the battery from being at an optimal temperature for a long time; The first reward includes: a reward for high system energy efficiency, a penalty for high operating costs, and a penalty for batteries not being at optimal operating temperature; The second reward is a cumulative penalty for the battery not being at the optimal operating temperature for a plurality of consecutive moments.
8. A control method based on an energy storage thermal management system, characterized in that: The control method comprises: Acquiring the current system state of the energy storage thermal management system; Inputting the current system state into a pre-trained reinforcement learning agent network to obtain a system control strategy; Controlling the operation of the energy storage thermal management system according to the system control strategy; Wherein, the reinforcement learning agent network is trained by the model training method described in any one of claims 5 to 7.
9. An energy storage thermal management device, characterized in that: The energy storage thermal management device includes a memory and a processor coupled to the memory; Wherein, the memory is used to store program data, and the processor is used to execute the program data to implement the model training method as described in any one of claims 1 to 7, and / or the control method as described in claim 8.
10. A computer storage medium, characterized in that: The computer storage medium is used to store program data, and when the program data is executed by a computer, it is used to implement the model training method as described in any one of claims 1 to 7, and / or the control method as described in claim 8.