Energy supply action intelligent control calculation method, device and equipment and storage medium
The DDPG algorithm addresses the inefficiencies in 5GDHC systems by optimizing energy supply actions, improving energy efficiency and stability through real-time adaptation and precise control adjustments.
Patent Information
- Application Number
- CN202510360062.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-25
- Publication Date
- 2025-07-15
AI Technical Summary
The existing 5GDHC system control methods have problems such as insufficient adaptability, low control accuracy and slow response in terms of multi-source integration, bidirectional energy flow and load staggering, which are difficult to meet the needs of efficient and stable operation.
A continuous control algorithm of deep deterministic strategy gradient (DDPG) is introduced, and by building an energy hub system, detecting and predicting loads and temperatures, and optimizing control actions using the ε-greedy method, it realizes refined management of the 5GDHC system.
It improves the operating efficiency and stability of the 5GDHC system, reduces the total energy consumption by about 10%, and reduces the proportion of time that end load does not meet the needs from more than 6% to less than 2%, improving the system's adaptability and energy efficiency.
Smart Images

Figure CN120317418A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of servers, and particularly to an intelligent control calculation method, device, computer equipment and storage medium for energy supply actions. Background Art
[0002] The Fifth-Generation District Heating and Cooling (5GDHC) system has characteristics such as low-temperature operation, energy sharing, decentralization, and two-way energy flow. Different from traditional district heating and cooling, 5GDHC realizes two-way circulation of heating and cooling through a low-temperature hot water network, avoiding heat loss and equipment durability problems of traditional high-temperature pipe networks. The system adopts a decentralized distributed energy station, such as heat pumps, refrigeration units, etc., which are connected to the interior of user buildings. Each user can be both a consumer and a producer of thermal energy, that is, the "prosumer" model. The existing fifth-generation district heating and cooling system has great flexibility. Existing control methods have their own advantages and limitations, and cannot fully unleash the system characteristics and energy-saving features of multi-source integration, load peak shaving, ultra-low water temperature, and two-way energy flow. It is necessary to introduce intelligent control algorithms to promote the overall improvement of the 5GDHC system.
[0003] First of all, existing control methods such as rule-based control, temperature control, and model predictive control show problems of insufficient adaptability, poor flexibility, and low refinement when dealing with multi-source integration and two-way energy flow of the 5GDHC system. These control strategies cannot make full use of real-time data and environmental variables, resulting in low energy utilization efficiency of the 5GDHC system and difficulty in fully exerting the energy-saving characteristics of the system.
[0004] Secondly, the decision-making ability of existing control methods in a high-dimensional continuous action space is limited. Especially for reinforcement learning control strategies based on value function methods such as DDQN, it is easy to have a loss of control accuracy due to the discretized action space, and cannot meet the refined requirements of the 5GDHC system for operations such as starting and stopping heat pumps and adjusting cold and hot flow rates. When facing large changes in end loads, the response is slow, and problems of over-energy supply or insufficient energy supply are likely to occur. Summary of the Invention
[0005] Based on this, an intelligent control calculation method, device, computer equipment and storage medium for energy supply actions are provided. By introducing a continuous control algorithm of Deep Deterministic Policy Gradient (DDPG), the limitations of existing control methods are overcome, thereby improving the operation efficiency and stability of the 5GDHC system.
[0006] On the one hand, an intelligent control calculation method for energy supply actions is provided. The method includes:
[0007] Construct an energy hub system through a 5GDHC model. The components in the energy hub system include a heat pump unit, a warm pipe and a cold pipe, a building load on the end-user side, a heat pump unit on the end-user side, and a water pump on the end-user side;
[0008] Detect and obtain the hourly heating and cooling loads of the building on the end-user side of the energy hub system, the hourly load provided by the heat pump unit, the heating energy efficiency ratio of the energy hub system, the average temperature of the warm pipe, and the average temperature of the cold pipe to form a state vector at the current moment;
[0009] Select the set outlet water temperature of the heat pump unit on the end-user side and the frequency of the water pump on the end-user side as the control action at the current moment according to the state vector at the current moment;
[0010] Predict the state vector and the reward at the next moment according to the control action at the current moment;
[0011] Input the reward at the next moment into the deep deterministic policy gradient model and compare it with the expected state of the energy hub system at the next moment. Select an action based on the ε-greedy method to correct the control action at the current moment, and execute the corrected control action at the current moment.
[0012] In one embodiment, the detecting and obtaining the hourly heating and cooling loads of the building on the end-user side of the energy hub system, the hourly load provided by the heat pump unit, the heating energy efficiency ratio of the energy hub system, the average temperature of the warm pipe, and the average temperature of the cold pipe to form a state vector at the current moment includes:
[0013] Set the current moment as represented by moment t. The formula for the state vector at the current moment is
[0014] In the formula: is the hourly heating and cooling load (kW) of the jth building on the end-user side at moment t;
[0015] is the hourly load (kW) provided by the jth heat pump unit on the end-user side at moment t;
[0016] COP sys (t) is the heating energy efficiency ratio of the energy hub system at moment t;
[0017] is the average temperature (°C) of the warm pipe at moment t;
[0018] is the average temperature (°C) of the cold pipe at moment t.
[0019] In one embodiment, selecting the set outlet water temperature of the heat pump unit on the end-user side and the frequency of the pump on the end-user side as the control action at the current moment according to the current moment state vector includes:
[0020] Set the current moment as represented by moment t, and the formula for the control action at the current moment is
[0021] where: f HP,j (t) is the hourly frequency (Hz) of the j-th pump at moment t;
[0022] is the set outlet water temperature (°C) of the j-th heat pump at moment t.
[0023] In one embodiment, predicting the state vector and the reward at the next moment according to the control action at the current moment includes:
[0024] Predict the state vector at the next moment according to the control action at the current moment. Set the next moment as represented by moment t + 1, and the state vector at the next moment is
[0025] where: is the hourly heating and cooling load (kW) of the j-th end-user side building at moment t + 1;
[0026] is the hourly load (kW) provided by the j-th heat pump unit on the end-user side at moment t + 1;
[0027] COP sys (t + 1) is the heating energy efficiency ratio of the energy hub system at moment t + 1;
[0028] is the average temperature (°C) of the warm pipe at moment t + 1;
[0029] is the average temperature (°C) of the cold pipe at moment t + 1;
[0030] Obtain the reward at the next moment according to the state vector at the next moment. The formula for the reward at the next moment is
[0031] where: E(t + 1) is the comprehensive energy consumption (kWh) of all heat pumps and pumps within moment t + 1;
[0032] It is indicated that if the difference between the hourly cooling and heating load of the j-th building and the hourly load provided by the j-th heat pump at time t+1 exceeds the deviation threshold ε, a penalty is imposed; (Preferably, the value of the deviation threshold ε is 2%-5% of the value;)
[0033] Δ action is the action smoothing term, which characterizes the magnitude of the difference from the previous action and is used to penalize frequent adjustment and startup / shutdown of equipment. Define Δ action =||a t -a t-1 || 2 ;
[0034] λ load is the penalty coefficient imposed when the hourly load provided by the heat pump unit on the end-user side exceeds the limit;
[0035] λ action is the penalty coefficient imposed when the action smoothing term provided by the heat pump unit on the end-user side exceeds the limit.
[0036] In one embodiment, before inputting the next moment reward into the Deep Deterministic Policy Gradient Model (DDPG) and comparing it with the expected state of the energy hub system at the next moment, it further includes:
[0037] An experience replay pool is set in the Deep Deterministic Policy Gradient Model to store historical transition data, and the historical transition data includes the state vector, control action, and reward at each moment in the energy hub system;
[0038] The Deep Deterministic Policy Gradient Model is trained using the historical transition data in the experience replay pool until the training converges or reaches a preset number of rounds to obtain the trained Deep Deterministic Policy Gradient Model.
[0039] In one embodiment, the training of the Deep Deterministic Policy Gradient Model using the historical transition data in the experience replay pool includes:
[0040] It is set that the Deep Deterministic Policy Gradient Model includes an online actor network, an online critic network, a target actor network, and a target critic network;
[0041] Set the online actor network as μ(s|θ μ ), set the online critic network as Q(s,μ(s|θ μ )|θ Q ), where s is the state vector, θ μ is the first network weight, and θ Q is the second network weight;
[0042] Set the target actor network representation as μ'(s|θ μ' ), and set the target critic network as Q'(s, μ'(s|θ μ' )|θ Q' ), where θ μ' is the third network weight and θ Q' is the fourth network weight;
[0043] Initialize the values of the first network weight and the second network weight, copy the value of the first network weight to the third network weight, and copy the value of the second network weight to the fourth network weight;
[0044] Input the simulation state vector of the energy hub system into the online actor network, add exploration noise in the online actor network to form a simulation control action, execute the simulation control action to simulate the next moment to obtain the next moment simulation state vector and the next moment simulation reward, and store the simulation state vector, the simulation control action, the next moment simulation state vector, and the next moment simulation reward in the experience replay pool to accumulate data samples;
[0045] Randomly sample a batch of N sampling data {s i , a i , r i , s i+1} from the historical transfer data in the experience replay pool to update the first network weight and the second network weight, where s i is the state vector at time i, a i is the control action at time i, r i is the reward at time i, and s i+1 is the state vector at time i+1;
[0046] When training the online critic network, the target Q value of the online critic network is y i = r i + γ·Q'(s i+1 , μ'(s i+1 |θ μ′ )|θ Q′ ), where γ is the discount factor, and the loss function is Update the second network weight θ Q by minimizing this loss through gradient descent;
[0047] When training the online actor network, the policy gradient of the online actor network is The goal of the online actor network is to make the selected action maximize the target Q value of the online critic network, and the update gradient is the average value of the policy gradient Taking the partial derivative of the input action in the online critic network and sending the updated gradient back to the online actor network to update the first network weight θ μ ;
[0048] After updating the online actor network, the third network weight of the target actor network is updated by θ Q′ ←τθ Q +(1 - τ)θ Q′ ; After updating the online critic network, the fourth network weight of the target critic network is updated by θ μ′ ←τθ μ +(1 - τ)θ μ′ ; where τ << 1.
[0049] In one embodiment, comparing the next - moment reward input into the Deep Deterministic Policy Gradient (DDPG) model with the expected state of the energy hub system at the next moment, and selecting an action based on the ε - greedy method to correct the current - moment control action, the execution of the corrected current - moment control action includes:
[0050] In response to the trained Deep Deterministic Policy Gradient model receiving the next - moment reward, obtaining the current - moment state vector, the current - moment control action, the current - moment reward, and the next - moment state vector;
[0051] Judging whether there is an offset between the next - moment reward and the expected state of the energy hub system at the next moment;
[0052] If there is an offset, adjusting the current - moment control action based on the ε - greedy method to make the next - moment reward the same as the expected state of the energy hub system at the next moment, and executing the corrected current - moment control action;
[0053] If there is no offset, executing the current - moment control action;
[0054] Obtaining the next - moment reward and the next - moment state vector after executing the control action, and storing the current - moment state vector, the current - moment control action, the next - moment reward, and the next - moment state vector in the experience replay pool.
[0055] On the other hand, an intelligent control calculation device for energy supply actions is provided. The device includes:
[0056] A model construction module for constructing an energy hub system through a 5GDHC model. The components in the energy hub system include a heat pump unit, a heating pipe and a cooling pipe, a building load on the end - user side, a heat pump unit on the end - user side, and a water pump on the end - user side;
[0057] A state vector acquisition module, which is used to detect and acquire the hourly heating and cooling loads of the building on the end-user side of the energy hub system, the hourly load provided by the heat pump unit, the heating energy efficiency ratio of the energy hub system, the average temperature of the heating pipe and the average temperature of the cooling pipe, and form a state vector at the current moment;
[0058] A control action generation module, which is used to select the set outlet water temperature of the heat pump unit on the end-user side and the frequency of the water pump on the end-user side as the control action at the current moment according to the state vector at the current moment;
[0059] An anticipation module, which is used to anticipate the state vector and the reward at the next moment according to the control action at the current moment;
[0060] A control action execution module, which is used to input the reward at the next moment into the deep deterministic policy gradient model for comparison with the anticipated state of the energy hub system at the next moment, select an action based on the ε-greedy method to correct the control action at the current moment, and execute the corrected control action at the current moment.
[0061] On the other hand, a computer device is provided, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the following steps are implemented:
[0062] Construct an energy hub system through a 5GDHC model. The components in the energy hub system include a heat pump unit, heating pipes and cooling pipes, building loads on the end-user side of the energy hub, heat pump units on the end-user side, and water pumps on the end-user side;
[0063] Detect and acquire the hourly heating and cooling loads of the building on the end-user side of the energy hub system, the hourly load provided by the heat pump unit, the heating energy efficiency ratio of the energy hub system, the average temperature of the heating pipe and the average temperature of the cooling pipe, and form a state vector at the current moment;
[0064] Select the set outlet water temperature of the heat pump unit on the end-user side and the frequency of the water pump on the end-user side as the control action at the current moment according to the state vector at the current moment;
[0065] Anticipate the state vector and the reward at the next moment according to the control action at the current moment;
[0066] Input the reward at the next moment into the deep deterministic policy gradient model for comparison with the anticipated state of the energy hub system at the next moment, select an action based on the ε-greedy method to correct the control action at the current moment, and execute the corrected control action at the current moment.
[0067] On yet another hand, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented:
[0068] Construct an energy hub system through the 5GDHC model. The components in the energy hub system include heat pump units, heating pipes and cooling pipes, building loads on the end-user side, heat pump units on the end-user side, and water pumps on the end-user side.
[0069] Detect and obtain the hourly cooling and heating loads of the building on the end-user side of the energy hub system, the hourly loads provided by the heat pump units, the heating energy efficiency ratio of the energy hub system, the average temperature of the heating pipes, and the average temperature of the cooling pipes to form the state vector at the current moment.
[0070] Select the set outlet water temperature of the heat pump unit on the end-user side and the frequency of the water pump on the end-user side as the control action at the current moment according to the state vector at the current moment.
[0071] Predict the state vector and the reward at the next moment according to the control action at the current moment.
[0072] Input the reward at the next moment into the Deep Deterministic Policy Gradient (DDPG) model and compare it with the expected state of the energy hub system at the next moment. Select an action based on the ε-greedy method to correct the control action at the current moment, and execute the corrected control action at the current moment.
[0073] The above intelligent control calculation method, device, computer equipment, and storage medium for energy supply actions can, by introducing the continuous control algorithm of Deep Deterministic Policy Gradient (DDPG), compare the reward at the next moment predicted according to the control action at the current moment with the expected state of the energy hub system at the next moment, select an action based on the ε-greedy method to correct the control action at the current moment, achieve the cooling and heating of the energy hub system to be close to the expected result, improve the accuracy of the energy hub system, and overcome the limitations of the existing control methods, thereby improving the operation efficiency and stability of the 5GDHC system. Brief Description of the Drawings
[0074] To more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0075] Figure 1 It is an application environment diagram of the intelligent control calculation method for energy supply actions in an embodiment of this application.
[0076] Figure 2It is a structural block diagram of an energy hub system constructed by a 5GDHC model under the TRNSYS software environment in an embodiment of the present application;
[0077] Figure 3 It is a framework diagram of a Markov decision process for optimizing the operation of a 5GDHC model in an embodiment of the present application;
[0078] Figure 4 It is a flowchart of an intelligent control calculation method for energy supply actions in an embodiment of the present application;
[0079] Figure 5 It is a control framework diagram of a 5GDHC system based on DDPG in an embodiment of the present application;
[0080] Figure 6 It is a simulation result diagram of the DDPG algorithm based on the TRNSYS software in an embodiment of the present application;
[0081] Figure 7 It is a structural block diagram of an intelligent control calculation device for energy supply actions in an embodiment of the present application;
[0082] Figure 8 It is an internal structural diagram of a computer device in an embodiment of the present application. Detailed implementation manners
[0083] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application, and are not used to limit the present application.
[0084] As described in the background art, since the 5GDHC system has great flexibility, it is necessary to be equipped with an advanced control system to ensure its efficient and stable operation. One of the control methods of 5GDHC is centralized control. Through a central control center, the entire cooling and heating network is uniformly scheduled and managed, and the operating status of each node and the energy flow are monitored in real time. This system relies on high-performance data processing and communication facilities to quickly respond to system load changes, optimize energy distribution, and improve the overall operating efficiency. The advantage of centralized control is that it has a global view and can perform coordinated optimization, but it lacks flexibility in dealing with local failures and system expansion, and the adaptive ability of the control system needs to be improved.
[0085] Another control method for 5GDHC is distributed control. The distributed control system disperses the control functions to each terminal subsystem. Each subsystem is equipped with an independent control unit, which autonomously adjusts the operation parameters according to the local energy consumption demand and environmental conditions, and at the same time exchanges information and collaborates with other terminal subsystems through a communication network. Distributed control has high system robustness and flexibility, and can more effectively cope with local failures and dynamic changes. However, from the system level, its coordination and optimization ability is inferior to that of the centralized control method.
[0086] Generally speaking, the existing fifth-generation district cooling and heating systems have great flexibility. Each existing control method has its own advantages and limitations, and cannot fully unleash the system characteristics and energy-saving features of multi-source integration, load peak shaving, ultra-low water temperature, and two-way energy flow. It is necessary to introduce intelligent control algorithms to promote the overall improvement of the 5GDHC system.
[0087] The existing fifth-generation district cooling and heating systems still face some technical challenges in practical applications, mainly manifested in problems such as insufficient optimization of control strategies and insufficient support of the control system for the overall stable operation.
[0088] (1) Rule-based control. 5GDHC adopts a traditional rule-based control method. The disadvantages are mainly reflected in the lack of adaptability, inflexible response, and difficulty in realizing the dynamic regulation of multiple variables such as water temperature, flow rate, and terminal equipment units. In rule-based control, the control logic usually relies on preset fixed rules. For example, the cold pipes and hot pipes of 5GDHC adjust the operation parameters of the energy station according to the set temperature threshold or according to the terminal user load range. However, 5GDHC adopts a system of multi-source energy supply and two-way energy flow, and it is difficult to use the rule-based control method to cope with the real-time changes in the external environment and the fluctuations in user demands. When the winter temperature drops suddenly or the summer temperature rises suddenly, the system cannot quickly adjust the energy supply according to the demand to stabilize the water temperature of the cold pipes and hot pipes, resulting in the failure to meet the terminal load in some areas. In addition, rule-based control is difficult to handle the influence of non-linear factors, such as the coupling effects of multiple variables such as solar radiation, wind speed, and user behavior, and it is difficult to achieve refined energy management.
[0089] (2) Control based on the water temperature of cold and hot main pipes. The control strategy based on temperature in the 5GDHC system is mainly reflected in aspects such as insufficient flexibility, limited energy efficiency, and difficulty in guaranteeing the terminal user load. The constant temperature control strategy usually adopts fixed temperature set points for the cold pipes and hot pipes, which easily leads to a decrease in the system operation efficiency and energy waste. For example, in a 5GDHC system in Germany that adopts constant temperature control, during the high-temperature period in summer, the cooling demand of the terminal heat pump could not be effectively met because the water supply temperature was not adjusted flexibly. In addition, constant temperature control is difficult to effectively integrate unstable renewable energy sources, such as shallow geothermal energy or industrial waste heat, which limits the system potential of multi-source integration of the 5GDHC system.
[0090] (3) Model Predictive Control. Model Predictive Control (MPC) highly depends on the accuracy of the model. It is difficult to establish a high-precision prediction model for the 5GDHC system that covers various environmental factors, user behaviors, and energy demands. Especially in a multi-source energy network and a two-way energy flow system, prediction errors will directly affect the control effect. Secondly, the computational complexity of MPC is relatively high. Especially in the 5GDHC system, real-time calculation and optimization may lead to control delays, affecting the rapid response ability of the system. In a 5GDHC system pilot project in a European city, the MPC control could not handle sudden load changes in real time, resulting in insufficient cooling in the system under extreme weather conditions. In addition, MPC also requires a large amount of real-time data and high-performance computing hardware facilities, increasing the deployment and maintenance costs.
[0091] (4) Reinforcement Learning Control. Using the Double Deep Q-Network (DDQN) in reinforcement learning for the control of 5GDHC has certain innovation, but there are also obvious deficiencies. First of all, DDQN belongs to the value function method and performs weakly in dealing with continuous action spaces. It needs to discretize the continuous action space before control, and this discretization leads to accuracy loss or the control strategy is not smooth enough. Secondly, DDQN is prone to the curse of dimensionality in high-dimensional state and action spaces and cannot efficiently handle a large amount of real-time data and the system state with multi-variable coupling. For example, in a 5GDHC project, DDQN is used to control the start-stop and load distribution of heat pumps. Due to the discretization of the action space, the system fails to accurately match user demands during the peak load period, resulting in excessive cooling. In addition, DDQN has a weak ability to adjust the exploration-exploitation trade-off and is prone to falling into local optima. Therefore, it is not suitable for the control of complex 5GDHC systems.
[0092] As Figure 1 shown, to solve the above problems, the embodiments of the present invention creatively propose an intelligent control calculation method for energy supply actions, which is applicable to the energy hub system for 5GDHC model construction. 5GDHC adopts a decentralized cooling and heating configuration method, and the water temperature used is close to the ambient temperature. It cools and heats according to the end load demand and realizes heat exchange between different types of buildings through the pipe network, playing a role in transferring cooling and heat among different users / prosumers. The magnitude and direction of the water flow in the warm pipe and the cold pipe are not fixed, but change according to the energy demand. The end users are both users and suppliers, that is, "prosumers".
[0093] As Figure 1As shown in the figure, in the 5GDHC system, there are data centers that require cooling throughout the year. In winter, water is taken from the cold pipe for cooling the condenser of the unit, and the heat is discharged into the warm pipe. Hotels, bath centers, etc. in the system require domestic hot water throughout the year. Heat is taken from the warm pipe for heating the condenser of the unit, and the cold water is discharged into the cold pipe. When the temperatures of the warm pipe and the cold pipe cannot be maintained within the working range, the energy hub uses traditional refrigeration units, gas boilers, or various forms of renewable cold and heat sources for supplementary cooling or heating. It is a decentralized regional energy system.
[0094] The water temperature of the 5GDHC pipe network generally maintains in the range of 12 - 30 °C. More low-grade renewable energy and waste heat resources can be utilized. Different spatially distributed decentralized resources are connected through the pipe network and integrated into the 5GDHC system to form an energy internet similar to the energy hub in the smart grid. Due to the low water supply temperature, plastic pipes without insulation can be used for the pipe network, and the heat loss is very low, enabling long-distance transportation. For the looped pipe network, the conveying pump can be omitted, thus further reducing energy consumption. The overall pipe network only requires one cold pipe and one warm pipe, which can supply cooling and heating simultaneously. Users have a high degree of freedom in using, and the amount of use depends entirely on user needs.
[0095] The following gives the detailed DDPG algorithm for achieving optimal control in the continuous action space in the 5GDHC system, including environment modeling, state definition, action definition, reward function, and the specific DDPG update equation and algorithm flow, aiming to optimize energy consumption and operating costs while meeting the cooling and heating demands and safety boundaries of the 5GDHC system.
[0096] As Figure 2 shown, a 5GDHC model is established in the TRNSYS software environment, mainly including components such as the heat pump unit of the energy hub, warm pipe and cold pipe, building loads of end-users, heat pump units on the user side, and pumps on the user side. To reduce the complexity of the model, the cooling loads of 6 typical buildings such as end-office buildings, data centers, and hospital buildings are saved as data files, read into the model using Type9e, and transmitted to Type682 for energy interaction with the chilled water loop. The heat pump unit uses Type666, and the performance data file is determined according to the technical manual data of the heat pump unit. The pump uses Type110 variable-frequency pump, and the flow rate and power of the pump are controlled according to the frequency signal. The meteorological file uses the Shanghai standard meteorological year data CHN_Shanghai.Shanghai.583620_CSWD.epw provided in the meteorological file database of the EnergyPlus website. The overall model is as Figure 2 shown.
[0097] Figure 2The modules in the top row respectively calculate the power, hourly energy consumption, daily energy consumption, and total energy consumption during the simulation period of the heat pump unit and the water pump. The COP is the heating energy efficiency ratio of this system, that is, the cooling capacity of the entire system / (heat pump energy consumption + water pump energy consumption). These values, including the data in other TRNSYS models, need to be transmitted to the DDPG module for algorithm operation.
[0098] The DDPG algorithm needs to be implemented under the Python framework. Therefore, the Type3157 component is introduced in the TRNSYS model for data communication between the TRNSYS environment and the Python environment. Data exchange is carried out through a nested Python dictionary structure TRNData. Each TRNSYS component has an independent entry in the dictionary for storing inputs and outputs. The Python module defines 5 specific functions corresponding to different stages of the simulation and interacts with TRNSYS through the dictionary. In this way, the 5GDHC system under the TRNSYS environment utilizes the powerful functions of Python to achieve complex mathematical modeling and DDPG algorithm integration, greatly enhancing the ability of the optimization control system.
[0099] TRNSYS will perform a data interaction with Python every time it simulates a time step. The RL controller, that is, the DDPG policy, generates an action α based on the state s at the current moment t t and inputs it into the TRNSYS model of the 5GDHC system. The model calculates the system state s, system energy consumption E, load satisfaction rate, etc. at the next moment t + 1, and feeds these results back to the RL algorithm for training or decision update. The framework diagram of the Markov decision process for optimizing the operation of the 5GDHC model is as t shown in t+1 t . Among them, r Figure 3 is the reward function value at the next moment t + 1 of the moment t, which is used to t+1 based on the action α t .
[0100] As Figure 4 shown, the intelligent control calculation method for the energy supply action includes the following steps:
[0101] Step S1, construct an energy hub system through the 5GDHC model. The components in the energy hub system include a heat pump unit, a heating pipe and a cooling pipe, a building load on the end-user side, a heat pump unit on the end-user side, and a water pump on the end-user side;
[0102] Step S2, detect and obtain the hourly heating and cooling loads of the building on the end-user side of the energy hub system, the hourly load provided by the heat pump unit, the heating energy efficiency ratio of the energy hub system, the average temperature of the heating pipe, and the average temperature of the cooling pipe to form the state vector at the current moment;
[0103] Step S3, select the set outlet water temperature of the heat pump unit on the end-user side and the frequency of the pump on the end-user side as the control action at the current moment according to the state vector at the current moment;
[0104] Step S4, predict the state vector and the reward at the next moment according to the control action at the current moment;
[0105] Step S5, input the reward at the next moment into the deep deterministic policy gradient model for comparison with the expected state of the energy hub system at the next moment, and select an action based on the ε-greedy method to correct the control action at the current moment, and execute the corrected control action at the current moment.
[0106] Among them, the Fifth-Generation District Heating and Cooling (5GDHC) system has the characteristics of low-temperature operation, energy sharing, decentralization, and two-way energy flow.
[0107] In one embodiment, the detection obtains the hourly cooling and heating loads of the buildings on the end-user side of the energy hub system, the hourly loads provided by the heat pump units, the heating energy efficiency ratio of the energy hub system, the average temperature of the warm pipe, and the average temperature of the cold pipe, and forms the state vector at the current moment, including:
[0108] Set the current moment as represented by moment t, and the formula for the state vector at the current moment is
[0109] In the formula: is the hourly cooling and heating load (kW) of the j-th building on the end-user side at moment t;
[0110] is the hourly load (kW) provided by the j-th heat pump unit on the end-user side at moment t;
[0111] COP sys (t) is the heating energy efficiency ratio of the energy hub system at moment t;
[0112] is the average temperature (°C) of the warm pipe at moment t;
[0113] is the average temperature (°C) of the cold pipe at moment t.
[0114] In one embodiment, the selection of the set outlet water temperature of the heat pump unit on the end-user side and the frequency of the pump on the end-user side as the control action at the current moment according to the state vector at the current moment includes:
[0115] Let the current moment be represented by the moment \(t\), and the formula for controlling the action at the current moment is
[0116] where: \(f\) HP,j (t) is the hourly frequency (Hz) of the \(j\)-th water pump at moment \(t\);
[0117] is the set outlet water temperature (°C) of the \(j\)-th heat pump at moment \(t\).
[0118] In one embodiment, the predicting the next moment state vector and the next moment reward according to the control action at the current moment includes:
[0119] Predicting the next moment state vector according to the control action at the current moment, setting the next moment to be represented by the moment \(t + 1\), and the next moment state vector is
[0120] where: is the hourly cooling and heating load (kW) of the \(j\)-th end-user side building at moment \(t + 1\);
[0121] is the hourly load (kW) provided by the \(j\)-th end-user side heat pump unit at moment \(t + 1\);
[0122] COP sys (t + 1) is the heating energy efficiency ratio of the energy hub system at moment \(t + 1\);
[0123] is the average temperature (°C) of the warm pipe at moment \(t + 1\);
[0124] is the average temperature (°C) of the cold pipe at moment \(t + 1\);
[0125] Obtaining the next moment reward according to the next moment state vector, and the formula for the next moment reward is
[0126] where: \(E(t + 1)\) is the comprehensive energy consumption (kWh) of all heat pumps and water pumps within the time \(t + 1\);
[0127] indicates that if the difference between the hourly cooling and heating load of the \(j\)-th building and the hourly load provided by the \(j\)-th heat pump at moment \(t + 1\) exceeds the deviation threshold \(\epsilon\), a penalty is generated;
[0128] \(\Delta\) action is the action smoothing term, which characterizes the magnitude of the difference from the previous action and is used to penalize frequent adjustment and start / stop of equipment. Define \(\Delta\) action \(= ||a\)t -a t-1 || 2 ;
[0129] λ load is the penalty coefficient for imposing a penalty when the hourly load of the heat pump unit on the end-user side exceeds the limit;
[0130] λ action is the penalty coefficient for imposing a penalty when the action smoothing term of the heat pump unit on the end-user side exceeds the limit.
[0131] Preferably, the deviation threshold ε takes a value of 2% - 5% of the value.
[0132] In one embodiment, before inputting the next moment reward into the Deep Deterministic Policy Gradient Model (DDPG) and comparing it with the expected state of the energy hub system at the next moment, it further includes:
[0133] An experience replay pool is set in the Deep Deterministic Policy Gradient Model to store historical transition data, and the historical transition data includes the state vector, control action, and reward at each moment in the energy hub system;
[0134] The Deep Deterministic Policy Gradient Model is trained using the historical transition data in the experience replay pool until the training converges or reaches a preset number of rounds to obtain the trained Deep Deterministic Policy Gradient Model.
[0135] In one embodiment, the training of the Deep Deterministic Policy Gradient Model using the historical transition data in the experience replay pool includes:
[0136] It is set that the Deep Deterministic Policy Gradient Model includes an online actor network, an online critic network, a target actor network, and a target critic network;
[0137] The online actor network is set as μ(s|θ μ ), and the online critic network is set as Q(s, μ(s|θ μ )|θ Q ), where s is the state vector, θ μ is the first network weight, and θ Q is the second network weight;
[0138] The target actor network is set as μ'(s|θ μ' ), and the target critic network is set as Q'(s, μ'(s|θ μ' )|θ Q' ), where θ μ' is the third network weight, and θ Q' is the fourth network weight;
[0139] Initialize the values of the first network weight and the second network weight, copy the value of the first network weight to the third network weight, and copy the value of the second network weight to the fourth network weight;
[0140] Input the simulation state vector of the energy hub system into the online actor network, add exploration noise in the online actor network to form a simulation control action, execute the simulation control action to simulate the next moment to obtain the next moment simulation state vector and the next moment simulation reward, and store the simulation state vector, the simulation control action, the next moment simulation state vector, and the next moment simulation reward in the experience replay pool to accumulate data samples;
[0141] Randomly sample a batch of N sampling data {s i , a i , r i , s i+1} from the historical transfer data in the experience replay pool to update the first network weight and the second network weight, where s i is the state vector at time i, a i is the control action at time i, r i is the reward at time i, and s i+1 is the state vector at time i + 1;
[0142] When training the online critic network, the target Q value of the online critic network is y i = r i + γ · Q'(s i+1 , μ'(s i+1 |θ μ′ )|θ Q′ ), where γ is the discount factor, and the loss function is Update the second network weight θ Q by minimizing this loss through gradient descent;
[0143] When training the online actor network, the policy gradient of the online actor network is The goal of the online actor network is to select an action to maximize the target Q value of the online critic network, and the update gradient is the average value of the policy gradient Take the partial derivative of the input action in the online critic network, and send the update gradient back to the online actor network to further update the first network weight θ μ ;
[0144] After updating the online actor network, for the third network weight of the target actor network through θQ′ ← τθ Q + (1 - τ)θ Q′ Update; after updating the online critic network, update the fourth network weight of the target critic network through θ μ′ ← τθ μ + (1 - τ)θ μ′ Update; where τ << 1, enabling the target network to smoothly track the online network and stabilizing the training process.
[0145] In one embodiment, inputting the next moment reward into the Deep Deterministic Policy Gradient model (DDPG) and comparing it with the expected state of the energy hub system at the next moment, and selecting an action based on the ε-greedy method to correct the current moment control action, and executing the corrected current moment control action includes:
[0146] In response to the trained Deep Deterministic Policy Gradient model receiving the next moment reward, obtain the current moment state vector, the current moment control action, the current moment reward, and the next moment state vector;
[0147] Determine whether there is an offset between the next moment reward and the expected state of the energy hub system at the next moment;
[0148] If there is an offset, adjust the current moment control action based on the ε-greedy method so that the next moment reward is the same as the expected state of the energy hub system at the next moment, and execute the corrected current moment control action;
[0149] If there is no offset, execute the current moment control action;
[0150] Obtain the next moment reward and the next moment state vector after executing the control action, and store the current moment state vector, the current moment control action, the next moment reward, and the next moment state vector in the experience replay pool.
[0151] When training and running the Deep Deterministic Policy Gradient model, the following steps are included.
[0152] (1) Initialization
[0153] Initialize the first network weight θ of the online actor network and the online critic network μ and the second network weight θ Q , using small random values;
[0154] Copy to the target network: θ μ′ ← θ μ ,θ Q′ ← θ Q ;
[0155] Initialize an empty experience replay pool D with a capacity of 10 4 records.
[0156] (2) Interact with the simulation environment
[0157] Input the system state s t into the online actor network: a t = μ(s t |θ μ ) + N t , where N t is exploration noise (Gaussian noise) to facilitate policy exploration;
[0158] Execute the action a t in the simulation model, simulate the system evolution at the next moment t + Δt, and obtain s t+1 , reward r t ;
[0159] Store {s i , a i , r i , s i+1} into the experience replay pool D;
[0160] Update t ← t + Δt, and repeat the process to accumulate data samples.
[0161] (3) Network training
[0162] Randomly sample N memory samples from the experience replay pool D; i = 1, …, N, i.e., 1 ≤ i ≤ N;
[0163] Use the aforementioned online critic network update equation and online actor network update equation to update θ μ , θ Q ;
[0164] Use soft update to update the target critic network and target actor network;
[0165] After each update of the online network, the parameters of the target network are soft updated:
[0166] θ Q′ ← τθ Q + (1 - τ)θ Q′ ;
[0167] θ μ′ ← τθ μ + (1 - τ)θ μ′ ;
[0168] Repeat the above process until the training converges or reaches the preset number of rounds.
[0169] (4) Deployed to real-time simulation
[0170] Remove exploration noise. In the TRNSYS system, the online actor network directly outputs the optimal action: a t = μ(s t |θ μ ).
[0171] As Figure 5 shown, Figure 5 is the control framework diagram of the 5GDHC system based on DDPG.
[0172] Among them, the hyperparameters are set as follows:
[0173] Learning rate η of the online actor network μ : 0.01;
[0174] Learning rate η of the online critic network Q : 0.1;
[0175] Batch size N: 128;
[0176] Discount factor γ: 0.99;
[0177] Soft update coefficient τ: 0.001;
[0178] Exploration noise: Gaussian noise, the initial noise variance is set to 0.2 and decays gradually with training;
[0179] Network structure: The online actor network and the online critic network are both 3-layer MLP, the number of neurons in the hidden layer is 128, the activation function uses ReLU, and the activation function of the output layer of the online actor network uses Sigmoid.
[0180] In summary of the above links, the workflow of DDPG in the 5GDHC system includes: environmental modeling, defining states, defining actions, designing reward functions, and setting hyperparameters, etc. The DDPG algorithm initializes the network and the replay pool, adds exploration noise, samples regularly during the interaction, updates the online actor network and the online critic network, and simultaneously soft-updates the target critic network and the target actor network. The overall workflow framework is as Figure 5 shown.
[0181] Using the previously established TRNSYS simulation platform for simulation, the changes in the total energy consumption and the total proportion of time when the end load is not satisfied during the operation of the 5GDHC system are analyzed in detail. The simulation results are as Figure 6 shown. It can be intuitively seen from the figure the changes in the algorithm performance and the gradual optimization process in different simulation stages.
[0182] When the first complete simulation is finished, since the DDPG algorithm is still in the initial learning stage, there is a large randomness in the action selection strategy at this time. The reinforcement learning controller cannot effectively optimize according to the environmental feedback. Therefore, the energy consumption and the non - satisfaction time ratio of the end - load of the 5GDHC system are both at a relatively high level. This is a common phenomenon in the initial stage of the reinforcement learning algorithm. The main reason is that the controller has not accumulated enough experience data, and the direction and intensity of strategy adjustment are somewhat blind.
[0183] Entering the second complete simulation stage, the DDPG algorithm makes full use of the experience data obtained in the first simulation process. Through effective learning of the reward signal and policy update, it can select control actions more accurately. Therefore, the total energy consumption of the system and the non - satisfaction time ratio of the end - load decrease rapidly. This stage demonstrates the strong fast - learning ability of the DDPG algorithm and its effective adaptability to the system operation state.
[0184] In the third and fourth complete simulation stages, the total energy consumption of the 5GDHC system continues to decrease, indicating that the controller performs excellently in energy - efficiency optimization and the overall energy - efficiency of the 5GDHC system is significantly improved. However, during this period, the non - satisfaction time ratio of the end - load increases slightly. This may be because the DDPG algorithm focuses more on optimizing the energy - consumption target and thus sacrifices part of the end - load satisfaction rate. This trade - off phenomenon reveals the conflict between the energy consumption of the 5GDHC system and fully satisfying the end - load, which is worthy of further consideration in practical applications.
[0185] As the simulation progresses, the DDPG controller further accumulates experience knowledge through continuous interaction with the environment and gradually optimizes the control strategy. In the fifth and subsequent simulations, the non - satisfaction time ratio of the end - load gradually returns to the previous level and further decreases. At the same time, the total energy consumption of the 5GDHC system also continues to decrease. This shows that the DDPG algorithm takes into account maximizing the satisfaction of the end - load while optimizing the energy consumption, and the system energy - efficiency is further improved. The simulation results of this stage reflect the coordination ability of the DDPG algorithm in multi - objective optimization.
[0186] After the sixth and seventh complete simulations, the DDPG controller enters the stable period. At this time, the total energy consumption of the system basically stabilizes. Although the non - satisfaction time ratio of the end - load has minor fluctuations, it remains at a relatively low level overall, indicating that the energy - efficiency of the 5GDHC system is close to the optimal state. In this stage, the control actions output by the DDPG algorithm gradually tend to be stable, and the decision - making ability of the controller is fully exercised, enabling it to effectively cope with various operating conditions.
[0187] At the end of the last simulation, the reinforcement learning controller has successfully learned the optimal control strategy, indicating that after multiple iterative learning, the DDPG algorithm already has the ability to be put into practical use.
[0188] In the above intelligent control calculation method for energy supply actions, by introducing a continuous control algorithm of Deep Deterministic Policy Gradient (DDPG), it is possible to compare the reward at the next moment expected according to the control action at the current moment with the expected state of the energy hub system at the next moment, select an action based on the ε-greedy method to correct the control action at the current moment, achieve the adjustment of the cooling and heating supply of the energy hub system to be close to the expected result, improve the accuracy of the energy hub system, and overcome the limitations of existing control methods, thereby improving the operation efficiency and stability of the 5GDHC system.
[0189] The present invention proposes a control method for a small-scale fifth-generation district cooling and heating system (5GDHC) based on deep reinforcement learning. It not only breaks through the limitations of existing rule-based and temperature-based control methods in terms of control strategy, but also can respond to system load changes in real time through deep reinforcement learning, improving the energy utilization efficiency and stability of the 5GDHC system. This method performs excellently in optimizing energy efficiency and load balance. After applying the DDPG control method of the present invention, the total system energy consumption is reduced by about 10%, and the proportion of the time when the terminal load is not satisfied has gradually decreased from more than 6% at the initial stage to less than 2%, demonstrating good optimization effects. In addition, the control method of the present invention has strong adaptability and scalability and can be applied in 5GDHC systems of different scales. Especially in the case of large system load fluctuations, multiple sources and multiple sinks in the system, it can better adjust and ensure the efficient distribution of energy. Economically, the present invention can effectively reduce the operating cost by improving energy efficiency and has good economic returns. The control method using deep reinforcement learning technology also provides important technical support for the intelligentization and automation of the 5GDHC system, contributing to promoting energy conservation and emission reduction.
[0190] In one embodiment, as Figure 7 shown, an intelligent control calculation device 10 for energy supply actions is provided, including: a model construction module 1, a state vector acquisition module 2, a control action generation module 3, an expectation module 4, and a control action execution module 5.
[0191] The model construction module 1 is used to construct an energy hub system through a 5GDHC model. The components in the energy hub system include a heat pump unit, a heating pipe and a cooling pipe, a building load on the terminal user side, a heat pump unit on the terminal user side, and a water pump on the terminal user side.
[0192] The state vector acquisition module 2 is used to detect and obtain the hourly cooling and heating loads of the building on the terminal user side of the energy hub system, the hourly load provided by the heat pump unit, the heating energy efficiency ratio of the energy hub system, the average temperature of the heating pipe, and the average temperature of the cooling pipe to form a state vector at the current moment.
[0193] The control action generation module 3 is configured to select the set outlet water temperature of the heat pump unit on the end-user side and the frequency of the pump on the end-user side as the control action at the current moment according to the state vector at the current moment.
[0194] The prediction module 4 is configured to predict the state vector and the reward at the next moment according to the control action at the current moment.
[0195] The control action execution module 5 is configured to input the reward at the next moment into the deep deterministic policy gradient model for comparison with the expected state of the energy hub system at the next moment, select an action based on the ε-greedy method to correct the control action at the current moment, and execute the corrected control action at the current moment.
[0196] In one embodiment, the detection obtains the hourly heating and cooling loads of the building on the end-user side of the energy hub system, the hourly loads provided by the heat pump units, the heating energy efficiency ratio of the energy hub system, the average temperature of the heating pipes, and the average temperature of the cooling pipes, and forms the state vector at the current moment including:
[0197] Set the current moment as represented by moment t, and the formula for the state vector at the current moment is
[0198] In the formula: is the hourly heating and cooling load (kW) of the j-th building on the end-user side at moment t;
[0199] is the hourly load (kW) provided by the j-th heat pump unit on the end-user side at moment t;
[0200] COP sys (t) is the heating energy efficiency ratio of the energy hub system at moment t;
[0201] is the average temperature (°C) of the heating pipes at moment t;
[0202] is the average temperature (°C) of the cooling pipes at moment t.
[0203] In one embodiment, the selection of the set outlet water temperature of the heat pump unit on the end-user side and the frequency of the pump on the end-user side as the control action at the current moment according to the state vector at the current moment includes:
[0204] Set the current moment as represented by moment t, and the formula for the control action at the current moment is
[0205] In the formula: f HP,j(t) is the hourly frequency (Hz) of the j-th water pump at time t;
[0206] is the set outlet water temperature (°C) of the j-th heat pump at time t.
[0207] In one embodiment, the control action according to the current moment to predict the next moment state vector and the next moment reward includes:
[0208] Predict the next moment state vector according to the control action at the current moment, and set the next moment to be represented by time t + 1. The next moment state vector is
[0209] Where: is the hourly cooling and heating load (kW) of the j-th end-user side building at time t + 1;
[0210] is the hourly load (kW) provided by the j-th end-user side heat pump unit at time t + 1;
[0211] COP sys (t + 1) is the heating energy efficiency ratio of the energy hub system at time t + 1;
[0212] is the average temperature (°C) of the warm pipe at time t + 1;
[0213] is the average temperature (°C) of the cold pipe at time t + 1;
[0214] Obtain the next moment reward according to the next moment state vector. The formula for the next moment reward is
[0215] Where: E(t + 1) is the comprehensive energy consumption (kWh) of all heat pumps and water pumps within time t + 1;
[0216] indicates that if the difference between the hourly cooling and heating load of the j-th building and the hourly load provided by the j-th heat pump at time t + 1 exceeds the deviation threshold ε, a penalty is generated; (Preferably, the value of the deviation threshold ε is 2% - 5% of the value; )
[0217] Δ action is the action smoothing term, which characterizes the magnitude of the difference from the previous action and is used to penalize frequent adjustment and startup / shutdown of equipment. Define Δ action =||a t -a t-1 || 2 ;
[0218] λ load Provide a penalty coefficient for imposing a penalty when the hourly load of the heat pump unit on the end-user side exceeds the limit;
[0219] λ action Provide a penalty coefficient for imposing a penalty when the action smoothing term of the heat pump unit on the end-user side exceeds the limit.
[0220] In one embodiment, before inputting the next moment reward into the Deep Deterministic Policy Gradient Model (DDPG) and comparing it with the expected state of the energy hub system at the next moment, it further includes:
[0221] Set an experience replay pool in the Deep Deterministic Policy Gradient Model to store historical transition data, where the historical transition data includes the state vector, control action, and reward at each moment in the energy hub system;
[0222] Use the historical transition data in the experience replay pool to train the Deep Deterministic Policy Gradient Model until the training converges or reaches a preset number of rounds to obtain the trained Deep Deterministic Policy Gradient Model.
[0223] In one embodiment, the training of the Deep Deterministic Policy Gradient Model using the historical transition data in the experience replay pool includes:
[0224] Set that the Deep Deterministic Policy Gradient Model includes an online actor network, an online critic network, a target actor network, and a target critic network;
[0225] Set the online actor network as μ(s|θ μ ), set the online critic network as Q(s, μ(s|θ μ )|θ Q ), where s is the state vector, θ μ is the first network weight, and θ Q is the second network weight;
[0226] Set the target actor network as μ'(s|θ μ' ), set the target critic network as Q'(s, μ'(s|θ μ' )|θ Q' ), where θ μ' is the third network weight, and θ Q' is the fourth network weight;
[0227] Initialize the values of the first network weight and the second network weight, copy the value of the first network weight to the third network weight, and copy the value of the second network weight to the fourth network weight;
[0228] Input the simulation state vector of the energy hub system into the online actor network, add exploration noise in the online actor network to form a simulation control action, execute the simulation control action to simulate the next moment to obtain the simulation state vector and the simulation reward at the next moment, and store the simulation state vector, the simulation control action, the simulation state vector at the next moment, and the simulation reward at the next moment in the experience replay pool to accumulate data samples;
[0229] Randomly sample a batch of N sampling data {s i , a i , r i , s i+1} from the historical transition data in the experience replay pool to update the first network weight and the second network weight, where s i is the state vector at time i, a i is the control action at time i, r i is the reward at time i, and s i+1 is the state vector at time i + 1;
[0230] When training the online critic network, the target Q value of the online critic network is y i = r i + γ·Q'(s i+1 , μ'(s i+1 |θ μ′ )|θ Q′ ), where γ is the discount factor, and the loss function is Update the second network weight θ Q by minimizing this loss through gradient descent;
[0231] When training the online actor network, the policy gradient of the online actor network is The goal of the online actor network is to select actions to maximize the target Q value of the online critic network, and the update gradient is the average of the policy gradients Take the partial derivative of the input action in the online critic network, and send the update gradient back to the online actor network to further update the first network weight θ μ ;
[0232] After updating the online actor network, update the third network weight of the target actor network through θ Q′ ← τθ Q + (1 - τ)θ Q′ ; After updating the online critic network, update the fourth network weight of the target critic network through θ μ′ ← τθ μ+(1 - τ)θ μ′ is updated; where τ << 1.
[0233] In one embodiment, the inputting the next - moment reward into the Deep Deterministic Policy Gradient (DDPG) model and comparing it with the expected state of the energy hub system at the next moment, selecting an action based on the ε - greedy method to correct the current - moment control action, and executing the corrected current - moment control action include:
[0234] In response to the trained Deep Deterministic Policy Gradient model receiving the next - moment reward, obtaining the current - moment state vector, the current - moment control action, the current - moment reward, and the next - moment state vector;
[0235] Judging whether there is an offset between the next - moment reward and the expected state of the energy hub system at the next moment;
[0236] If there is an offset, adjusting the current - moment control action based on the ε - greedy method to make the next - moment reward the same as the expected state of the energy hub system at the next moment, and executing the corrected current - moment control action;
[0237] If there is no offset, executing the current - moment control action;
[0238] Obtaining the next - moment reward and the next - moment state vector after executing the control action, and storing the current - moment state vector, the current - moment control action, the next - moment reward, and the next - moment state vector in the experience replay pool. In the above - mentioned intelligent control calculation device for energy supply actions, by introducing the continuous control algorithm of Deep Deterministic Policy Gradient (DDPG), it can compare the next - moment reward expected according to the current - moment control action with the expected state of the energy hub system at the next moment, select an action based on the ε - greedy method to correct the current - moment control action, achieve the result that the cooling and heating of the energy hub system are close to the expected result, improve the accuracy of the energy hub system, and overcome the limitations of the existing control methods, thereby improving the operation efficiency and stability of the 5GDHC system.
[0239] For the specific limitations of the intelligent control calculation device for energy supply actions, reference can be made to the limitations of the intelligent control calculation method for energy supply actions in the above text, which will not be elaborated here. Each module in the above - mentioned intelligent control calculation device for energy supply actions can be implemented in whole or in part by software, hardware, and their combinations. The above - mentioned modules can be embedded in the processor of the computer device in hardware form or independent of it, or stored in the memory of the computer device in software form, so that the processor can call and execute the operations corresponding to the above - mentioned modules.
[0240] In one embodiment, a computer program product is provided, including a computer program which, when executed by a processor, implements the following steps:
[0241] Construct an energy hub system through a 5GDHC model. The components in the energy hub system include a heat pump unit, a heating pipe and a cooling pipe, a building load on the end-user side, a heat pump unit on the end-user side, and a water pump on the end-user side;
[0242] Detect and obtain the hourly heating and cooling loads of the building on the end-user side of the energy hub system, the hourly load provided by the heat pump unit, the heating energy efficiency ratio of the energy hub system, the average temperature of the heating pipe, and the average temperature of the cooling pipe, and form a state vector at the current moment;
[0243] Select the set outlet water temperature of the heat pump unit on the end-user side and the frequency of the water pump on the end-user side as the control action at the current moment according to the state vector at the current moment;
[0244] Predict the state vector and the reward at the next moment according to the control action at the current moment;
[0245] Input the reward at the next moment into the deep deterministic policy gradient model for comparison with the expected state of the energy hub system at the next moment, and select an action based on the ε-greedy method to correct the control action at the current moment, and execute the corrected control action at the current moment.
[0246] For the specific limitations on the steps implemented when the computer program is executed by the processor, reference can be made to the limitations on the method for intelligent control calculation of energy supply actions in the above text, which will not be elaborated here.
[0247] In one embodiment, a computer device is provided. The computer device can be a server, and its internal structure diagram can be as Figure 8 shown. The computer device includes a processor, a memory, a network interface, and a database connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store intelligent control calculation data for energy supply actions. The network interface of the computer device is used to communicate with an external terminal through a network connection. The computer program, when executed by the processor, implements a method for intelligent control calculation of energy supply actions.
[0248] Those skilled in the art can understand that Figure 8The structure shown is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.
[0249] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the following steps are implemented:
[0250] Construct an energy hub system through a 5GDHC model. The components in the energy hub system include a heat pump unit, a heating pipe and a cooling pipe, a building load on the end-user side, a heat pump unit on the end-user side, and a water pump on the end-user side;
[0251] Detect and obtain the hourly heating and cooling loads of the building on the end-user side of the energy hub system, the hourly load provided by the heat pump unit, the heating energy efficiency ratio of the energy hub system, the average temperature of the heating pipe, and the average temperature of the cooling pipe, and form a state vector at the current moment;
[0252] Select the set outlet water temperature of the heat pump unit on the end-user side and the frequency of the water pump on the end-user side as the control action at the current moment according to the state vector at the current moment;
[0253] Predict the state vector and reward at the next moment according to the control action at the current moment;
[0254] Input the reward at the next moment into the deep deterministic policy gradient model for comparison with the expected state of the energy hub system at the next moment, and select an action based on the ε-greedy method to correct the control action at the current moment, and execute the corrected control action at the current moment.
[0255] For the specific limitations on the steps implemented when the processor executes the computer program, reference can be made to the limitations on the method for intelligent control calculation of energy supply actions in the above text, which will not be elaborated here.
[0256] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented:
[0257] Construct an energy hub system through a 5GDHC model. The components in the energy hub system include a heat pump unit, a heating pipe and a cooling pipe, a building load on the end-user side, a heat pump unit on the end-user side, and a water pump on the end-user side;
[0258] Detect and obtain the hourly heating and cooling loads of the building on the end-user side of the energy hub system, the hourly loads provided by the heat pump units, the heating energy efficiency ratio of the energy hub system, the average temperature of the heating pipe, and the average temperature of the cooling pipe to form the state vector at the current moment;
[0259] Select the set outlet water temperature of the heat pump unit on the end-user side and the frequency of the water pump on the end-user side as the control action at the current moment according to the state vector at the current moment;
[0260] Predict the state vector and the reward at the next moment according to the control action at the current moment;
[0261] Input the reward at the next moment into the deep deterministic policy gradient model and compare it with the expected state of the energy hub system at the next moment. Select an action based on the ε-greedy method to correct the control action at the current moment, and execute the corrected control action at the current moment.
[0262] For the specific limitations on the implementation steps when the computer program is executed by the processor, reference can be made to the limitations on the method for intelligent control calculation of the energy supply action in the above text, which will not be elaborated here.
[0263] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to the memory, storage, database, or other media used in the various embodiments provided in the present application can include non-volatile and / or volatile memories. Non-volatile memories can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memories can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.
[0264] The technical features of the above embodiments can be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in this specification.
[0265] The above-described embodiments merely represent several implementation manners of the present application. The description is relatively specific and detailed, but it should not be construed as a limitation on the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application shall be subject to the appended claims.
Claims
1. An intelligent control calculation method for energy supply actions, characterized in that Including: Constructing an energy hub system through a 5GDHC model, where the components in the energy hub system include a heat pump unit, a heating pipe and a cooling pipe, a building load on the end-user side, a heat pump unit on the end-user side, and a water pump on the end-user side; Detecting and obtaining the hourly heating and cooling loads of the building on the end-user side of the energy hub system, the hourly load provided by the heat pump unit, the heating energy efficiency ratio of the energy hub system, the average temperature of the heating pipe, and the average temperature of the cooling pipe to form a state vector at the current moment; Selecting the set outlet water temperature of the heat pump unit on the end-user side and the frequency of the water pump on the end-user side as the control action at the current moment according to the state vector at the current moment; Predicting the state vector and the reward at the next moment according to the control action at the current moment; Inputting the reward at the next moment into the deep deterministic policy gradient model and comparing it with the expected state of the energy hub system at the next moment, and selecting an action based on the ε-greedy method to correct the control action at the current moment, and executing the corrected control action at the current moment.
2. The intelligent control calculation method for energy supply actions according to claim 1, characterized in that The detecting and obtaining the hourly heating and cooling loads of the building on the end-user side of the energy hub system, the hourly load provided by the heat pump unit, the heating energy efficiency ratio of the energy hub system, the average temperature of the heating pipe, and the average temperature of the cooling pipe to form a state vector at the current moment includes: Setting the current moment to be represented by moment t, and the formula for the state vector at the current moment is In the formula: is the hourly cooling and heating load of the j-th end-user building at time t; The hourly load provided for the j-th end-user side heat pump unit at time t; COP sys (t) is the heat production energy efficiency ratio of the energy hub system at time t; is the average temperature of the warm-up pipe at time t; is the average temperature of the cold tube at time t.
3. The intelligent control calculation method for energy supply actions according to claim 1, wherein The selecting the set outlet water temperature of the heat pump unit on the end-user side and the frequency of the water pump on the end-user side as the control action at the current moment according to the state vector at the current moment includes: Set the current moment to be represented by moment t, and the formula for controlling actions at the current moment is where: f HP,j (t) is the hourly frequency of the jth pump at time t; is the set outlet water temperature of the j-th heat pump at time t.
4. The intelligent control calculation method for energy supply actions according to claim 2, characterized in that, The predicting the state vector and the reward at the next moment according to the control action at the current moment includes: Predicting the state vector at the next moment according to the control action at the current moment, setting the next moment to be represented by moment t + 1, and the state vector at the next moment is In the formula: is the hourly cooling and heating load of the j-th end-user building at time t + 1; The hourly load provided for the j-th end-user side heat pump unit at time t+1; COP sys (t + 1) is the heat production energy efficiency ratio of the energy hub system at time t + 1; is the average temperature for warming the pipe at time t + 1; is the average temperature of the cold tube at time t+1; Obtaining the reward at the next moment according to the state vector at the next moment, and the formula for the reward at the next moment is In the formula: E(t + 1) is the comprehensive energy consumption of all heat pumps and water pumps within the time of t + 1; It means that if the difference between the hourly cooling and heating load of the j-th building and the hourly load provided by the j-th heat pump at time t+1 exceeds the deviation threshold ε, a penalty will be imposed. Δ action is the motion smoothness term, which characterizes the magnitude of the difference from the previous motion and is used to penalize the frequent adjustment and startup / shutdown of the device. Define Δ action =||a t -a t-1 || 2 ; λ load Provide a penalty coefficient for imposing penalties when the hourly load of the heat pump unit on the end-user side exceeds the limit. λ action The penalty coefficient for imposing a penalty when the operation smoothing term for the heat pump unit on the end-user side exceeds the limit.
5. The intelligent control calculation method for energy supply actions according to claim 1, characterized in that Before inputting the reward at the next moment into the deep deterministic policy gradient model and comparing it with the expected state of the energy hub system at the next moment, it further includes: Setting up an experience replay pool in the deep deterministic policy gradient model to store historical transition data, where the historical transition data includes the state vector, control action, and reward at each moment in the energy hub system; Training the deep deterministic policy gradient model with the historical transition data in the experience replay pool until the training converges or reaches a preset number of rounds to obtain the trained deep deterministic policy gradient model.
6. The intelligent control calculation method for energy supply operation according to claim 5, characterized in that, The training the deep deterministic policy gradient model with the historical transition data in the experience replay pool includes: Setting that the deep deterministic policy gradient model includes an online actor network, an online critic network, a target actor network, and a target critic network; Set the online actor network as μ(s|θ μ ), and set the online critic network as Q(s, μ(s|θ μ )|θ Q ), where s is the state vector, θ μ is the first network weight, and θ Q is the second network weight; Set the target actor network representation as μ'(s|θ μ' ), and set the target critic network as Q'(s, μ'(s|θ μ' )|θ Q' ), where θ μ' is the third network weight, and θ Q' is the fourth network weight; Initializing the values of the weights of the first network and the values of the weights of the second network, copying the values of the weights of the first network to the weights of the third network, and copying the values of the weights of the second network to the weights of the fourth network; Input the simulation state vector of the energy hub system into the online actor network, add exploration noise in the online actor network to form a simulation control action, execute the simulation control action to simulate the next moment to obtain the simulation state vector and simulation reward at the next moment, and store the simulation state vector, the simulation control action, the simulation state vector at the next moment, and the simulation reward at the next moment in the experience replay pool to accumulate data samples; Randomly extract a batch of sampling data {s with a quantity of N from the historical transfer data in the experience replay pool i , a i , r i , s i+1} to update the first network weight and the second network weight, where s i is the state vector at time i, a i is the control action at time i, r i is the reward at time i, and s i+1 is the state vector at time i + 1; When training the online critic network, the target Q-value of the online critic network is y i = r i + γ·Q'(s i+1 , μ'(s i+1 |θ μ′ )|θ Q′ ), where γ is the discount factor, and the loss function is The second network weights θ are updated by gradient descent to minimize this loss Q ; During the training of the online actor network, the policy gradient of the online actor network is The goal of the online actor network is to select actions to maximize the target Q value of the online critic network, and the update gradient is the average value of the policy gradient In the online critic network, take the partial derivative of the input action, and transmit the update gradient back to the online actor network, thereby updating the first network weight θ μ ; After updating the online actor network, the third network weight of the target actor network is updated by θ Q′ ← τθ Q′ +(1 - τ)θ Q′ After updating the online critic network, the fourth network weight of the target critic network is updated by θ μ′ ← τθ μ +(1 - τ)θ μ′ where τ << 1.
7. The intelligent control calculation method for energy supply operation according to claim 5, wherein Input the reward at the next moment into the deep deterministic policy gradient model and compare it with the expected state of the energy hub system at the next moment. Select an action based on the ε-greedy method to correct the control action at the current moment. Executing the corrected control action at the current moment includes: In response to the trained deep deterministic policy gradient model receiving the reward at the next moment, obtain the state vector at the current moment, the control action at the current moment, the reward at the current moment, and the state vector at the next moment; Determine whether there is an offset between the reward at the next moment and the expected state of the energy hub system at the next moment; If there is an offset, adjust the control action at the current moment based on the ε-greedy method so that the reward at the next moment is the same as the expected state of the energy hub system at the next moment, and execute the corrected control action at the current moment; If there is no offset, execute the control action at the current moment; Obtain the reward at the next moment and the state vector at the next moment after executing the control action, and store the state vector at the current moment, the control action at the current moment, the reward at the next moment, and the state vector at the next moment in the experience replay pool.
8. An intelligent control calculation device for energy supply actions, characterized in that, The device includes: A model construction module for constructing an energy hub system through a 5GDHC model. The components in the energy hub system include a heat pump unit, a heating pipe and a cooling pipe, a building load on the end-user side, a heat pump unit on the end-user side, and a water pump on the end-user side; A state vector acquisition module for detecting and obtaining the hourly heating and cooling loads of the building on the end-user side of the energy hub system, the hourly load provided by the heat pump unit, the heating energy efficiency ratio of the energy hub system, the average temperature of the heating pipe, and the average temperature of the cooling pipe to form a state vector at the current moment; A control action generation module for selecting the set outlet water temperature of the heat pump unit on the end-user side and the frequency of the water pump on the end-user side as the control action at the current moment according to the state vector at the current moment; An expectation module for expecting the state vector and reward at the next moment according to the control action at the current moment; A control action execution module for inputting the reward at the next moment into the deep deterministic policy gradient model and comparing it with the expected state of the energy hub system at the next moment, selecting an action based on the ε-greedy method to correct the control action at the current moment, and executing the corrected control action at the current moment.
9. A computer device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 7.
10. A computer-readable storage medium, having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the method according to any one of claims 1 to 7.