An Energy Management Method for Integrated Energy Systems Based on Improved Deep Reinforcement Learning
By improving the deep reinforcement learning method and the equivalent encapsulation model of long and short-term memory neural networks, the limitations of mathematical optimization methods in energy management of integrated energy systems are solved and the problem that heuristic algorithms are difficult to ensure global optimality, achieving more efficient and stable energy management.
Patent Information
- Application Number
- CN202210965022.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-08-12
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2042-08-12
AI Technical Summary
When the prior art deals with energy management of integrated energy systems, mathematical optimization methods have limitations in large-scale nonlinear programming problems, and heuristic algorithms are difficult to ensure the global optimality of solutions.
The improved deep reinforcement learning method is adopted to build an equivalent encapsulation model of the comprehensive energy system through long and short-term memory neural networks, simplify the complex iteration process, and use the k-priority sampling strategy and the improved deep reinforcement learning algorithm for online learning in large-scale action space.
It reduces the difficulty of solving energy management solutions, improves convergence and stability, and can realize the adaptive learning evolution of thermal and electrical multi-energy management strategies in complex scenarios, improving the operational economy of the comprehensive energy system.
Smart Images

Figure CN115409645B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of integrated energy system control, and particularly relates to an energy management method for an integrated energy system based on improved deep reinforcement learning. Background Art
[0002] In order to promote the process of global low-carbon transformation, the energy and power industry, which accounts for a relatively large proportion of carbon emissions, has brought new challenges. The integrated energy system can realize the complementarity of multiple energies such as electricity, heat, and gas, and is an important starting point for optimizing and transforming the energy structure and promoting the realization of the low-carbon development goal. The construction direction of the integrated energy system is gradually developing from the "source-source" horizontal multi-energy complementary system to the "source-network-load-storage" vertical integration direction. Reasonable energy management of the integrated energy system is an effective way to reduce the impact of distributed energy fluctuations on the power grid, promote the development and application of renewable energy, alleviate the shortage of fossil energy, and reduce carbon emissions. Therefore, configuring a reasonable and effective energy management method for the integrated energy system is of great significance for accelerating the construction of a low-carbon integrated energy system.
[0003] At present, there have been a large number of studies on the energy management and optimal scheduling of integrated energy systems. The mainstream methods include mathematical optimization methods represented by nonlinear programming, second-order cone programming, mixed integer programming, etc., and heuristic algorithms represented by genetic algorithms and particle swarm algorithms. Chinese invention patent CN111969602A provides a day-ahead stochastic optimal scheduling method and device for an integrated energy system, which uses a parallel optimization method of dynamic programming to solve the day-ahead stochastic optimal scheduling model aiming at minimizing the expected cost of the integrated energy system operation; although the mathematical optimization method has clear theory and can guarantee the optimality of the solution to a certain extent, such mathematical programming models usually appropriately simplify the constraint conditions of the energy supply system and have limitations in dealing with large-scale nonlinear programming problems. Chinese invention patent CN111463773A provides an energy management optimization method and device for a regional integrated energy system, which uses the Monte Carlo method for sampling and combines the genetic algorithm for solution to construct an optimization model aiming at the lowest energy management cost of the regional integrated energy system; although such heuristic algorithms are convenient to solve and can guarantee better results within polynomial time, the obtained results are difficult to guarantee the global optimality of the solution. Summary of the Invention
[0004] To overcome the shortcomings of the existing technologies, the present invention proposes an energy management method for an integrated energy system based on improved deep reinforcement learning. The present invention simplifies the complex iterative process during the interaction of multiple integrated energy systems through the equivalent modeling of a long short-term memory neural network, reduces the difficulty of solving the energy management scheme, and at the same time, the improved deep reinforcement learning algorithm can reduce the access frequency to low-reward actions during the exploration of a large-scale action space, and has better convergence and stability. In addition, the present invention does not require detailed understanding of the detailed parameter information of the devices in each park, and can also achieve the adaptive learning and evolution of thermal and electrical multi-energy management strategies in complex and changing scenarios, improving the operating economy of the integrated energy system.
[0005] To achieve the above object, the technical solution adopted by the present invention is as follows:
[0006] An energy management method for an integrated energy system based on improved deep reinforcement learning mainly includes the following steps:
[0007] Step (1): Based on the historical operation data of the integrated energy system, build an equivalent encapsulation model of the integrated energy system using a long short-term memory neural network;
[0008] Step (2): Construct a reinforcement learning environment required for learning and training the energy management strategy of the integrated energy system;
[0009] Step (3): Adopt a k-priority sampling strategy and perform online learning on the energy management strategy of the integrated energy system based on an improved deep reinforcement learning algorithm.
[0010] Further, in the step (1), based on the historical operation data of the integrated energy system, build an equivalent encapsulation model of the integrated energy system using a long short-term memory neural network. The steps are as follows:
[0011] Step (1-1): Select the input variables and output variables of the long short-term memory neural network model
[0012] The historical operation data of the integrated energy system mainly includes: the output of uncontrollable distributed renewable energy generating units such as wind turbines and photovoltaic units, the output of controllable distributed generating units such as micro gas turbines and fuel cells, electrical loads, thermal loads, electricity trading prices, heat trading prices, electricity trading volumes, and heat trading volumes. For the needs of optimal operation and coordinated operation, the output variables are selected as the electricity trading volume and heat trading volume of the integrated energy system, and the remaining variables are used as input variables;
[0013] Step (1-2): Data processing, statistically analyze the historical operation data of the integrated energy system, and perform preprocessing such as data normalization and division of training sets and test sets;
[0014]
[0015] In Equation (1), D represents a data set composed of historical operation data; X represents a column vector composed of a set of all variables, d represents the d-th day, M represents the total number of days; t represents the t-th time period in a day, and N is usually 24, representing 24 time periods in a day; D u represents the normalized historical data; min(·) represents the minimum value function, and max(·) represents the maximum value function; represents the training set taken from the historical data after normalization, represents the test set taken from the historical data after normalization, and ε represents the proportion of the training set in the total data set;
[0016] Step (1-3): Train the long short-term memory neural network model:
[0017] Adopt a long short-term memory neural network and use the mini-batch gradient descent method based on backpropagation to learn and train the training set data:
[0018]
[0019] In Equation (2), x t represents the data set taken from the training data set at the t-th time period; h t-1 represents the accumulation before the t-th time period; f t represents the output of the forget gate corresponding to the t-th time period in the current iteration, w f and b f are the weight coefficients and bias coefficients of each neuron in the forget layer, σ(·) represents the sigmoid curve function, i t represents the output of the input layer at the t-th time period, w i and b i are the weight coefficients and bias coefficients of each neuron in the input layer, represents the predicted output of the convolutional layer at the t-th time period, w c and b c are the weight coefficients and bias coefficients of each neuron in the convolutional layer, tanh(·) represents the hyperbolic tangent function, c t represents the actual output of the convolutional layer at the t-th time period, o t represents the output of the output layer at the t-th time period, w o and b o are the weight coefficients and bias coefficients of each neuron in the output layer, h t represents the actual output at the t-th time period;
[0020] Step (1-4): Evaluate the effect of the long short-term memory neural network model:
[0021] The long short-term memory neural network model is tested using the test set, and the root mean square error is used for effect evaluation;
[0022]
[0023] In Equation (3), RMSE represents the root mean square error between the model prediction value and the true value, x test represents the input variable of the network in the test set, y test represents the output variable of the network in the test set, and net represents the trained network function.
[0024] Furthermore, in step (2), the steps for constructing the reinforcement learning environment required for learning and training the energy management strategy of the integrated energy system are as follows:
[0025] Step (2-1): Set the state space:
[0026] Regarding the control center of each integrated energy system as an agent, the observable state space of the agent is:
[0027] S = S C ×S X ×S T (4)
[0028] In Equation (4), S C represents the controllable observable, S X represents the uncontrollable observable, S T represents the time series information observable;
[0029] The controllable observables include the state of charge SoC of distributed energy storage in the integrated energy system t , the state of charge SoT of the TCL load t and the market price level C t , and the observables are shown as follows:
[0030] S C = [SoC t , SoT t , C b t (5)
[0031] In Equation (5), the uncontrollable observables include the temperature T t , the electric energy G provided by distributed energy t , the thermal energy H provided by distributed energy t , the energy trading price with other integrated energy systems and the electric load and the heat load The unobservable quantities are shown in Equation (6):
[0032]
[0033] The time-series information observation quantity includes the current day t d , the current hour t h , as shown in Equation (7):
[0034] S T = [t d , t h (7)
[0035] Step (2-2): Set the action space:
[0036] The action space of the agent is a 10-dimensional discrete space, and this action space mainly includes the control of electric energy A e and the control of thermal energy A h , as shown in Equation (8):
[0037] A = A e × A h (8)
[0038] The control action for electric energy is:
[0039] A e = [a tcl , a l , a c , a G , a p , a s (9)
[0040] In Equation (9), a tcl is the control signal of the TCL load, a l is the control information of the price-responsive electric load, a c is the charge and discharge control signal of the distributed energy storage tank, a G is the power generation power control signal of the gas turbine, a p is the electric energy trading price control signal, a s is the electric energy trading order control signal;
[0041] The control action for thermal energy is:
[0042] A h = [a hc , a hG , a hp , a hs (10)
[0043] In Equation (10), a hc is the control signal of the heat storage tank, a hG is the supplementary combustion control signal of the boiler, a hp is the thermal energy trading price control signal, ahs It is the control signal for the thermal energy trading sequence.
[0044] Step (2-3): Set the reward function:
[0045] In order to maximize the objective of the energy management scheme of each integrated energy system to load its own interests, the set reward function is as follows:
[0046] R t = S t - C t + Pen t (11)
[0047] In formula (11), S t is the revenue from selling energy, C t is the cost of obtaining energy, and Pen t is the penalty term;
[0048]
[0049] In formula (12), the revenue S from selling energy t mainly comes from the internal users of the integrated energy system and other integrated energy systems; N l is the number of internal load users of the integrated energy system, and L i t is the electrical load magnitude of the i-th user at time t, and L i h,t is the thermal load magnitude of the i-th user at time t, and P t is the electricity selling price at time t, and P h,t is the thermal energy selling price at time t; N a is the number of tradable integrated energy systems, and P j t is the electricity selling price to the j-th integrated energy system at time t, and E j t is the electricity magnitude sold to the j-th integrated energy system at time t, and P j h,t is the thermal energy selling price to the j-th integrated energy system at time t, and H j t is the thermal energy magnitude sold to the j-th integrated energy system at time t;
[0050]
[0051] In formula (13), the cost C of obtaining energy t mainly comes from the power generation and heat production costs of distributed energy and the purchase cost from other integrated energy systems; C e is the power generation cost, and Gt is the power generation of the micro gas turbine at time t, C h is the heat energy cost, H t is the heat energy provided by supplementary combustion of the boiler at time t, P k t is the electricity purchase price from the k-th integrated energy system at time t, E k t is the amount of electricity purchased from the k-th integrated energy system at time t, P k h,t is the heat energy purchase price from the k-th integrated energy system at time t, H k t is the amount of heat energy purchased from the k-th integrated energy system at time t;
[0052]
[0053] In Equation (14), λ is the penalty coefficient. The penalty term is always 0 at non-starting times of each day, and is determined according to the difference in SoC from the initial time of the day at the last time of each day.
[0054] Furthermore, in step (3), the k-priority sampling strategy is adopted, and the steps for online learning of the energy management strategy of the integrated energy system based on the improved deep reinforcement learning algorithm are as follows:
[0055] Step (3-1): Initialize the experience pool and Q-network parameters:
[0056] Randomly initialize the actions of the agent, record the state transition process of the agent, and store the current state, the currently taken action, the next state, and the reward function of the agent in the experience pool until the experience pool is full. At the same time, initialize the weights of the target Q-network;
[0057] Step (3-2): Obtain the current environmental state s t :
[0058] Take the output of the wind turbine, the output of the photovoltaic unit, the state of the distributed energy storage, the size of the electrical load, the size of the thermal load, the real-time electricity trading price, and the real-time heat trading price in the integrated energy system during the current period as the environmental state s observable by the agent t ;
[0059] Step (3-3): Improve the deep reinforcement learning algorithm with the k-priority sampling strategy and select the current action a t :
[0060] The k - priority sampling strategy first selects k candidate actions with the highest Q - values among all actions, then calculates the normalized scores of the k candidate actions according to the softmax function, and finally selects an action according to the probability distribution that conforms to the normalized scores.
[0061] The mathematical expression of the k - priority sampling strategy is:
[0062]
[0063] In Equation (15), s is the current state of the agent; a is the action that the agent can choose; π(a|s) is the policy function, which is used to describe the probability of choosing action a in state s; Q(s,a) is the action - value function composed of state s and action a; a k ∈A * , A * is the set composed of the k actions with the highest action - value Q(s,a) among all action - values, and its expression is:
[0064]
[0065] In Equation (16), represents the k actions with the largest action - value function in the set of all actions;
[0066] Step (3 - 4): Update the experience pool:
[0067] Execute the current action a obtained according to the k - priority adoption strategy t , and obtain the state s at the next moment t+1 and the reward value r t , and store the state transition process in the form of (s t ,a t ,r t ,s t+1 ) in the experience pool. If the experience pool is full, delete the earliest experience record. If the experience pool is not full, proceed to the next step;
[0068] Step (3 - 5): Update the Q - network parameters:
[0069] Randomly extract N data (si,ai,ri,si + 1) from the experience pool and calculate the target network prediction value:
[0070] y i =r i +γmax a Q ω′ (s i+1 ,a) (17)
[0071] In Equation (17), y irepresents the predicted value of the target network for the i-th sample, γ is the attenuation coefficient, and Q ω′ (s i+1 , a) is the action value function for state s i+ 1 calculated by the target network, represents the target network parameters;
[0072] Update the Q-network parameters using the gradient descent method, and minimize the loss function as:
[0073]
[0074] In Equation (18), Q ω (s i , a i ) is the action value function for state s i calculated by the evaluation network, represents the evaluation network parameters;
[0075] Finally, repeat steps (3-2) to (3-5) until the maximum number of training times is reached.
[0076] Beneficial effects:
[0077] The present invention simplifies the complex iterative process during the interaction of multiple integrated energy systems through the equivalent modeling of the long short-term memory neural network, reduces the difficulty of solving the energy management scheme, and at the same time, the improved deep reinforcement learning algorithm can reduce the access frequency to low-reward actions during the exploration of a large-scale action space, with better convergence and stability; in addition, the present invention does not require detailed knowledge of the detailed parameter information of the equipment in each integrated energy system, and can also achieve the adaptive learning and evolution of the thermal and electrical multi-energy management strategies in complex and changing scenarios, improving the operating economy of the integrated energy system. Compared with traditional mathematical optimization methods, the present invention does not need to simplify the constraint conditions of the integrated energy system, can fully reflect the dynamic characteristics of the integrated energy system, and the solution results are more accurate, and can be applied to complex non-linear scenarios; compared with heuristic algorithms, the present invention has better convergence performance, can be applied to different scenarios at the same time, does not need to retrain the model, and can realize the function of real-time energy management. Description of the Drawings
[0078] Figure 1 is the flow chart of the integrated energy system management method based on the improved deep reinforcement learning algorithm of the present invention;
[0079] Figure 2 is the flow chart of the improved deep reinforcement learning algorithm of the present invention. Detailed Embodiments
[0080] To make the objectives, technical solutions and advantages of the present invention more clear and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention. In addition, the technical features involved in the various embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.
[0081] The energy management method of the park integrated energy system based on the improved deep reinforcement learning algorithm of the present invention mainly includes the following steps:
[0082] Step 1: Based on the historical operation data of the integrated energy system, a long short-term memory neural network is used to build an equivalent encapsulation model of the integrated energy system;
[0083] Step 2: Construct a reinforcement learning environment required for learning and training the energy management strategy of the integrated energy system;
[0084] Step 3: Adopt the k-priority sampling strategy and perform online learning on the energy management strategy of the integrated energy system based on the improved deep reinforcement learning algorithm.
[0085] The specific implementation process of the present invention is as Figure 1 shown and includes the following steps:
[0086] Step 1. Based on the historical operation data of the integrated energy system, a long short-term memory neural network is used to build an equivalent encapsulation model of the integrated energy system, which specifically includes:
[0087] (1-1) Select the input variables and output variables of the long short-term memory neural network model.
[0088] The historical operation data of the integrated energy system mainly includes: the output of uncontrollable distributed renewable energy generating units such as wind turbines and photovoltaic units, the output of controllable distributed generating units such as micro gas turbines and fuel cells, electrical loads, heat loads, electricity trading prices, heat trading prices, electricity trading volumes, and heat trading volumes. For the needs of optimal operation and coordinated operation, the output variables are selected as the electricity trading volume and heat trading volume of the integrated energy system, and the remaining variables are used as input variables;
[0089] (1-2) Data processing, counting the historical operation data of each integrated energy system, and performing preprocessing such as data normalization and division of training set and test set;
[0090]
[0091] In Equation (1), D represents the data set composed of historical operation data; X represents a column vector composed of a set of all variables, d represents the d-th day, M represents the total number of days; t represents the t-th period in a day, and N is usually 24, representing 24 periods in a day; D u represents the normalized historical data; min(·) represents the minimum value function, and max(·) represents the maximum value function; represents the training set extracted from the historical data after normalization, represents the test set extracted from the historical data after normalization, and ε represents the proportion of the training set in the total data set;
[0092] (1-3) Train the long short-term memory neural network model.
[0093] Adopt the long short-term memory neural network and use the mini-batch gradient descent method based on backpropagation to learn and train the data in the training set:
[0094]
[0095] In Equation (2), x t represents the data set extracted from the training data set at the t-th period; h t-1 represents the accumulation before the t-th period; f t represents the output of the forgetting gate corresponding to the t-th period of the current iteration, w f and b f are the weight coefficients and bias coefficients of each neuron in the forgetting layer, σ(·) represents the sigmoid function, i t represents the output of the input layer at the t-th period, w i and b i are the weight coefficients and bias coefficients of each neuron in the input layer, represents the estimated output of the convolutional layer at the t-th period, w c and b c are the weight coefficients and bias coefficients of each neuron in the convolutional layer, tanh(·) represents the hyperbolic tangent function, c t represents the actual output of the convolutional layer at the t-th period, o t represents the output of the output layer at the t-th period, w o and b o are the weight coefficients and bias coefficients of each neuron in the output layer, h t represents the actual output at the t-th period;
[0096] (1-4) Evaluate the effect of the long short-term memory neural network model.
[0097] Use the test set to test the long short-term memory neural network model and adopt the root mean square error for effect evaluation;
[0098]
[0099] In Equation (3), RMSE represents the root mean square error between the model prediction value and the true value, and x test represents the input variable of the network in the test set, and y test represents the output variable of the network in the test set, and net represents the trained network function.
[0100] Step 2: Construct the reinforcement learning environment required for learning and training the energy management strategy of the integrated energy system, specifically including:
[0101] (2-1) Set the state space:
[0102] The state space observable by the agent is:
[0103] S = S C × S X × S T (4)
[0104] In Equation (4), S C represents the controllable observable, S X represents the uncontrollable observable, and S T represents the time series information observable;
[0105] The controllable observables include the state of charge SoC of distributed energy storage in the integrated energy system t , the state of charge SoT of the TCL load t and the market price level C t , and the observables are shown as follows:
[0106] S C = [SoC t , SoT t , C b t (5)
[0107] The uncontrollable observables include the temperature T t , the electric energy G provided by distributed energy t , the thermal energy H provided by distributed energy t , the energy trading price with other integrated energy systems and the electric load and the heat load The unobservable quantities are shown in Equation (6):
[0108]
[0109] The time series information observables include the current day t d , the current hour t h , as shown in Equation (7):
[0110] S T = [t d , t h (7)
[0111] (2 - 2) Set the action space:
[0112] Regard the control center of each integrated energy system as an agent, and its action space is a 10 - dimensional discrete space. This action space A mainly includes the control of electric energy A e and the control of thermal energy A h , as shown in Equation (8):
[0113] A = A e × A h (8)
[0114] The control action for electric energy is:
[0115] A e = [a tcl , a l , a c , a G , a p , a s (9)
[0116] In Equation (9), a tcl is the control signal of the TCL load, a l is the control information of the price - responsive electric load, a c is the charge - discharge control signal of the distributed energy storage tank, a G is the power generation control signal of the gas turbine, a p is the electric energy trading price control signal, a s is the electric energy trading order control signal;
[0117] The control action for thermal energy is:
[0118] A h = [a hc , a hG , a hp , a hs (10)
[0119] In Equation (10), a hc is the control signal of the heat storage tank, a hG is the supplementary combustion control signal of the boiler, a hp is the thermal energy trading price control signal, a hs is the thermal energy trading order control signal.
[0120] (2 - 3) Set the reward function:
[0121] To maximize the objective of the energy management scheme of each integrated energy system in terms of its own interests, the reward function is set as follows:
[0122] R t = S t - C t + Pen t (11)
[0123] In Equation (11), S t is the revenue from selling energy, C t is the cost of obtaining energy, and Pen t is the penalty term;
[0124]
[0125] In Equation (12), the revenue S from selling energy t mainly comes from the internal users of the integrated energy system and other integrated energy systems; N l is the number of internal load users in the integrated energy system, and L i t is the electrical load magnitude of the i-th user at time t, and L i h,t is the thermal load magnitude of the i-th user at time t, and P t is the electricity selling price at time t, and P h,t is the thermal energy selling price at time t; N a is the number of tradable integrated energy systems, and P j t is the electricity selling price to the j-th integrated energy system at time t, and E j t is the electricity magnitude sold to the j-th integrated energy system at time t, and P j h,t is the thermal energy selling price to the j-th integrated energy system at time t, and H j t is the thermal energy magnitude sold to the j-th integrated energy system at time t;
[0126]
[0127] In Equation (13), the cost C of obtaining energy t mainly comes from the power generation and heat production costs of distributed energy and the purchase cost from other integrated energy systems; C e is the power generation cost, and G t is the power generation amount of the micro gas turbine at time t, and C h is the thermal energy cost, and H t is the thermal energy provided by the boiler supplementary combustion at time t, and P k tThe electricity purchase price from the k-th integrated energy system at time t, E k t The amount of electricity purchased from the k-th integrated energy system at time t, P k h,t The heat purchase price from the k-th integrated energy system at time t, H k t The amount of heat purchased from the k-th integrated energy system at time t;
[0128]
[0129] In Equation (14), λ is the penalty coefficient. The penalty term is always 0 at non-starting times of each day and is determined according to the difference in SoC from the initial time of the day at the last time of each day.
[0130] Step 3. Use the k-priority sampling strategy to replace the ε-greedy strategy to improve the deep reinforcement learning algorithm, and online learn the energy management strategy of the integrated energy system based on the improved deep reinforcement learning algorithm, specifically including:
[0131] (3-1) Initialize the experience pool and Q-network parameters:
[0132] Randomly initialize the actions of the energy management agent of the integrated energy system, record the state transition process of the agent, and store the current state, the current action taken, the next state, and the reward function of the energy management agent of the integrated energy system in the experience pool until the experience pool is full. At the same time, initialize the weights of the Q-network; in reinforcement learning, the Q(s,a) function is used to represent the cumulative expected return that can be obtained by taking action a in state s. In the case of a continuous state space, it is usually impossible to effectively maintain the Q-table, and it is necessary to use the method of approximating the value function to approximate the Q function. The Q-network is a method of using a neural network to approximate the Q value. At the same time, to avoid the instability of the Q value caused by frequent network updates, two sets of Q-networks are used for alternating updates. Among them, the parameters of the evaluation Q-network are initialized to The parameters of the target Q-network are initialized to The evaluation Q-network is updated at each step, and the target Q-network is updated at regular intervals.
[0133] (3-2) Obtain the current environmental state s t :
[0134] Take the output of the wind turbine, the output of the photovoltaic unit, the distributed energy storage state, the electricity load size, the heat load size, the real-time electricity trading price, and the real-time heat trading price in the integrated energy system during the current period as the environmental state s observable by the agent t ;
[0135] (3-3) Improve the deep reinforcement learning algorithm using the k-priority sampling strategy and select the current action a t :
[0136] Traditional deep reinforcement learning methods use the ε-greedy strategy, that is, when selecting an action each time, the optimal action is selected with a probability of 1-ε, and other actions are explored with a probability of ε. Its policy function is:
[0137]
[0138] In formula (15), a * = argmax a Q(s,a), representing the greedy action; the ε-greedy strategy helps to traverse the action space in a small-scale action space and balance the exploration rate and utilization rate of the strategy; s is the current state of the agent; a is the action that the agent can choose; π(a|s) is the policy function, which is used to describe the probability of selecting action a in state s. This strategy is only applicable to the reinforcement learning environment of low-dimensional discrete action spaces. When facing large-scale discrete action spaces, it will face problems such as low exploration efficiency, slow convergence speed, and easy convergence to sub-optimal solutions. This is because in high-dimensional discrete action spaces, the traditional ε-greedy strategy is too inefficient when taking non-greedy strategies to explore and cannot effectively update the Q-value network parameters. Therefore, the present invention proposes a k-priority sampling strategy for large-scale discrete action spaces.
[0139] The flowchart of the improved deep reinforcement learning algorithm of the present invention is as Figure 2 shown:
[0140] The k-priority sampling strategy first selects k candidate actions with the highest Q-values according to the Q-values of all actions, then calculates the normalized scores of the k candidate actions according to the softmax function, and finally completes the selection of actions according to the probability distribution that conforms to the normalized scores.
[0141] The mathematical expression of the k-priority sampling strategy is:
[0142]
[0143] In formula (16), s is the current state of the agent; a is the action that the agent can choose; π(a|s) is the policy function, which is used to describe the probability of selecting action a in state s; Q(s,a) is the action value function composed of state s and action a; a k ∈A * , A * is the set composed of the k actions with the highest action values Q(s,a) among all action values, and its expression is:
[0144]
[0145] In formula (17), represents the k actions with the largest action value functions in the set of all actions;
[0146] (3-4) Update the experience pool:
[0147] Execute the current action a obtained according to the k-priority adoption strategy t , and obtain the state s at the next moment t+1 and the reward value r t , and store the state transition process in the form of (s t , a t , r t , s t+1 ) in the experience pool. If the experience pool is already full, delete the earliest experience record. If the experience pool is not full, proceed to the next step;
[0148] (3-5) Update the Q-network parameters:
[0149] Randomly extract N data (s i , a i , r i , s i+1 ) from the experience pool, and calculate the target network prediction value:
[0150] y i = r i + γ max a Q ω′ (s i+1 , a) (18)
[0151] In formula (18), y i represents the target network prediction value of the i-th sample, γ is the attenuation coefficient, and Q ω′ (s i+1 , a) is the action value function in the state s calculated by the target network i+1 , represents the target network parameters;
[0152] Update the Q-network parameters using the gradient descent method, and minimize the loss function as:
[0153]
[0154] In formula (19), Q ω (s i , a i ) is the action value function in the state s calculated by the evaluation network i , represents the evaluation network parameters;
[0155] Finally, repeat steps (3-2) to (3-5) until the maximum number of training times is reached.
[0156] Those skilled in the art can easily understand that the above are only the preferred embodiments of the present invention, and are not intended to limit the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.
Claims
1. An energy management method for an integrated energy system based on improved deep reinforcement learning, characterized in that, It includes the following steps: Step (1): Based on the historical operation data of the integrated energy system, a long short-term memory neural network is used to build an equivalent encapsulation model of the integrated energy system; the outputs of the equivalent encapsulation model of the integrated energy system are the electricity transaction volume and the heat transaction volume; Step (2): Construct a reinforcement learning environment required for learning and training the energy management strategies of each integrated energy system; Step (3): Adopt the k-priority sampling strategy and conduct online learning of the energy management strategies of each integrated energy system based on the improved deep reinforcement learning algorithm, including: Step (3-1) Initialize the experience pool and Q-network parameters: Randomly initialize the actions of the intelligent agent, record the state transition process of the intelligent agent, and store the current state, the currently taken action, the next state, and the reward function of the intelligent agent in the experience pool until the experience pool is full; meanwhile, initialize the weights of the target Q-network; Step (3-2) obtains the current environmental state s t : Take the output of wind turbines, the output of photovoltaic units, the state of distributed energy storage, the magnitude of electrical load, the magnitude of thermal load, the real-time electricity trading price, and the real-time heat trading price in the integrated energy system during the current period as the environmental state s observable by the agent t ; Step (3-3) improves the deep reinforcement learning algorithm with a k-priority sampling strategy and selects the current action a t : The k-priority sampling strategy first selects k candidate actions with the highest Q-values according to the Q-values of all actions, then calculates the normalized scores of the k candidate actions according to the softmax function, and finally completes the selection of actions according to the probability distribution that conforms to the normalized scores; The mathematical expression of the k-priority sampling strategy is: ; where s is the state of the current agent; a is the action available to the agent; is the policy function, which is used to describe the probability of selecting action a in state s; Q(s,a) is the action-value function composed of state s and action a; , is the set composed of the top k actions among all action values Q(s,a), and its expression is: ; In the formula, represents the k actions with the largest action value functions in the set of all actions; Step (3-4) Update the experience pool: The current action a obtained by executing the k-priority strategy t , obtain the state s at the next moment t+1 and the reward value r t , store the state transition process in the form of (s t , a t , r t , s t+1 ) into the experience pool. If the experience pool is already full, delete the earliest experience record. If the experience pool is not full, proceed to the next step; Step (3-5) Update the Q-network parameters: Randomly extract N data (s i , a i , r i , s i+1 ) from the experience pool, and calculate the predicted value of the target network: ; In the formula, represents the predicted value of the target network for the i-th sample, is the attenuation coefficient, is the action value function calculated by the target network under the state, represents the target network parameters; Update the Q-network parameters using the gradient descent method, and minimize the loss function as: ; In the formula, is the action value function in the state calculated by the evaluation network, represents the evaluation network parameters; Finally, repeat steps (3-2) to (3-5) until the maximum number of training times is reached.
2. The energy management method for an integrated energy system based on improved deep reinforcement learning according to claim 1, characterized in that, The specific steps of step (1) are as follows: Step (1-1) Select the input variables and output variables of the long short-term memory neural network model: The historical operation data of the integrated energy system includes the outputs of uncontrollable distributed renewable power generation units such as wind turbines and photovoltaic units, the outputs of controllable distributed power generation units such as micro gas turbines and fuel cells, electrical loads, heat loads, electricity trading prices, heat trading prices, electricity transaction volumes, and heat transaction volumes; Step (1-2) Conduct data processing, count the historical operation data of each integrated energy system, and perform data normalization and division of the training set and the test set; ; In the formula, represents the data set composed of historical operation data; represents the column vector composed of a set of all variables, represents the th day, represents the total number of days; represents the th time period in a day, is 24, indicating 24 time periods in a day; represents the normalized historical data; min(·) represents the minimum value function, and max(·) represents the maximum value function; represents the training set extracted from the historical data after normalization, represents the test set extracted from the historical data after normalization, represents the proportion of the training set in the total data set; Step (1-3) Train the long short-term memory neural network model: Adopt a long short-term memory neural network and conduct learning and training on the training set data based on the mini-batch gradient descent method of backpropagation; ; Wherein, represents the data set taken from the training data set in the th time period; represents the accumulation before the th time period; represents the output of the forgetting gate corresponding to the th time period in the current iteration, and are the weight coefficient and bias coefficient of each neuron in the forgetting layer, represents the sigmoid function, represents the output of the input layer in the th time period, and are the weight coefficient and bias coefficient of each neuron in the input layer, represents the estimated output of the convolutional layer in the th time period, and are the weight coefficient and bias coefficient of each neuron in the convolutional layer, represents the hyperbolic tangent function, represents the actual output of the convolutional layer when the th time period, represents the output of the output layer in the th time period, and are the weight coefficient and bias coefficient of each neuron in the output layer, represents the actual output when the th time period; Step (1-4) Evaluate the effect of the long short-term memory neural network model: Use the test set to test the long short-term memory neural network model, and use the root mean square error for effect evaluation; ; In the formula, represents the root mean square error between the model predicted value and the true value, represents the input variables of the network in the test set, represents the output variables of the network in the test set, represents the trained network function.
3. The integrated energy system energy management method based on improved deep reinforcement learning according to claim 2, characterized in that, The specific steps in step (2) are as follows: Step (2-1) Set the state space: Regard the control center of each integrated energy system as an intelligent agent, and the observable state space of the intelligent agent is: ; In the formula, represents the controllable observable quantity, represents the uncontrollable observable quantity, represents the observable quantity of time series information; The controllable observables include the state of charge (SoC) of distributed energy storage within the integrated energy system t , the state of temperature (SoT) of the TCL load t and the market price level C t , and the observables are shown as follows: ; The uncontrollable observables include the temperature T t , the electric energy G provided by distributed energy t , the thermal energy H provided by distributed energy t , the energy trading prices with different integrated energy systems , and the electric load and the thermal load , and the unobservable quantities are shown as follows: ; The time-series information observation quantity includes the current day t d , the current hour t h , as shown in the following formula: ; Step (2-2) Set the action space: The action space of the agent is a 10-dimensional discrete space, and this action space A includes the control of electric energy A e and the control of thermal energy A h , as shown in the following formula: ; The control actions for electricity are: ; where a tcl is the control signal of the TCL load, a l is the control information of the price-responsive electric load, a c is the charge and discharge control signal of the distributed energy storage tank, a G is the power generation control signal of the gas turbine, a p is the electricity trading price control signal, a s is the electricity trading sequence control signal; The control actions for heat are: ; Wherein, a hc is the control signal of the heat storage tank, a hG is the supplementary combustion control signal of the boiler, a hp is the heat energy trading price control signal, a hs is the heat energy trading sequence control signal; Step (2-3) Set the reward function: In order to maximize the target of the energy management plan of each integrated energy system to meet its own interests, the reward function is set as follows: ; Where S t is the revenue from selling energy, C t is the cost of obtaining energy, and Pen t is the penalty term; ; In the formula, the revenue S from selling energy t comes from the internal users of the integrated energy system and other integrated energy systems; N l is the number of internal load users in the integrated energy system, L i t is the electrical load of the i-th user at time t, L i h,t is the heat load of the i-th user at time t, P t is the electricity selling price at time t, P h,t is the heat energy selling price at time t; N a is the number of tradable integrated energy systems, P j t is the electricity selling price to the j-th integrated energy system at time t, E j t is the amount of electricity sold to the j-th integrated energy system at time t, P j h,t is the heat energy selling price to the j-th integrated energy system at time t, H j t is the amount of heat energy sold to the j-th integrated energy system at time t; ; Among them, the cost C of obtaining energy t The power generation and heat production costs from distributed energy sources and the purchase cost from other integrated energy systems; C e is the power generation cost, G t is the power generation amount of the micro gas turbine at time t, C h is the heat energy cost, H t is the heat energy provided by supplementary combustion of the boiler at time t, P k t is the electricity purchase price from the k-th integrated energy system at time t, E k t is the electricity amount purchased from the k-th integrated energy system at time t, P k h,t is the heat energy purchase price from the k-th integrated energy system at time t, H k t is the heat energy amount purchased from the k-th integrated energy system at time t; ; In the formula, is the penalty coefficient. The penalty term is always 0 at non-starting moments of each day, and is determined according to the SoC difference from the initial moment of the day at the last moment of each day.
Citation Information
Patent Citations
Energy management optimization method and device for regional integrated energy system
CN111463773A
Day-ahead random optimization scheduling method and device for integrated energy system
CN111969602A
Multi-park energy scheduling method and system based on deep reinforcement learning
CN114091879A
Multi-microgrid system collaborative optimization method based on multi-agent reinforcement learning
CN114611772A