A variable volume water heater demand response optimization method based on reinforcement learning

By using a demand response optimization method for variable capacity water heaters based on deep reinforcement learning and establishing an optimization model using the DQN algorithm, the energy management problem of water heaters under uncertain environments is solved, achieving efficient energy utilization and automatic adaptability, and reducing energy consumption and electricity waste.

CN115879362BActive Publication Date: 2026-05-12XIANGTAN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
XIANGTAN UNIV
Filing Date
2022-09-16
Publication Date
2026-05-12

AI Technical Summary

Technical Problem

Demand response optimization for water heaters needs to address the uncertainties of the equipment itself and its operating environment. Individual water heaters have low value density, and it is difficult for users to deploy human-assisted demand response after purchasing them, resulting in energy waste and insufficient flexibility.

Method used

A demand response optimization method for variable capacity water heaters based on deep reinforcement learning is adopted. An optimization model is established through the DQN algorithm. Combining user comfort temperature, sterilization temperature and safety requirements, and utilizing the probability distribution of power grid and hot water demand, the method automatically adapts to uncertainty and optimizes heating and water filling operations.

Benefits of technology

It enables efficient energy management of variable capacity water heaters under uncertain environments, reduces energy consumption, lowers electricity costs, improves users' electricity waste management, automatically adapts to differences in operating conditions, and reduces computing costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115879362B_ABST
    Figure CN115879362B_ABST
Patent Text Reader

Abstract

The application discloses a variable-volume water heater demand response optimization method based on reinforcement learning, and core steps are as follows: 1) establishing water temperature and water quantity safety constraints of variable-volume water heater operation; 2) acquiring energy price data of a future 24 hours of a power grid and predicted hot water demand of a future 24 hours of a user, generating a future energy price trend signal and a future hot water demand trend signal; 3) establishing a reinforcement learning model of variable-volume water heater demand response optimization based on DQN; and 4) maintaining actual variable-volume water heater operation according to a control action output by reinforcement learning. The method disclosed by the application meets real-time requirements, can cope with uncertainty in a dispatch process, realizes automatic optimization of variable-volume water heater demand response, and not only saves energy, but also reduces electricity charges of a user.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application relates to a power user demand response optimization method, in particular to a variable-capacity water heater demand response optimization method based on reinforcement learning, and belongs to the technical field of smart home control. BACKGROUND

[0002] Water heaters are widely used in buildings for bathing, heating, cleaning, etc., and their energy consumption accounts for 25% of the total energy consumption of the residential sector. Water heaters usually have a storage tank that can store heat, and under the premise of not affecting user comfort, can realize load interruption, shifting and reduction. Throughout the years, water heaters have been excellent demand response resources.

[0003] However, so far, most households use water heaters with the following two characteristics: fixed water volume in the storage tank and tracking artificial set temperature. Such water heaters must maintain full tank of hot water when water demand is low, which results in serious energy waste and weak flexibility of demand response. Therefore, variable-capacity water heaters have attracted attention, and the water volume in the storage tank of such water heaters can be dynamically adjusted, and the water temperature can be flexibly set.

[0004] In the demand response optimization of water heaters, the uncertainty of the device itself and its operating environment needs to be addressed. The uncertainty of the device itself mainly comes from the influence of scale on heating efficiency. Studies have shown that 1mm thick scale reduces the thermal efficiency of the water heater by 15%, and when the scale is 7mm thick, the thermal efficiency is reduced by 40%. The reasons and processes of heater scale formation are complex and have significant randomness. The causes of scale formation in general domestic water heaters are extremely complex, and detection and cleaning are difficult. The accumulation state of scale has uncertainty and is difficult to model. The uncertainty of the operating environment of the water heater lies in the inherent dynamics, randomness and prediction error of the water demand, energy price and environmental water temperature, etc. The value density of a single water heater is not high, and it is difficult to deploy artificial demand response after users purchase it. Therefore, using the pre-built optimization strategy of the device to make demand response automatically adapt to uncertainty is a prerequisite for water heaters to participate in demand response and promote social energy saving and emission reduction.

[0005] Currently, there are three main categories of demand response optimization methods for water heaters. The first category is optimization methods. This is the mainstream method for demand response optimization of electric water heaters. This method obtains decisions by establishing and solving optimization equations. However, optimization methods emphasize pre-setting solutions for various operating conditions, while variable-capacity water heaters have numerous uncertainties and their combinations, making it difficult to comprehensively cover all operating conditions in advance, thus limiting optimization accuracy. The second category is fuzzy logic rules. This method maps uncertainties to fuzzy variables and combines expert knowledge to establish fuzzy control rules, handling various possible situations and enhancing the ability to capture subjective and probabilistic uncertainties during the optimization process. However, fuzzy logic-based optimization scheduling methods often fail to achieve optimal control results. The required fuzzy rules and membership functions are derived from trial and error and expert experience, while variable-capacity water heaters are still in the development stage and have not yet been widely applied. Related research is limited, making it difficult to obtain mature fuzzy rules. The third category is reinforcement learning. Reinforcement learning is increasingly used for the optimization control of built environments. It can continuously adapt to the behavior of random occupants, time-varying environmental conditions, and system aging without requiring a strict mathematical representation of the system, and can automatically adapt to the uncertainties of the controlled system. Deep reinforcement learning builds upon this by using deep neural networks to express optimization knowledge, further enhancing its ability to cope with system uncertainties.

[0006] In summary, this patent designs a reinforcement learning-based demand response optimization method for variable capacity water heaters, which can cope with the uncertainties in the demand response optimization process of variable capacity water heaters, automatically adapt to the optimization strategy of different operating conditions of variable capacity water heaters, and realize the demand response optimization of variable capacity water heaters. Summary of the Invention

[0007] This invention addresses the problem of addressing uncertainties in demand response optimization for water heaters, particularly regarding the inherent limitations of the equipment and its operating environment. Since individual water heaters have low value density, it's difficult for users to implement manually assisted demand response after purchase. Therefore, utilizing pre-built optimization strategies to automatically adapt demand response to uncertainty is a prerequisite for water heaters to participate in demand response and promote energy conservation and emission reduction. This patent designs a reinforcement learning-based demand response optimization method for variable-capacity water heaters. This method can handle uncertainties in the demand response optimization process, automatically adapt optimization strategies to differences in the operating conditions of variable-capacity water heaters, and ultimately achieve demand response optimization for variable-capacity water heaters.

[0008] To achieve the above objectives, the technical solution of this invention is as follows: a demand response optimization method for variable capacity water heaters based on reinforcement learning, the main core technologies of which include: a deep reinforcement learning model for demand response optimization involving variable capacity water heaters, and the DQN reinforcement learning algorithm. The specific technical solution includes the following steps:

[0009] Step S1: Establish hot water comfort temperature constraints based on user water comfort temperature, establish sterilization temperature constraints based on sterilization temperature requirements, and establish water volume constraints based on the safety usage requirements of variable capacity water heaters. Among them, the sterilization temperature requirement is to maintain the water temperature of the water heater at least 60℃ or above once a day for at least 11 minutes.

[0010] Step S2: Obtain the probability distribution of energy price data for the power grid in the next 24 hours and the probability distribution of the predicted hot water demand of users in the next 24 hours. Generate a future energy price trend signal based on the energy price data and a future hot water demand trend signal based on the predicted hot water demand. The hot water demand and energy price have great uncertainty and follow a normal distribution in each interval.

[0011] Step S3: Obtain the current hot water temperature and hot water volume data of the variable capacity water heater, and combine the future energy price trend data and hot water demand trend data. With the goal of minimizing the electricity expenditure in the next 24 hours while meeting the user's comfort, sterilization temperature and safety requirements, establish a reinforcement learning model for demand response optimization of variable capacity water heater based on DQN, where DQN refers to deep Q network.

[0012] Step S4: After the reinforcement learning model for demand response optimization of variable capacity water heater based on DQN is trained to convergence, the reinforcement learning outputs the heating switch and water filling switch control actions to the variable capacity water heater to control its operation.

[0013] Preferably, step S1, which involves "establishing a hot water comfort temperature constraint based on the user's water comfort temperature, establishing a sterilization temperature constraint based on the sterilization temperature requirement, and establishing a water volume constraint based on the safety requirements for the variable capacity water heater, wherein the sterilization temperature requirement is to maintain the water temperature of the water heater at least once a day above 60°C for at least 11 minutes," includes the following steps:

[0014] Step S101: Let the control cycle of the variable capacity water heater be t, where the value of t is not greater than 2 hours and not less than 0.25 hours, and the default value is 0.5 hours; divide the next 24 hours of the scheduling into segments with t as the interval time, and number the time segments in natural number order, and record the total number of time segments as N;

[0015] Step S102: Design the constraints of hot water comfort temperature, water volume, and sterilization temperature for reinforcement learning. The hot water comfort temperature constraint is Tmin < Tn < Tmax, the water volume constraint is Vmin < Vn < Vmax, and the sterilization temperature constraint is max{Tn, n = 1, 2, …, N} > Th. Here, Tn is the hot water temperature of the water heater in the nth period, Vn is the hot water volume in the nth period. In the hot water comfort temperature constraint, Tmin is the lower limit of the hot water comfort temperature, Tmax is the upper limit of the hot water comfort temperature. In the water volume constraint, Vmin is the lower limit of the water volume, Vmax is the upper limit of the water volume. In the sterilization temperature constraint, Th ≥ 60°C and is between the upper and lower limits of the hot water comfort temperature.

[0016] Preferably, the step described in Step S2, "Obtain the probability distribution of the future 24-hour energy price data of the power grid and the probability distribution of the predicted future 24-hour hot water demand of users, generate a future energy price trend signal based on the energy price data, and generate a future hot water demand trend signal based on the predicted hot water demand. The hot water demand and the energy price have great uncertainties and follow a normal distribution in each interval" includes the following steps:

[0017] Step S201: Obtain the mean value μ of the predicted hot water demand data of users in the future 24 hours from other dedicated devices D and the mean value μ of the energy price data E , the hot water demand and the energy price have great uncertainties and follow a normal distribution in each interval respectively , , where and are the variances of the uncertain quantities, and take 0.05μ D and 0.05μ E respectively;

[0018] Step S202: In each iteration, randomly sample according to the normal distribution to obtain the uncertain hot water demand data and energy price data as D n , E n respectively; Let the length of the trend interval be m, m is an integer multiple of the control period t of the water heater, and the hot water demand trend data R D,n in the nth period is the difference between the mean value μ of the hot water demand data in the time period [n + 1, n + m] D and the hot water demand data D in the nth period n ; Similarly, the energy price trend data R E,n in the nth period is the difference between the mean value μ of the energy price data in the time period [n + 1, n + m] E and the energy price data E in the nth period n .

[0019] Preferably, step S3, "establishing a reinforcement learning model for demand response optimization of variable capacity water heaters based on DQN," includes the following steps:

[0020] Step S301: Obtain the hot water temperature T of the variable capacity water heater during n time periods. n and hot water volume V n Using the user hot water demand data D and energy price data E obtained in step S1 for the next 24 hours, the user hot water demand data and energy price data for N time periods are summed and averaged to calculate the overall user hot water demand average A for the round. D,n And the overall average energy price A E,n Then calculate the energy price deviation L over time period n. E,n The difference between the user's hot water demand and L D,n Among them, the energy price deviation L E,n Energy price data E for the nth time period n Compared with the overall average energy price A in the round E,n The difference, the deviation L of user hot water demand. D,n D represents the hot water demand data for the nth time period. n Compared with the average hot water demand of all users in the round A D,n The difference;

[0021] Step S302: Design the state variables for reinforcement learning, the state variable s for the nth time period. n Designed for s n ={n, T n V n A E A D , L E,n , L D,n , R E,n , R D,n}, including the current time period n, and the hot water temperature T of the variable capacity water heater in the nth time period. n Hot water volume V n The average overall energy price in round A E The average hot water demand of all users in the round A D Energy price deviation L E,n User hot water demand deviation L D,n Hot water demand trend data D E , n And future energy price trend data R E,n ;

[0022] Step S303: Design the action space for reinforcement learning, and the control action combination a in the nth time period. n Including the action of controlling the heating switch X 1,n And the action of controlling the water filling switch X2,n The action of controlling the heating switch X 1,n There are two states, X 1,n When the value is 1, the heater of the variable capacity water heater heats at its rated power, X 1,n When the value is 0, heating is not performed; the action of the water filling switch X is controlled. 2,n There are six states, X 2,n When X = 0, no water is added. 2,n =1 hour, add 1 gallon of water to the tank of the variable capacity water heater, X 2,n =2 hours, add 2 gallons of water to the variable capacity water heater tank, X 2,n =3 hours, add 3 gallons of water to the variable capacity water heater tank, X 2,n =4 hours, add 4 gallons of water to the tank of the variable capacity water heater, X 2,n =5 gallons of water are added to the tank of the variable capacity water heater at 5 o'clock; the different state combinations of the two control actions constitute the action space of the deep reinforcement learning model, and there are 12 action state combinations in the action space;

[0023] Step S304: Obtain the water temperature T of the variable capacity water heater during the n time periods through S301. n and hot water volume V n The constraint of hot water comfort temperature constraint C was violated. 1,n and violation of water quantity constraint mark C 2,n When T is satisfied n <T min or T max <T n At that time, the constraint of the hot water comfort temperature constraint mark C is violated. 1,n =1, otherwise C 1,n =0; when V is satisfied n <V min or V max <V n At that time, the water volume constraint sign C was violated. 2,n =1, otherwise C 2,n =0; After a round ends, the hot water temperature T for all time periods within the round is obtained. n For n=1,2,…,N, we obtain the violation of sterilization temperature constraint flag C. 3,n When max{T} is satisfied n ,n=1,2,…,N} < T h At that time, C 3,n =1, otherwise C 3,n =0;

[0024] Step S305: Design the reward function for reinforcement learning, where the immediate reward r in the nth time period is... n =-tPE n X 1,n -α (C 1,n +C2,n ), tPE n X 1,n Part of it is energy cost, α (C 1,n + C 2,n The first part represents the penalties for violating the constraints on comfortable hot water temperature and water volume, where P is the rated power of the variable capacity water heater, and E... n Let be the energy price in the nth time period, and α be the penalty coefficient, which takes a positive value and is greater than tPE. n The default value is 3tPE. n Total reward for the k-th training round , To determine whether a violation of the sterilization temperature constraint has occurred, β is the penalty coefficient. The sterilization temperature constraint can only be determined after one round of training has been completed.

[0025] Step S306: Design a deep reinforcement learning network for demand response optimization of variable capacity water heaters based on DQN. The characteristics of the designed reinforcement learning network are: the entire deep reinforcement learning network is based on DNN (Deep Neural Networks), and the deep reinforcement learning network consists of two parts: an evaluation network and a target network. The evaluation network and the target network have the same structure, and the structural characteristics of these two networks are:

[0026] 1) The network is a DNN (Deep Neural Network), which includes one input layer, four hidden layers, and one output layer, all of which are fully connected neural networks;

[0027] 2) The input layer has 9 neurons, each neuron connected one-to-one with s. n The elements in the array are n, T, and T respectively. n V n A E,n A D,n , L E,n , L D,n , R E,n , R D,n ;

[0028] 3) Hidden layer 1 has W1 neurons. Hidden layer 1 is fully connected to the input layer and hidden layer 2, but not connected to other layers. W1 is 32 by default.

[0029] 4) Hidden layer 2 has W2 neurons. Hidden layer 2 is fully connected to hidden layer 1 and hidden layer 3, but not connected to other layers. W2 is 128 by default.

[0030] 5) Hidden layer 3 has W3 neurons. Hidden layer 3 is fully connected to hidden layer 2 and hidden layer 4, but not connected to other layers. W3 is 128 by default.

[0031] 6) Hidden layer 4 has W4 neurons. Hidden layer 4 is fully connected to hidden layer 3 and the output layer, but not connected to other layers. W4 is 32 by default.

[0032] 7) The output layer has 12 neurons. The output layer is fully connected to hidden layer 4, but not connected to other layers. In the nth time interval, the 12 outputs correspond to Q(s) respectively. n ,a n |X 1,n =0, X 2,n =0), Q(s) n ,a n |X 1,n =0, X 2,n =1), Q(s) n ,a n |X 1,n =0, X 2,n =2), Q(s) n ,a n |X 1,n =0, X 2,n =3), Q(s) n ,a n |X 1,n =0, X 2,n =4), Q(s) n ,a n |X 1,n =0, X 2,n =5), Q(s) n ,a n |X 1,n =1, X 2,n =0), Q(s) n ,a n |X 1,n =1, X 2,n =1), Q(s) n ,a n |X 1,n =1, X 2,n =2), Q(s) n ,a n |X 1,n =1, X 2,n =3), Q(s) n ,a n |X 1,n =1, X 2,n =4), Q(s) n ,a n |X 1,n =1, X 2,n =5), where Q(s) n ,a n ) is in s n Control action a in staten The value expression;

[0033] Step S307: Design a deep reinforcement learning training algorithm for demand response optimization of variable capacity water heaters based on DQN. The training algorithm is characterized by performing the following steps:

[0034] Step 1: Randomly initialize the evaluation network parameters ω, and copy ω to the target network parameters. ω Based on steps S2 and S3, the current status s of the variable capacity water heater is obtained. n Action a selected through the ε-greedy strategy n * The calculated reward r n And the state s at the next moment n+1 Generate reinforcement learning samples (s) n ,a n * ,r n ,s n+1 Add samples to the sample pool, and extract M samples for training each time. M defaults to 64. Initialize the working period n=1; initialize the number of training rounds k=1.

[0035] Step 2: Selecting dominant strategy actions and evaluating network convergence. The method is to perform the following steps in sequence:

[0036] 1) Acquire and record the dominant policy control actions of the evaluation network. If the current time period is n, then the dominant policy control action a for the nth time period is... n For: a n =argmaxQ(s n ), where argmax represents taking s n Determine the action corresponding to the maximum Q value under a given state, and calculate the immediate reward r in the nth time period. n Its formula is r n =-tPE n X 1,n -α (C 1,n + C 2,n );

[0037] 2) The convergence of the network is evaluated by calculating the total reward of the dominant policy action in each training round and the total reward of the dominant policy control action in the k-th training round. If there are L consecutive training rounds R k If the variance of L is less than φ and no default occurs, the evaluation network is considered to have converged; otherwise, the evaluation network is considered to have not converged. Here, φ is a natural number not less than 10, with a default value of 15; L is a natural number not less than 10, with a default value of 20.

[0038] 3) If the network has converged in the nth time interval, then exit the reinforcement learning training and set a n Send to the actual variable capacity water heater to control its operation;

[0039] 4) If the network fails to converge in the nth time period, proceed to step 3;

[0040] Step 3: Obtain M samples from the sample pool, where M is 64 by default;

[0041] Step 4: Explore and utilize coordination, which is achieved by employing an ε-greedy strategy. At the start of training, the probabilities are initialized, and ε is set to... t =ε in As the number of training rounds k increases, ε t The value is set in step size ε. step Decrease until it reaches the set minimum value ε. min , where ε t The calculation formula is With probability ε t Choose a random action from the action space with probability 1-ε t The action with the highest Q-value obtained from the evaluation network is selected as action a. n * ;ε min The value is greater than 0;

[0042] Step 5: Calculate the i-th sample (s) based on the output of the target network. i ,a i ,r i ,s i+1 The target value y) i Its calculation formula is Where γ is the discount factor, taking values ​​in the range (0,1), and the four elements in (si,ai,ri,si+1) correspond to the state, action, immediate reward, and next state of the i-th sample, respectively. Is the target network at input s i+1 Output at time;

[0043] Step 6: Calculate the loss function Loss, the formula of which is: , where x is the number of samples used in one training iteration, It is to evaluate the network on the input s i At that time, action a i The corresponding output;

[0044] Step 7: ω t and ω t+1 These are the evaluation network parameter sets before and after the update, l rThe learning rate is used to update the network parameters using the gradient backpropagation algorithm. The network parameter update formula is: ;

[0045] Step 8: Let n = n + 1; if n > N, then proceed to step 9; otherwise, go back to step 2.

[0046] Step 9: Let n=1, k=k+1. When the result of k / Z is an integer, update the parameters of the target network using the parameters ω of the evaluation network. ω If the result of k / Z is not an integer, then jump directly to step 2; where Z is the target network update period, which is a positive integer not less than 10, and the default value is 50.

[0047] Compared with the prior art, the present invention has the following advantages:

[0048] 1) This invention provides a demand response optimization method for variable capacity water heaters, which can reduce energy consumption, improve the situation of electricity waste for users, and reduce electricity expenditure; 2) Compared with other optimization methods, the demand response optimization method for variable capacity water heaters based on reinforcement learning proposed in this invention can better cope with uncertainty and automatically adapt to the optimization strategy of different operating conditions of variable capacity water heaters; 3) This invention does not require modeling, only needs to obtain existing data, and has low computational cost. Attached Figure Description

[0049] Figure 1 Flowchart of a reinforcement learning-based demand response optimization method for variable capacity water heaters

[0050] Figure 2 This is a schematic diagram of a reinforcement learning model for demand response optimization of variable capacity water heaters based on DQN.

[0051] Figure 3 This is a diagram illustrating the working principle of a convergent reinforcement learning model governing a variable capacity water heater.

[0052] Figure 4 This is a structural diagram of the target network and the evaluation network. Detailed Implementation

[0053] The invention will be further described below with reference to the accompanying drawings. The core steps of the invention are: 1) Establishing safety constraints on water temperature and water volume for the operation of the variable-capacity water heater; 2) Obtaining energy price data from the power grid for the next 24 hours and predicting the user's hot water demand for the next 24 hours, generating future energy price trend signals and future hot water demand trend signals; 3) Establishing a reinforcement learning model based on DQN for demand response optimization of the variable-capacity water heater; 4) Maintaining the actual operation of the variable-capacity water heater according to the control actions output by the reinforcement learning. The flowchart of steps 1)-4) is shown below. Figure 1 As shown, step 3) is the reinforcement model of demand response for variable capacity water heaters based on DQN.Figure 2 As shown, step 4) controls the actual variable capacity water heater working principle diagram according to the optimization result output by reinforcement learning as Figure 4 shown.

[0054] Based on Figure 1 the flowchart, the specific implementation of the present invention includes the following steps:

[0055] Step S1: Establish a hot water comfort temperature constraint according to the user's comfortable water temperature for use, establish a sterilization temperature constraint according to the sterilization temperature requirement, and establish a water volume constraint according to the safe use requirements of the variable capacity water heater. Among them, the sterilization temperature requirement is to keep the water temperature of the water heater at least once a day above 60 °C and the duration is more than 11 minutes. Specifically, it includes the following steps:

[0056] Step S101: Set the control period of the variable capacity water heater as t, where the value of t is not greater than 2 hours and not less than 0.25 hours, and the default is 0.5 hours; segment the future 24 hours participating in the scheduling at intervals of t, number the time periods in sequence according to the natural numbers, and record the total number of time periods as N;

[0057] Step S102: Design the hot water comfort temperature constraint, water volume constraint and sterilization temperature constraint of reinforcement learning. The hot water comfort temperature constraint is Tmin <Tn< Tmax, the water volume constraint is Vmin <Vn< Vmax, and the sterilization temperature constraint is max{Tn,n=1,2,…,N} > Th. Among them, Tn is the hot water temperature of the water heater in the nth time period, Vn is the hot water volume in the nth time period. In the hot water comfort temperature constraint, Tmin is the lower limit of the hot water comfort temperature, Tmax is the upper limit of the hot water comfort temperature. In the water volume constraint, Vmin is the lower limit of the water volume, Vmax is the upper limit of the water volume. In the sterilization temperature constraint, Th ≥60 °C and is between the upper and lower limits of the hot water comfort temperature.

[0058] Step S2: Obtain the probability distribution of the future 24-hour energy price data of the power grid and the probability distribution of the predicted future 24-hour hot water demand of users, generate a future energy price trend signal according to the energy price data, and generate a future hot water demand trend signal according to the predicted hot water demand. Among them, the hot water demand and energy price have great uncertainties and follow a normal distribution in each interval; specifically, it includes the following steps:

[0059] Step S201: Obtain the mean value μ D of the predicted future 24-hour hot water demand data of users and the mean value μ E of the energy price data from other special equipment. The hot water demand and energy price have great uncertainties and follow a normal distribution respectively in each interval , , among which and The variance of the uncertainty is taken as 0.05 μ. D and 0.05μ E ;

[0060] Step S202: In each iteration, randomly sample according to a normal distribution to obtain uncertain hot water demand data and energy price data, respectively D. n E n Let the length of the trend interval be m, where m is an integer multiple of the water heater control cycle t. Let R be the trend data of hot water demand in the nth time period. D,n Let μ be the mean value of hot water demand data over the time period [n+1, n+m]. D Hot water demand data D for time period n n The difference; similarly, the energy price trend data R for the nth time period. E,n Let μ be the mean of energy price data over the time period [n+1, n+m]. E Energy price data E for time period n n The difference.

[0061] Step S3: Obtain current hot water temperature and volume data for the variable capacity water heater. Integrate future energy price trends and hot water demand trends, and with the goal of minimizing electricity costs over the next 24 hours while meeting user comfort, sterilization temperature, and safety requirements, establish a reinforcement learning model for demand response optimization of the variable capacity water heater based on Deep Q Network (DQN). The established reinforcement learning model for demand response optimization of the variable capacity water heater based on DQN is as follows: Figure 3 As shown, the specific steps include:

[0062] Step S301: Obtain the hot water temperature T of the variable capacity water heater during n time periods. n and hot water volume V n Using the user hot water demand data D and energy price data E obtained in step S1 for the next 24 hours, the user hot water demand data and energy price data for N time periods are summed and averaged to calculate the overall user hot water demand average A for the round. D,n And the overall average energy price A E,n Then calculate the energy price deviation L over time period n. E,n The difference between the user's hot water demand and L D,n Among them, the energy price deviation L E,n Energy price data E for the nth time period n Compared with the overall average energy price A in the round E,n The difference, the deviation L of user hot water demand. D,n D represents the hot water demand data for the nth time period. n Compared with the average hot water demand of all users in the round AD,n The difference;

[0063] Step S302: Design the state variables for reinforcement learning, the state variable s for the nth time period. n Designed for s n ={n, T n V n A E A D , L E,n , L D,n , R E,n , R D,n}, including the current time period n, and the hot water temperature T of the variable capacity water heater in the nth time period. n Hot water volume V n The average overall energy price in round A E The average hot water demand of all users in the round A D Energy price deviation L E,n User hot water demand deviation L D,n Hot water demand trend data D E , n And future energy price trend data R E,n ;

[0064] Step S303: Design the action space for reinforcement learning, and the control action combination a in the nth time period. n Including the action of controlling the heating switch X 1,n And the action of controlling the water filling switch X 2,n The action of controlling the heating switch X 1,n There are two states, X 1,n When the value is 1, the heater of the variable capacity water heater heats at its rated power, X 1,n When the value is 0, heating is not performed; the action of the water filling switch X is controlled. 2,n There are six states, X 2,n When X = 0, no water is added. 2,n =1 hour, add 1 gallon of water to the tank of the variable capacity water heater, X 2,n =2 hours, add 2 gallons of water to the variable capacity water heater tank, X 2,n =3 hours, add 3 gallons of water to the variable capacity water heater tank, X 2,n =4 hours, add 4 gallons of water to the tank of the variable capacity water heater, X 2,n =5 gallons of water are added to the tank of the variable capacity water heater at 5 o'clock; the different state combinations of the two control actions constitute the action space of the deep reinforcement learning model, and there are 12 action state combinations in the action space;

[0065] Step S304: Obtain the water temperature T of the variable capacity water heater during the n time periods through S301. n and hot water volume V nThe constraint of hot water comfort temperature constraint C was violated. 1,n and violation of water quantity constraint mark C 2,n When T is satisfied n <T min or T max <T n At that time, the constraint of the hot water comfort temperature constraint mark C is violated. 1,n =1, otherwise C 1,n =0; when V is satisfied n <V min or V max <V n At that time, the water volume constraint sign C was violated. 2,n =1, otherwise C 2,n =0; After a round ends, the hot water temperature T for all time periods within the round is obtained. n For n=1,2,…,N, we obtain the violation of sterilization temperature constraint flag C. 3,n When max{T} is satisfied n ,n=1,2,…,N} < T h At that time, C 3,n =1, otherwise C 3,n =0;

[0066] Step S305: Design the reward function for reinforcement learning, where the immediate reward r in the nth time period is... n =-tPE n X 1,n -α (C 1,n +C 2,n ), tPE n X 1,n Part of it is energy cost, α (C 1,n + C 2,n The first part represents the penalties for violating the constraints on comfortable hot water temperature and water volume, where P is the rated power of the variable capacity water heater, and E... n Let be the energy price in the nth time period, and α be the penalty coefficient, which takes a positive value and is greater than tPE. n The default value is 3tPE. n Total reward for the k-th training round , To determine whether a violation of the sterilization temperature constraint has occurred, β is the penalty coefficient. The sterilization temperature constraint can only be determined after one round of training has been completed.

[0067] Step S306: Design a deep reinforcement learning network for demand response optimization of variable capacity water heaters based on DQN. The designed reinforcement learning network is characterized by the following: the entire deep reinforcement learning network is based on DNN (Deep Neural Networks), and the deep reinforcement learning network consists of two parts: an evaluation network and a target network. The evaluation network and the target network have the same structure, such as... Figure 3 As shown, the structural characteristics of both networks are:

[0068] 1) The network is a DNN (Deep Neural Network), which includes one input layer, four hidden layers, and one output layer, all of which are fully connected neural networks;

[0069] 2) The input layer has 9 neurons, each neuron connected one-to-one with s. n The elements in the array are n, T, and T respectively. n V n A E,n A D,n , L E,n , L D,n , R E,n , R D,n ;

[0070] 3) Hidden layer 1 has W1 neurons. Hidden layer 1 is fully connected to the input layer and hidden layer 2, but not connected to other layers. W1 is 32 by default.

[0071] 4) Hidden layer 2 has W2 neurons. Hidden layer 2 is fully connected to hidden layer 1 and hidden layer 3, but not connected to other layers. W2 is 128 by default.

[0072] 5) Hidden layer 3 has W3 neurons. Hidden layer 3 is fully connected to hidden layer 2 and hidden layer 4, but not connected to other layers. W3 is 128 by default.

[0073] 6) Hidden layer 4 has W4 neurons. Hidden layer 4 is fully connected to hidden layer 3 and the output layer, but not connected to other layers. W4 is 32 by default.

[0074] 7) The output layer has 12 neurons. The output layer is fully connected to hidden layer 4, but not connected to other layers. In the nth time interval, the 12 outputs correspond to Q(s) respectively. n ,a n |X 1,n =0, X 2,n =0), Q(s) n ,a n |X 1,n =0, X 2,n =1), Q(s) n ,an |X 1,n =0, X 2,n =2), Q(s) n ,a n |X 1,n =0, X 2,n =3), Q(s) n ,a n |X 1,n =0, X 2,n =4), Q(s) n ,a n |X 1,n =0, X 2,n =5), Q(s) n ,a n |X 1,n =1, X 2,n =0), Q(s) n ,a n |X 1,n =1, X 2,n =1), Q(s) n ,a n |X 1,n =1, X 2,n =2), Q(s) n ,a n |X 1,n =1, X 2,n =3), Q(s) n ,a n |X 1,n =1, X 2,n =4), Q(s) n ,a n |X 1,n =1, X 2,n =5), where Q(s) n ,a n ) is in s n Control action a in state n The value expression;

[0075] Step S307: Deep reinforcement learning training algorithm for a reinforcement learning model of demand response optimization for variable capacity water heaters based on DQN design. The training algorithm is characterized by performing the following steps:

[0076] Step 1: Randomly initialize the evaluation network parameters ω, and copy ω to the target network parameters. ω Based on steps S2 and S3, the current status s of the variable capacity water heater is obtained. n Action a selected through the ε-greedy strategy n * The calculated reward r n And the state s at the next momentn+1 Generate reinforcement learning samples (s) n ,a n * ,r n ,s n+1 Add samples to the sample pool, and extract M samples for training each time. M defaults to 64. Initialize the working period n=1; initialize the number of training rounds k=1.

[0077] Step 2: Selecting dominant strategy actions and evaluating network convergence. The method is to perform the following steps in sequence:

[0078] 1) Acquire and record the dominant policy control actions of the evaluation network. If the current time period is n, then the dominant policy control action a for the nth time period is... n For: a n =argmaxQ(s n ), where argmax represents taking s n Determine the action corresponding to the maximum Q value under a given state, and calculate the immediate reward r in the nth time period. n Its formula is r n =-tPE n X 1,n -α (C 1,n + C 2,n );

[0079] 2) The convergence of the network is evaluated by calculating the total reward of the dominant policy action in each training round and the total reward of the dominant policy control action in the k-th training round. If there are L consecutive training rounds R k If the variance of L is less than φ and no default occurs, the evaluation network is considered to have converged; otherwise, the evaluation network is considered to have not converged. Here, φ is a natural number not less than 10, with a default value of 15; L is a natural number not less than 10, with a default value of 20.

[0080] 3) If the network has converged in the nth time interval, then exit the reinforcement learning training and set a n Send to the actual variable capacity water heater to control its operation;

[0081] 4) If the network fails to converge in the nth time period, proceed to step 3;

[0082] Step 3: Obtain M samples from the sample pool, where M is 64 by default;

[0083] Step 4: Explore and utilize coordination, which is achieved by employing an ε-greedy strategy. At the start of training, the probabilities are initialized, and ε is set to... t =ε in As the number of training rounds k increases, ε t The value is set in step size ε.step Decrease until it reaches the set minimum value ε. min , where ε t The calculation formula is With probability ε t Choose a random action from the action space with probability 1-ε t The action with the highest Q-value obtained from the evaluation network is selected as action a. n * ;ε min The value is greater than 0;

[0084] Step 5: Calculate the i-th sample (s) based on the output of the target network. i ,a i ,r i ,s i+1 The target value y) i Its calculation formula is , where γ is the discount factor, taking values ​​in the range (0,1), (s i ,a i ,r i ,s i+1 The four elements in the formula correspond to the state, action, immediate response, and next state of the i-th sample, respectively. Is the target network at input s i+1 Output at time;

[0085] Step 6: Calculate the loss function Loss, the formula of which is: , where x is the number of samples used in one training iteration, It is to evaluate the network on the input s i At that time, action a i The corresponding output;

[0086] Step 7: ω t and ω t+1 These are the evaluation network parameter sets before and after the update, l r The learning rate is used to update the network parameters using the gradient backpropagation algorithm. The network parameter update formula is: ;

[0087] Step 8: Let n = n + 1; if n > N, then proceed to step 9; otherwise, go back to step 2.

[0088] Step 9: Let n=1, k=k+1. When the result of k / Z is an integer, update the parameters of the target network using the parameters ω of the evaluation network. ω If the result of k / Z is not an integer, then jump directly to step 2; where Z is the target network update period, which is a positive integer not less than 10, and the default value is 50.

[0089] Step S4: After the reinforcement learning model for demand response optimization of variable capacity water heaters based on DQN has converged, the reinforcement learning outputs the heating switch and water filling switch control actions to the variable capacity water heater, thus controlling its operation. Figure 4 As shown.

Claims

1. A demand response optimization method for variable capacity water heaters based on reinforcement learning, characterized in that, Includes the following steps: Step S1: Establish hot water comfort temperature constraints based on user water comfort temperature, establish sterilization temperature constraints based on sterilization temperature requirements, and establish water volume constraints based on the safety usage requirements of variable capacity water heaters. Among them, the sterilization temperature requirement is to maintain the water temperature of the water heater at least 60℃ or above once a day for at least 11 minutes. Step S2: Obtain the probability distribution of energy price data for the power grid for the next 24 hours and the probability distribution of predicted hot water demand from users for the next 24 hours. Generate a future energy price trend signal based on the energy price data and a future hot water demand trend signal based on the predicted hot water demand. This includes the following steps: Step S201: Obtain the average value μ of the predicted user hot water demand data for the next 24 hours from other dedicated equipment. D and the mean of energy price data μ E The demand for hot water and energy prices are highly uncertain, and they follow a normal distribution in each interval. , ,in and The variance of the uncertainty is taken as 0.05 μ. D and 0.05μ E ; Step S202: In each iteration, randomly sample according to a normal distribution to obtain uncertain hot water demand data and energy price data, respectively D. n E n Let the length of the trend interval be m, where m is an integer multiple of the water heater control cycle t. Let R be the trend data of hot water demand in the nth time period. D,n Let μ be the mean value of hot water demand data over the time period [n+1, n+m]. D Hot water demand data D for time period n n The difference; similarly, the energy price trend data R for the nth time period. E,n Let μ be the mean of energy price data over the time period [n+1, n+m]. E Energy price data E for time period n n The difference; Step S3: Includes the following steps: Step S301: Obtain the hot water temperature T of the variable capacity water heater during n time periods. n and hot water volume V n Using the user hot water demand data D and energy price data E obtained in step S2 for the next 24 hours, the user hot water demand data and energy price data for N time periods are summed and averaged to calculate the overall user hot water demand average A for the round. D And the overall average energy price A E Then calculate the energy price deviation L over time period n. E,n The difference between the user's hot water demand and L D,n Among them, the energy price deviation L E,n Energy price data E for the nth time period n Compared with the overall average energy price A in the round E,n The difference, the deviation L of user hot water demand. D,n D represents the hot water demand data for the nth time period. n Compared with the average hot water demand of all users in the round A D,n The difference; Step S302: Design the state variables for reinforcement learning, the state variable s for the nth time period. n Designed for s n ={n, T n V n A E A D ,L E,n , L D,n , R E,n , R D,n }, including the current time period n, and the hot water temperature T of the variable capacity water heater in the nth time period. n Hot water volume V n The average overall energy price in round A E The average hot water demand of all users in the round A D Energy price deviation L E,n User hot water demand deviation L D,n Hot water demand trend data D E , n And future energy price trend data R E,n ; Step S303: Design the action space for reinforcement learning, and the control action combination a in the nth time period. n Including the action of controlling the heating switch X 1,n And the action of controlling the water filling switch X 2,n ; Step S304: Obtain the water temperature T of the variable capacity water heater during the n-time period through S301. n and hot water volume V n The constraint of hot water comfort temperature constraint C was violated. 1,n and violation of water quantity constraint mark C 2,n When T is satisfied n <T min or T max <T n At that time, the constraint of the hot water comfort temperature constraint mark C is violated. 1,n =1, otherwise C 1,n =0; when V is satisfied n <V min or V max <V n At that time, the water volume constraint sign C was violated. 2,n =1, otherwise C 2,n =0; After a round ends, the hot water temperature T for all time periods within the round is obtained. n For n=1,2,…,N, the violation of sterilization temperature constraint C is obtained. 3,n When max{T} is satisfied n ,n=1,2,…,N} < T h At that time, C 3,n =1, otherwise C 3,n =0; Step S305: Design the reward function for reinforcement learning, where the immediate reward r in the nth time period is... n =-tPE n X 1,n -α (C 1,n + C 2,n ), tPE n X 1,n Part of it is energy cost, where t is a control period, α (C 1,n + C 2,n The first part represents the penalties for violating the constraints on comfortable hot water temperature and water volume, where P is the rated power of the variable capacity water heater, and E... n Let α be the energy price in time period n, and α be the penalty coefficient, which takes a positive value and is greater than tPE. n The default value is 3tPE. n Total reward for the k-th training round , To determine whether a violation of the sterilization temperature constraint has occurred, β is the penalty coefficient. The sterilization temperature constraint can only be determined after one round of training has been completed. Step S306: Establish a deep reinforcement learning network based on DQN for demand response optimization of variable capacity water heaters; Step S307: Establish a deep reinforcement learning training algorithm for demand response optimization of variable capacity water heaters based on DQN; Step S4: After the reinforcement learning model for demand response optimization of variable capacity water heater based on DQN is trained to convergence, the reinforcement learning outputs the heating switch and water filling switch control actions to the variable capacity water heater to control its operation.

2. The method for optimizing the demand response of a variable-capacity water heater based on reinforcement learning according to claim 1, characterized in that, Step S1, which involves "establishing hot water comfort temperature constraints based on user water comfort temperature, establishing sterilization temperature constraints based on sterilization temperature requirements, and establishing water volume constraints based on the safety requirements for variable capacity water heaters, wherein the sterilization temperature requirement is to maintain the water temperature of the water heater at least once a day above 60°C for at least 11 minutes," includes the following steps: Step S101: Let the control cycle of the variable capacity water heater be t, where the value of t is not greater than 2 hours and not less than 0.25 hours, and the default value is 0.5 hours; divide the next 24 hours of the scheduling into segments with t as the interval time, and number the time segments in natural number order, and record the total number of time segments as N; Step S102: Design reinforcement learning constraints for hot water comfort temperature, water flow rate, and sterilization temperature. The hot water comfort temperature constraint is T. min <T n < T max Water volume constraint is V min <V n < V max The sterilization temperature constraint is max{T} n n=1,2,…,N}> T h , among which, T n V represents the water temperature of the water heater during time period n. n Given the hot water volume over time period n and the comfortable hot water temperature constraint, T min T is the lower limit of the comfortable temperature for hot water. max V represents the upper limit of comfortable hot water temperature, with water volume constraints. min V is the lower limit of water volume. max The upper limit of water volume and the sterilization temperature constraint are T. h ≥60℃, which is between the upper and lower limits of the comfortable hot water temperature.

3. The method for optimizing the demand response of a variable-capacity water heater based on reinforcement learning according to claim 1, characterized in that, The "control action combination a" in step S303 n Including the action of controlling the heating switch X 1,n And the action of controlling the water filling switch X 2,n It has the following characteristics: X controls the action of the heating switch. 1,n There are two states, X 1,n When the value is 1, the heater of the variable capacity water heater heats at its rated power, X 1,n When the value is 0, heating is not performed; the action of the water filling switch X is controlled. 2,n There are six states, X 2,n When X = 0, no water is added. 2,n =1 hour, add 1 gallon of water to the tank of the variable capacity water heater, X 2,n =2 hours, add 2 gallons of water to the variable capacity water heater tank, X 2,n =3 hours, add 3 gallons of water to the variable capacity water heater tank, X 2,n =4 hours, add 4 gallons of water to the tank of the variable capacity water heater, X 2,n =5 hours to add 5 gallons of water to the tank of the variable capacity water heater; the different state combinations of the two control actions constitute the action space of the deep reinforcement learning model, and there are 12 action state combinations in the action space.

4. The method for optimizing the demand response of a variable-capacity water heater based on reinforcement learning according to claim 1, characterized in that, The deep reinforcement learning network established in step S306 consists of two parts: an evaluation network and a target network. The evaluation network and the target network have the same structure, and the structural features of these two networks are as follows: 1) The network is a DNN (Deep Neural Network), which includes one input layer, four hidden layers, and one output layer, all of which are fully connected neural networks; 2) The input layer has 9 neurons, each neuron connected one-to-one with s. n The elements in the array are n, T, and T respectively. n V n A E,n A D,n , L E,n , L D,n , R E,n , R D,n ; 3) Hidden layer 1 has W1 neurons. Hidden layer 1 is fully connected to the input layer and hidden layer 2, but not connected to other layers. W1 is 32 by default. 4) Hidden layer 2 has W2 neurons. Hidden layer 2 is fully connected to hidden layer 1 and hidden layer 3, but not connected to other layers. W2 is 128 by default. 5) Hidden layer 3 has W3 neurons. Hidden layer 3 is fully connected to hidden layer 2 and hidden layer 4, but not connected to other layers. W3 is 128 by default. 6) Hidden layer 4 has W4 neurons. Hidden layer 4 is fully connected to hidden layer 3 and the output layer, but not connected to other layers. W4 is 32 by default. 7) The output layer has 12 neurons. The output layer is fully connected to the hidden layer 4, but not connected to other layers. In the nth time interval, the 12 outputs correspond to Q(s) respectively. n ,a n |X 1,n =0, X 2,n =0), Q(s) n ,a n |X 1,n =0, X 2,n =1), Q(s) n ,a n |X 1,n =0, X 2,n =2), Q(s) n ,a n |X 1,n =0, X 2,n =3), Q(s) n ,a n |X 1,n =0, X 2,n =4), Q(s) n ,a n |X 1,n =0, X 2,n =5), Q(s) n ,a n |X 1,n =1,X 2,n =0), Q(s) n ,a n |X 1,n =1,X 2,n =1), Q(s) n ,a n |X 1,n =1,X 2,n =2), Q(s) n ,a n |X 1,n =1,X 2,n =3), Q(s) n ,a n |X 1,n =1,X 2,n =4), Q(s) n ,a n |X 1,n =1,X 2,n =5), where Q(s) n ,a n ) is in s n Control action a in state n The value expression of .

5. The method for optimizing the demand response of a variable-capacity water heater based on reinforcement learning according to claim 1, characterized in that, The deep reinforcement learning training algorithm established in step S307 includes the following steps: Step 1: Randomly initialize the evaluation network parameters ω, and copy ω to the target network parameters ω; based on steps S2 and S3, obtain the current state s of the variable capacity water heater. n Action a selected through the ε-greedy strategy n * The calculated reward r n And the state s at the next moment n+1 Generate reinforcement learning samples (s) n ,a n * ,r n ,s n+1 Add samples to the sample pool, and extract M samples for training each time. M defaults to 64. Initialize the working period n=1; initialize the number of training rounds k=1. Step 2: Selecting dominant strategy actions and evaluating network convergence. The method is to perform the following steps in sequence: 1) Acquire and record the dominant policy control actions of the evaluation network. If the current time period is n, then the dominant policy control action a for the nth time period is... n For: a n =argmaxQ(s n ), where argmax represents taking s n Determine the action corresponding to the maximum Q value under a given state, and calculate the immediate reward r in the nth time period. n Its formula is r n =-tPE n X 1,n -α (C 1,n + C 2,n ); 2) The convergence of the network is evaluated by calculating the total reward of the dominant policy action in each training round and the total reward of the dominant policy control action in the k-th training round. If there are L consecutive training rounds R k If the variance of L is less than φ and no default occurs, the evaluation network is considered to have converged; otherwise, the evaluation network is considered to have not converged. Here, φ is a natural number not less than 10, with a default value of 15; L is a natural number not less than 10, with a default value of 20. 3) If the network has converged in the nth time interval, then exit the reinforcement learning training and set a n Send to the actual variable capacity water heater to control its operation; 4) If the network fails to converge in the nth time period, proceed to step 3; Step 3: Obtain M samples from the sample pool, where M is 64 by default; Step 4: Explore and utilize coordination, which is achieved by employing an ε-greedy strategy. At the start of training, the probabilities are initialized, and ε is set to... t =ε in As the number of training rounds k increases, ε t The value is set in step size ε. step Decrease until it reaches the set minimum value ε. min , where ε t The calculation formula is With probability ε t Choose a random action from the action space with probability 1-ε t The action with the highest Q-value obtained from the evaluation network is selected as action a. n * ;ε min The value is greater than 0; Step 5: Calculate the i-th sample (s) based on the output of the target network. i ,a i ,r i ,s i+1 The target value y) i Its calculation formula is , where γ is the discount factor, taking values ​​in the range (0,1), (s i ,a i ,r i ,s i+1 The four elements in the equation correspond to the state, action, immediate response, and next state of the i-th sample, respectively. Is the target network at input s i+1 Output at time; Step 6: Calculate the loss function Loss, the formula of which is: , where x is the number of samples used in one training iteration, It is to evaluate the network on the input s i At that time, action a i The corresponding output; Step 7: and These are the evaluation network parameter sets before and after the update, l r The learning rate is used to update the network parameters using the gradient backpropagation algorithm. The network parameter update formula is: ; Step 8: Let n = n + 1; if n > N, then proceed to step 9; otherwise, go back to step 2. Step 9: Let n=1, k=k+1. When the result of k / Z is an integer, update the parameter ω of the target network with the parameter ω of the evaluation network and jump to step 2; when the result of k / Z is not an integer, jump directly to step 2; where Z is the update period of the target network, which is a positive integer not less than 10, and the default value is 50.