A Model-Free Reinforcement Learning Method for Energy Demand Response Management

By constructing a model-free reinforcement learning method, combining dispatchable and undispatchable electrical appliances and plug-in electric vehicles, and employing multiple algorithms to optimize retail prices, the complexity of electricity demand planning and model inaccuracy in residential demand response management are solved, achieving efficient electricity demand response under dynamic electricity pricing.

CN116227806BActive Publication Date: 2025-10-31SOUTHWEST UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211562407.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-12-07
Publication Date
2025-10-31
Estimated Expiration
2042-12-07

AI Technical Summary

Technical Problem

Existing residential demand response management systems are unable to effectively cope with dynamic electricity prices and the uncertainty of electricity consumption by appliances, resulting in complex electricity demand planning and inaccurate models, which cannot efficiently respond to changes in the electricity market.

Method used

A model-free reinforcement learning method is constructed, combining schedulable and unschedulable electrical appliances and plug-in electric vehicles. The Q-table, Q-network algorithm combining deep learning and Q-learning, and Actor-Critic algorithm are used to optimize retail prices to balance the interests of residents and retailers and maximize social welfare.

Benefits of technology

By integrating the interests of residents and retailers and optimizing the retail price series, the complexities of residential demand response management are resolved, enabling efficient electricity demand planning in an uncertain electricity market environment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116227806B_ABST
    Figure CN116227806B_ABST
Patent Text Reader

Abstract

This invention provides a model-free reinforcement learning method for energy demand response management, comprising: constructing a residential appliance model; determining social welfare by combining the comprehensive cost of residential electricity consumption and the profit of electricity retailers; balancing the comprehensive cost of residential electricity consumption and the profit of retailers based on the social welfare, wherein the social welfare is represented as a price-based non-convex optimization problem for residential demand response management; constructing reinforcement learning solutions for the price-based non-convex optimization problem for residential demand response management using Q-learning algorithms based on Q-tables, Q-network algorithms combining deep learning and Q-learning, and Actor-Critic algorithms, respectively, based on the transmission data of the power grid; and determining the optimal solution and the optimal retail price sequence based on the three reinforcement learning solutions for the price-based non-convex optimization problem for residential demand response management. This invention can use three algorithms to perform modeling respectively, achieving optimal retail price planning under unknown electricity market conditions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of energy demand response management technology, and in particular to a model-free reinforcement learning method based on energy demand response management. Background Technology

[0002] With the rapid development of smart grids as a typical cyber-physical system in the information age, home energy management systems are one of the key technologies for deploying energy demand response. As a crucial component of home energy management systems, residential demand response management aims to utilize changes in load energy use to respond to time-varying electricity prices or incentive measures to achieve cost reduction or other benefits. However, due to the randomness and elasticity of residential electricity consumption, developing effective residential demand response management strategies for households is very challenging. Specifically, due to residents' living habits, the timing and frequency of appliance on / off are uncertain and difficult to predict. The complexity of residential demand response management increases further when appliances are further classified as dispatchable or undispatchable based on the transferability of their energy consumption. All of these factors make it difficult for residential demand response management to effectively plan the timing of electricity demand in response to dynamic electricity prices. Furthermore, to achieve effective load operation, it is necessary to determine accurate appliance models and parameters in a timely manner to simulate the power characteristics and operational dynamics of these appliances. However, ordinary households do not always have access to this professional knowledge.

[0003] To address the aforementioned difficulties in residential demand response management (DRT), scholars have proposed a series of methods. Early DRT efforts primarily focused on minimizing household electricity costs. For example, some combined mixed-integer linear programming models with appliance demand response to reduce daily household energy consumption, but failed to consider the elasticity of appliance usage and dynamic electricity pricing; others minimized worst-case daily bill payments by considering the uncertainty of consumer behavior. Simultaneously, a chance-constrained optimization model was developed to ensure the probabilistic satisfaction of appliance operating constraints. Generally, price-based DRT encourages loads to adjust their energy consumption according to time-based pricing mechanisms; common strategies include real-time pricing and usage-time pricing.

[0004] From a benefit perspective, current research on price-based residential demand response management can be divided into three parts. For individual interests, there's a tendency to choose appropriate pricing mechanisms to reduce electricity costs or bring other benefits to customers. For energy companies, maximizing company profits or minimizing generation costs is their goal. However, to meet the reasonable needs of social development, integrating the relative interests of both has become a trend, thus maximizing social benefits has become a new research hotspot.

[0005] However, unresolved issues remain in the aforementioned work: 1) The system needs to identify steps, i.e., explicitly optimize the model, predictor, and solver. Developing model-based demand response strategies requires model building and parameter identification; performance may degrade due to model inaccuracies. 2) Existing price-based residential demand response management works largely rely on deterministic pricing models, such as time-based pricing and real-time pricing, which fail to reflect the uncertainty and flexibility of dynamic electricity markets. 3) The short-sightedness of the power grid leads to a focus only on the immediate response of loads to current pricing strategies, without predicting the impact of all subsequent responses. Therefore, developing a method based on models of unknown residential environments to address price-based residential demand response management problems in smart grids is of great significance.

[0006] In recent years, reinforcement learning has been widely applied in industry. It can overcome the aforementioned problems by leveraging the end-to-end learning capabilities of neural networks and has achieved significant success in many complex decision-making applications, such as distributed economic dispatch in smart grids. As one of the energy dispatch problems, this model-free reinforcement learning algorithm has inspired researchers to study reinforcement learning-based residential demand response management. For example, developing a community smart home energy management scheme to minimize energy costs and user thermal discomfort; proposing a demand response dispatch method based on Deep Q-Learning Network (DQN) for indoor air temperature control and thermal comfort management; developing a DQN-based method to optimize electric vehicle charging scheduling in smart homes to minimize charging costs; and proposing an online building energy optimization method using DQN to schedule time-series and time-shifted loads.

[0007] As can be seen, most of the aforementioned reinforcement learning-based works utilize the Q-learning framework to solve the demand response problem. However, the structure based on the Q-learning framework is prone to overestimation and the curse of dimensionality, resulting in a lack of further modeling extensions and cross-sectional comparisons of the convergence and performance of reinforcement learning, thus failing to efficiently solve the demand response problem. Summary of the Invention

[0008] Therefore, it is necessary to provide a model-free reinforcement learning method for energy demand response management to address the aforementioned technical problems.

[0009] A model-free reinforcement learning method for energy demand response management includes the following steps: constructing a residential appliance model, wherein the residential appliances include dispatchable appliances, non-dispatchable appliances, and plug-in electric vehicles; determining social welfare by combining the comprehensive residential electricity cost and retailer profit, and balancing the comprehensive residential electricity cost and retailer profit according to the social welfare, wherein the social welfare is represented as a price-based non-convex optimization problem for residential demand response management; constructing reinforcement learning solutions for the price-based non-convex optimization problem for residential demand response management using Q-learning algorithm based on Q-tables, Q-network algorithm combining deep learning and Q-learning, and Actor-Critic algorithm, respectively, based on the transmission data of the power grid; and determining the optimal solution and the optimal retail price sequence based on the three reinforcement learning solutions for the price-based non-convex optimization problem for residential demand response management.

[0010] In one embodiment, constructing the residential appliance model specifically includes: assuming the set of dispatchable appliances is N. d ={1,...,D}, and the set of unschedulable appliances is N. n = {1,...,N}, and the set of plug-in electric vehicles is N. p If the set of all electrical appliances is {1,...,P}, then the set of all electrical appliances can be represented as N. z =N d ∪N n ∪N p ={1,…Z}.

[0011] In one embodiment, modeling the schedulable appliances specifically includes: classifying the schedulable appliances d∈N d The actual electricity consumption is expressed as follows:

[0012]

[0013] In the formula, T represents the total time slots, and R... d,t It is the expected energy demand of dispatchable appliances, expressed in kWh, E d,t The actual electricity consumption of dispatchable appliances, expressed in kWh, ρ d,t The retail price of electricity determined by retailers, θ t The wholesale price at which retailers purchase energy from the energy market, and ρ d,t ≥θ t δ t The price elasticity coefficient, which is less than 0, represents the relationship between energy demand and electricity retail prices. The difference between expected energy demand and actual electricity consumption, i.e., the demand error, represents the residents' electricity consumption happiness index, as shown in the formula:

[0014]

[0015] In the formula, C d,t Let E represent the electricity happiness function, which indicates the actual consumption of electricity by residents during time period t at a specific retail electricity price. d,t The closer to R d,t The stronger the sense of happiness; d1 and h d2 All are happiness coefficients related to electrical appliances; the usable range of the electricity happiness function is limited as follows:

[0016]

[0017] In the formula, This is the lower bound of the demand error. This is the upper bound of the demand error.

[0018] In one embodiment, modeling the unschedulable electrical appliance specifically includes ensuring that the unschedulable electrical appliance satisfies the constant relationship between energy demand and actual electricity consumption in all time periods, i.e.:

[0019] R n,t =E n,t (4)

[0020] In the formula, R n,t E represents the expected energy demand of non-dispatchable appliances. n,t This indicates the actual power consumption of unschedulable electrical appliances.

[0021] In one embodiment, modeling the plug-in electric vehicle specifically includes: classifying the plug-in electric vehicle p∈N p The actual electricity consumption is expressed as follows:

[0022]

[0023] In the formula, R p,t For the anticipated energy demand of plug-in electric vehicles, E p,t This refers to the actual electricity consumption of plug-in electric vehicles, and in E p,t When <0, it indicates discharge, at E p,t When ρ > 0, it indicates charging; p,t For the retail price of electricity for electric vehicles, θ t For wholesale prices, δ t The price elasticity coefficient, which is less than 0, represents the relationship between energy demand and electricity retail prices; the electric happiness function C of the plug-in electric vehicle. p,t Represented as:

[0024]

[0025] In the formula, h d1 and h d2All of these are happiness coefficients related to plug-in electric vehicles; the plug-in electric vehicles have corresponding rated power limits for each time slot, namely:

[0026]

[0027] In the formula, This indicates the rated power of the plug-in electric vehicle; based on the charging and discharging characteristics and battery capacity of the plug-in electric vehicle, the rated power of the plug-in electric vehicle is subject to the following limitations:

[0028]

[0029] In the formula, and These represent the minimum and maximum battery capacity, respectively. Represents the initial energy level, e p Indicates charging or discharging efficiency; the cost of battery degradation in the plug-in electric vehicle is:

[0030] Deg p,t =υ|E p,t |(9)

[0031] In the formula, υ is the degradation coefficient; and the retail price of electricity is subject to the following constraints:

[0032] ρ min ≤ρ n,t ,ρ p,t ,ρ d,t ≤ρ max (10)

[0033] In the formula, ρ min ρ represents the minimum retail price. max This represents the maximum retail price.

[0034] In one embodiment, determining social welfare by combining the comprehensive residential electricity cost and retailer profits, and balancing the comprehensive residential electricity cost and retailer profits based on the social welfare, specifically includes: obtaining the energy consumption of residential appliances, then the comprehensive residential electricity cost is:

[0035]

[0036] In the formula, EC t This represents the total cost of all electrical appliances during time period t. For the electricity cost of plug-in electric vehicles, BD p,t The price of electricity sold during battery discharge is given; given the total energy consumption of residential electricity, the retail price of electricity, and the wholesale price of electricity, the retailer's profit is:

[0037]

[0038] In the formula, EP t Let represent the retailer's profit during time period t; determine the actual electricity consumption based on the retail price, and maximize the social welfare by balancing residential electricity costs and retailer profits. Maximizing social welfare refers to a weight balance problem with a two-level optimization objective, namely, a price-based residential demand response management non-convex optimization problem, expressed as:

[0039]

[0040] In the formula, n,d,p∈N z ,t∈T,ω is the relative social value weight balancing commercial profits and residential energy consumption, andP is the vector of retail electricity prices consisting of residential appliances.

[0041] In one embodiment, prior to constructing a reinforcement learning solution for the price-based residential demand response management nonconvex optimization problem, the method further includes: constructing a retailer-resident electricity trading model based on a model-free reinforcement learning algorithm, wherein the basic elements of the model-free reinforcement learning algorithm include a quintuple. <S,A,R,T t The corresponding retailer-resident electricity trading model is: State S = {s1,...,s...} T}, R i,t Energy demand and E i,t-1 The actual power consumption is generated by the residence; Action A = {a1,...,a M Retail price ρ i,t Determined by the electricity retailer, M represents the discrete retail price range [ρ min ,ρ max The set number after [r1, ..., r]; the reward R = {r1, ..., r T Social welfare F t (P); State transition function T t , which is related to retail price; discount factor γ∈[0,1], which is the importance weight of future social welfare.

[0042] In one embodiment, the reinforcement learning solution for the price-based housing demand response management nonconvex optimization problem constructed using a Q-table-based Q-learning algorithm specifically includes: updating the q-value based on the Bellman equation and a greedy policy, wherein the q-value update formula is:

[0043]

[0044] In the formula, k represents the training index, and lr represents the learning rate; when the Q-table converges, a greedy strategy is used to obtain the optimal retail price, which is:

[0045]

[0046] The Q-value function is used to approximate the Q-table, that is:

[0047]

[0048] In the formula, α represents the weight of the Q-network; in supervised learning, the q-value of the next time slot is estimated by the network with the current training index k as the label, and the target q-value can be expressed as:

[0049]

[0050] The loss function of the Q-network can be expressed as:

[0051]

[0052] In the formula, the gradient descent method is used to iteratively update the weight α.

[0053] In one embodiment, the reinforcement learning solution for constructing a price-based residential demand response management non-convex optimization problem using a Q-network algorithm combining deep learning and Q-learning specifically includes: employing DDQN with two identical Q-network structures for each appliance: the current Q-network α... k and target Q-network And respectively used for decision-making and q-function-based estimation; based on formula (17), the target q-value is expressed as:

[0054]

[0055] In the formula, This represents the q-value of the target Q-network. Let q represent the estimated q-value of the current Q-network, then the loss function is expressed as:

[0056]

[0057] Using formulas (19) and (20), electricity retailers determine the current Q-network α. k The system interacts with the residence for F days, stores the observations in an experience replay buffer D, and extracts M sets of observations from D to train the current Q-network. The target Q-network weights are... It is then updated at a fixed period C.

[0058] In one embodiment, the reinforcement learning solution for constructing a price-based residential demand response management non-convex optimization problem using the Actor-Critic algorithm specifically includes: distributing policies by adding an Actor-Critic policy network to the Q-network.

[0059]

[0060] In the formula, Pr represents the probability distribution; the output of the Actor-Critic network is expressed as:

[0061] η i,t =φ a (s i,t ,β i )(twenty two)

[0062] J i,t =φ c (s i,t ,a i,t ,α i )(twenty three)

[0063] In the formula, β i and α i These are the weights of the Actor neural network and the Critic neural network, respectively, φ. a and φ c For the activation function, the estimated action η i,t It is the output of the Actor network, the action-value function J. i,t If the output of the Critic network is in time slot t, then the timing difference error is:

[0064] TD i,t =F t (P)+γJ i,t+1 -J i,t (twenty four)

[0065] In the formula, when the discount factor γ = 0, it means that future state values ​​are ignored; when γ = 1, it means that the learning algorithm gives fair consideration to the rewards of all time periods; then the loss function of the Critic neural network is defined as:

[0066]

[0067] Using the temporal difference error as the evaluation function of the Actor network, the update formula for the Actor-Critic network based on backpropagation is expressed as:

[0068]

[0069]

[0070] In the formula, la and lc represent the learning rates of Actor and Critic, respectively.

[0071] Compared to existing technologies, the advantages and beneficial effects of this invention are as follows: By classifying residential appliances into dispatchable appliances, non-dispatchable appliances, and plug-in electric vehicles, and constructing a residential appliance model, the comprehensive electricity cost for residents and the profit for retailers are calculated. Social welfare is determined by combining the comprehensive electricity cost for residents and the profit for retailers. The comprehensive electricity cost for residents and the profit for retailers are then balanced based on the social welfare. Social welfare is used to represent the price-based non-convex optimization problem of residential demand response management. Based on the transmission data from the power grid, a strengthened algorithm for the price-based non-convex optimization problem of residential demand response management is constructed using a Q-learning algorithm based on Q-tables, a Q-network algorithm combining deep learning and Q-learning, and an Actor-Critic algorithm, respectively. The proposed solution, based on reinforcement learning solutions for three price-based residential demand response management (RTD) non-convex optimization problems, determines the optimal solution and obtains the optimal retail price sequence. It fully considers electric vehicles with charging and discharging characteristics. Furthermore, it models the RTD non-convex optimization problem using three different algorithms, compares their performance, and selects the optimal algorithm model for the current environment to obtain the corresponding optimal retail price sequence. In addition, it achieves maximum social welfare by integrating the interests of residents and electricity retailers, thus providing an efficient solution for price-based residential demand response management and enabling the planning of optimal retail prices under unknown electricity market conditions. Attached Figure Description

[0072] Figure 1 This is a flowchart illustrating a model-free reinforcement learning method for energy demand response management in one embodiment.

[0073] Figure 2 This is a block diagram of DQN in price-based residential demand response management in one embodiment;

[0074] Figure 3 This is a block diagram of DDQN with an experience pool in price-based residential demand response management in one embodiment;

[0075] Figure 4 An Actor-Critic block diagram in price-based residential demand response management in one embodiment;

[0076] Figure 5 This is a schematic diagram comparing the energy demand and consumption of dispatchable appliances within a day in one embodiment.

[0077] Figure 6 This is a schematic diagram comparing the demand and consumption of plug-in electric vehicles in one day in one embodiment;

[0078] Figure 7 This is a schematic diagram illustrating the convergence effect of the q-table for four types of plug-in electric vehicles over 24 time periods in one embodiment.

[0079] Figure 8 for Figure 6 A schematic diagram of the convergence evolution of DQN weights and time-series difference error for p3 in a plug-in electric vehicle.

[0080] Figure 9 This is a schematic diagram illustrating the actual energy consumption and retail price planning of four plug-in electric vehicles over 24 time periods under the DDQN and DQN algorithms in one embodiment.

[0081] Figure 10 This is a schematic diagram illustrating the changes in action probability of the Actor network for a plug-in electric vehicle before and after training in one embodiment.

[0082] Figure 11 for Figure 5 Dispatchable electrical appliances D3 and Figure 6 A schematic diagram illustrating the evolution of action values ​​in the Actor network of a plug-in electric vehicle p2;

[0083] Figure 12 This is a schematic diagram illustrating the average social welfare level of four algorithms after 10 trials in one embodiment. Detailed Implementation

[0084] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to specific embodiments and accompanying drawings. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention.

[0085] In one embodiment, such as Figure 1 As shown, a model-free reinforcement learning method for energy demand response management is provided, including the following steps:

[0086] Step S101: Construct a residential appliance model, which includes dispatchable appliances, non-dispatchable appliances, and plug-in electric vehicles.

[0087] Specifically, based on the transferability and economy of energy consumed by the load, residential appliances are generally divided into dispatchable appliances (such as washing machines and dishwashers), non-dispatchable appliances (such as lights and refrigerators), and plug-in electric vehicles. Because plug-in electric vehicles have charging and discharging characteristics, they require special consideration.

[0088] Typically, electricity retailers receive residents' actual energy consumption in the previous period and their expected energy demand in the current period, and dynamically adjust their retail pricing strategies while taking into account commercial interests. Residents, in turn, proactively change their energy consumption based on changes in retail prices in the current period.

[0089] Therefore, before constructing the residential appliance model, we can assume that the set of dispatchable appliances is N. d ={1,...,D}, and the set of unschedulable appliances is N. n = {1,...,N}, and the set of plug-in electric vehicles is N. p If the set of all electrical appliances is {1,...,P}, then the set of all electrical appliances can be represented as N. z =N d ∪N n ∪N p ={1,…Z}.

[0090] Specifically, the model for schedulable appliances includes: defining schedulable appliances d∈N. d The actual energy consumption is expressed as:

[0091]

[0092] In the formula, T represents the total time slots, and R... d,t It is the expected energy demand of dispatchable appliances, expressed in kWh, E d,t The actual electricity consumption of dispatchable appliances, expressed in kWh, ρ d,t The retail price of electricity determined by retailers, θ t The wholesale price at which retailers purchase energy from the energy market, and ρ d,t ≥θ t δ t The price elasticity coefficient, which is less than 0, represents the relationship between energy demand and electricity retail prices. The difference between expected energy demand and actual electricity consumption, i.e., the demand error, represents the residents' electricity consumption happiness index, as shown in the formula:

[0093]

[0094] In the formula, C d,t Let E represent the electricity happiness function, which indicates the actual consumption of electricity by residents during time period t at a specific retail electricity price. d,t The closer to R d,t The stronger the sense of happiness; d1 and h d2 All are happiness coefficients related to electrical appliances; the usable range of the electricity happiness function is limited as follows:

[0095]

[0096] In the formula, This is the lower bound of the requirement error. This is the upper bound of the requirement error.

[0097] Specifically, in a daily-based electricity trading model, users send signals (Rt) to electricity retailers during time period t.d,t E d,t Retailers obtain corresponding profit estimates from feedback signals and make decisions regarding ρ. d,t+1 The adjustment decision. The essence of formula (1) is that R d,t To meet the maximum consumption of the residence, when θ t Approximately ρ d,t At this time, users are willing to consume more electricity but cannot accept higher retail prices, thus reducing consumption. Therefore, the demand error (R) d,t -E d,t This can represent the happiness index of residents' electricity consumption. d,t This reveals the actual consumption E of residents over time period t at a reasonable retail price. d,t The closer to R d,t The higher the energy consumption, the greater their sense of well-being. However, when residents are limited by high retail prices in their motivation to consume energy, their sense of well-being decreases, thus restricting the usable range of the electricity-based happiness function. and between.

[0098] Specifically, the model for unschedulable appliances includes ensuring that the unschedulable appliances maintain the identity relationship between expected energy demand and actual electricity consumption in all time periods, namely:

[0099] R n,t =E n,t (4)

[0100] In the formula, R n,t E represents the expected energy demand of non-dispatchable appliances. n,t This indicates the actual power consumption of unschedulable electrical appliances.

[0101] Specifically, taking a refrigerator, an unschedulable appliance, as an example, its expected energy demand will not change without external influence. Therefore, the expected energy demand is equal to the actual electricity consumption in all time periods.

[0102] Specifically, when constructing the model for plug-in electric vehicles, the following steps are involved: representing plug-in electric vehicles p∈N p The actual electricity consumption is expressed as follows:

[0103]

[0104] In the formula, R p,t For the anticipated energy demand of plug-in electric vehicles, E p,t This refers to the actual electricity consumption of plug-in electric vehicles, and in E p,t When <0, it indicates discharge, at E p,t When ρ > 0, it indicates charging; p,t For the retail price of electricity for electric vehicles, θt For wholesale prices, δ t The price elasticity coefficient, which is less than 0, represents the relationship between energy demand and electricity retail prices; the electric happiness function C of the plug-in electric vehicle. p,t Represented as:

[0105]

[0106] In the formula, h d1 and h d2 All of these are happiness coefficients related to plug-in electric vehicles; the plug-in electric vehicles have corresponding rated power limits for each time slot, namely:

[0107]

[0108] In the formula, This indicates the rated power of a plug-in electric vehicle; based on the charging and discharging characteristics and battery capacity of plug-in electric vehicles, the rated power of plug-in electric vehicles is limited as follows:

[0109]

[0110] In the formula, and These represent the minimum and maximum battery capacity, respectively. Represents the initial energy level, e p Indicates charging or discharging efficiency; the cost of battery degradation in the plug-in electric vehicle is:

[0111] Deg p,t =υ|E p,t | (9)

[0112] In the formula, υ is the degradation coefficient; and the retail price of electricity is subject to the following constraints:

[0113] ρ min ≤ρ n,t ,ρ p,t ,ρ d,t ≤ρ max (10)

[0114] In the formula, ρ min ρ represents the minimum retail price. max This represents the maximum retail price.

[0115] Specifically, plug-in electric vehicles (PHEVs) possess all the characteristics of dispatchable electrical appliances; therefore, their actual electricity consumption is determined by expected energy demand, retail electricity prices, wholesale electricity prices, and price elasticity coefficients. Furthermore, electric vehicles with onboard batteries require corresponding rated power limits in each time slot to ensure safety; therefore, their actual electricity consumption must be within a certain range. This range.

[0116] Furthermore, considering the charging and discharging characteristics and battery capacity of plug-in electric vehicles, the system is also limited by the minimum and maximum battery capacity. In addition, the impact of frequent charging and discharging on battery life is taken into account, thus quantifying the cost of battery degradation. Retail electricity prices are subject to restrictions set by the International Organization for Standardization (ISO) and must fall between minimum and maximum values.

[0117] Step S102: Determine social welfare by combining residential comprehensive electricity costs and retailer profits. Balance residential comprehensive electricity costs and retailer profits based on social welfare. Social welfare is expressed as a non-convex optimization problem of price-based residential demand response management.

[0118] Specifically, price-based residential demand response management is treated as a two-tiered optimization problem. From the user's perspective, the goal is to achieve optimal actual energy consumption by minimizing the overall residential cost; from the electricity retailer's perspective, the goal is to maximize profits. By considering both expectations simultaneously, the relative social value of commercial profits and overall user costs is revealed through maximizing social welfare.

[0119] Among these, social welfare is determined by combining the comprehensive cost of residential electricity consumption and retailer profits. Specifically, this includes obtaining the energy consumption of residential appliances. The comprehensive cost of residential electricity consumption is then:

[0120]

[0121] In the formula, EC t This represents the total cost of all electrical appliances during time period t. For the electricity cost of plug-in electric vehicles, BD p,t This represents the electricity sales price during battery discharge; given the total energy consumption of residential electricity, the retail price of electricity, and the wholesale price of electricity, the retailer's profit is:

[0122]

[0123] In the formula, EP t Let represent the retailer's profit during time period t; determine actual electricity consumption based on retail prices; maximizing social welfare refers to a weight balance problem with a two-level optimization objective, namely, a price-based residential demand response management non-convex optimization problem, expressed as:

[0124]

[0125] In the formula, n,d,p∈N z ,t∈T,ω is the relative social value weight balancing commercial profits and residential energy consumption, andP is the vector of retail electricity prices consisting of residential appliances.

[0126] Specifically, according to formulas (1) and (5), the actual electricity consumption is determined by the retail price. Maximizing social welfare refers to a weight balance problem with a two-level optimization objective, which can represent the non-convex optimization problem of housing demand management based on price. This makes it easier to calculate the optimal retail price sequence based on the non-convex optimization problem.

[0127] Step S103: Based on the data transmitted from the power grid, reinforcement learning solutions for the non-convex optimization problem of housing demand response management based on price are constructed using Q-learning algorithm based on Q-tables, Q-network algorithm combining deep learning and Q-learning, and Actor-Critic algorithm, respectively.

[0128] Specifically, based on model-free reinforcement learning algorithms, power grid transmission data is used to make decisions based on policy exploration. Reinforcement learning solutions for price-based residential demand response management non-convex optimization problems are constructed by employing Q-learning algorithms based on Q-tables, Q-network algorithms combining deep learning and Q-learning, and the Actor-Critic algorithm, respectively. The advantages and disadvantages of the three algorithms in the modeling process are analyzed, and simulation experiments are used to verify the effectiveness and environmental applicability of the three algorithms.

[0129] Before constructing a reinforcement learning solution for the nonconvex optimization problem, the process also includes: building a retailer-resident electricity trading model based on a model-free reinforcement learning algorithm, wherein the basic element of the model-free reinforcement learning includes a quintuple. <S,A,R,T t The corresponding retailer-resident electricity trading model is: State S = {s1,...,s...} T}, R i,t Energy demand and E i,t-1 The actual power consumption is generated by the residence; Action A = {a1,...,a M Retail price ρ i,t Determined by the electricity retailer, M represents the discrete retail price range [ρ min ,ρ max The set number after [r1, ..., r]; the reward R = {r1, ..., r T Social welfare F t (P); State transition function T t , which is related to retail price; discount factor γ∈[0,1], which is the importance weight of future social welfare.

[0130] Among them, the Q-learning algorithm based on Q-tables constructs a reinforcement learning solution for the price-based residential demand response management non-convex optimization problem, specifically including: updating the q-value based on the Bellman equation and a greedy policy, with the q-value update formula being:

[0131]

[0132] In the formula, k represents the training index, and lr represents the learning rate; when the Q-table converges, a greedy strategy is used to obtain the optimal retail price, which is:

[0133]

[0134] The Q-value function is used to approximate the Q-table, that is:

[0135]

[0136] In the formula, 'a' represents the weight of the Q-network; in supervised learning, the q-value of the next time slot is estimated using the network at the current training index k as the label, and the target q-value can be expressed as:

[0137]

[0138] The loss function of the Q-network can be expressed as:

[0139]

[0140] In the formula, the gradient descent method is used to iteratively update the weight α.

[0141] Specifically, Q-learning algorithms require pre-creating a corresponding Q-table for each appliance. However, the increasing number of appliances, discrete retail prices, and time periods place a significant burden on storage and computation. Therefore, Q-learning algorithms based on Q-tables must consider an appropriate discrete action space and residential area size when solving residential demand response management problems. To address the curse of dimensionality caused by excessively large Q-tables, Q-functions are used to approximate them. These are typically constructed using deep neural networks. Compared to the q-values ​​in the Q-table and the optimal retail price sequence under a greedy strategy, only the convergence effect of the weight α needs to be considered. The q-function, also known as the action-value function, is an evaluation function that measures a pair of (s) based on the prediction of future rewards. t ,a t The good or bad of ).

[0142] Among them, a reinforcement learning solution for the non-convex optimization problem of housing demand response management based on price is constructed using a Q-network algorithm that combines deep learning and Q-learning. Specifically, this includes using DDQN with two identical Q-network structures for each appliance: the current Q-network α... k and target Q-network And respectively used for decision-making and q-function-based estimation; based on formula (17), the target q-value is expressed as:

[0143]

[0144] In the formula, This represents the q-value of the target Q-network. Let q represent the estimated q-value of the current Q-network, then the loss function is expressed as:

[0145]

[0146] Using formulas (19) and (20), electricity retailers determine the current Q-network α. k The system interacts with the residence for F days, stores the observations in an experience replay buffer D, and extracts M sets of observations from D to train the current Q-network. The target Q-network weights are... It is then updated at a fixed period C.

[0147] Specifically, the linear approximation of the q-function by a three-layer deep neural network is as follows: Figure 2 As shown. For offline DQN, historical information from the exploration is directly utilized by the Q-network. However, the correlation of the data easily leads to local solutions. Furthermore, while maximizing the operation can quickly bring the q-value close to the possible optimization objective, it easily overestimates the q-value, leading to overestimation. Therefore, the DDQN algorithm can be used to eliminate the overestimation problem by decoupling the action selection and the calculation of the target q-value. The expected reward of DDQN has been proven to be an unbiased estimate.

[0148] Specifically, the DDQN algorithm uses two identical Q-network structures for each appliance, with the current Q-network α... k and target Q-network And respectively used for decision-making and q-function-based estimation. Using formulas (19) and (20), it can be seen that the electricity retailer, based on the current Q-network α k The system interacts with the residence for F days, storing the observations in an experience replay buffer D. Then, M sets of observations are extracted from D to train the current Q-network, while the target Q-network weights are... Updates are then performed at a fixed period C. This delayed update reduces the parameter dependency between the target Q-network and the current Q-network.

[0149] Therefore, DDQN eliminates the overestimation problem by decoupling the two steps of selecting the current action and calculating the target Q-value. The detailed structure of the algorithm, such as the empirical replay pool and the network model, is as follows: Figure 3 As shown. It's important to note that when the capacity D of the experience replay pool is set to 1, the online DDQN algorithm is used, meaning that data mined from the current Q-network can be directly and in real-time utilized by the target network. Furthermore, the iterative formula for Q-learning is derived from the Bellman equation and a greedy policy:

[0150]

[0151] For ease of understanding, let's assume the q-value is optimal. Then we have the ideal optimal q-function as:

[0152]

[0153] In the formula, E(·) represents the expectation. However, the optimal q-function should satisfy the Bellman equation:

[0154]

[0155] Therefore, the overestimation problem can be simply attributed to the inequality:

[0156]

[0157] The results show that fitting the q-function using gradient descent yields a larger expected value. In summary, due to the adoption of greedy strategies (such as the ò-greedy strategy) to maximize cumulative reward in unknown environments to maintain efficient exploration, overestimation is unavoidable for reinforcement learning algorithms based on the Q-learning framework.

[0158] The solution employs the Actor-Critic algorithm to construct a reinforcement learning solution for the price-based residential demand response management non-convex optimization problem. Specifically, this involves policy distribution through the addition of an Actor-Critic policy network to the Q-network.

[0159]

[0160] In the formula, Pr represents the probability distribution; the output of the Actor-Critic network is expressed as:

[0161] η i,t =φ a (s i,t ,β i )(twenty two)

[0162] J i,t =φ c (s i,t ,a i,t ,α i )(twenty three)

[0163] In the formula, β i and α i These are the weights of the Actor neural network and the Critic neural network, respectively, φ. a and φ c For the activation function, the estimated action η i,tIt is the output of the Actor network, the action value function J. i,t If the output of the Critic network is in time slot t, then the timing difference error is:

[0164] TD i,t =F t (P)+γJ i,t+1 -J i,t (twenty four)

[0165] In the formula, when the discount factor γ = 0, it means that future state values ​​are ignored; when γ = 1, it means that the learning algorithm gives fair consideration to the rewards of all time periods; then the loss function of the Critic neural network is defined as:

[0166]

[0167] Using the temporal difference error as the evaluation function of the Actor network, the update formula for the Actor-Critic network based on backpropagation is expressed as:

[0168]

[0169]

[0170] In the formula, la and lc represent the learning rates of Actor and Critic, respectively.

[0171] Specifically, in reinforcement learning, Q-learning-based methods are called a class of value-based learning algorithms, which maximize reward expectation by selecting the best action based on a deterministic policy. In the optimization problem of residential demand response management, stochastic policies must be relied upon to handle the continuous action space to obtain a more accurate energy dispatching scheme. Therefore, a price-based residential demand response management learning algorithm based on policy values, namely the Actor-Critic method, can be adopted. This method approximates the policy distribution by adding a policy network to the Q-network, and its structure is as follows: Figure 4 As shown.

[0172] Specifically, the policy network is represented as an Actor network, and the Critic network corresponds to a Q-network. Generally, the input to an Actor network is a state vector, and the output is an estimated action. The structure of the Critic network is consistent with that of the Q-network. It is important to note that the Actor-Critic algorithm can be divided into three steps: 1) Calculate the (s) of the current time slot and the next time slot respectively. t ,a t 2) From the calculated action value function J i,t and J i,t+1 Determine the evaluation function TD i,t 3) Utilizing TDi,t Update the Actor-Critic parameter α i and β i .

[0173] Step S104: Based on the reinforcement learning solutions for three price-based residential demand response management non-convex optimization problems, determine the optimal solution and the optimal retail price sequence.

[0174] Specifically, reinforcement learning solutions for the non-convex optimization problem of price-based housing demand response management are constructed based on the three algorithms mentioned above. The optimal solution is then determined by considering the current environment. For example, if a fast convergence algorithm is required, a solution based on a Q-table and Q-learning algorithm can be used; if reducing the impact of overestimation on overall decision-making is needed, a solution combining deep learning and Q-learning (Q-network algorithm) can be used; and if maximizing social welfare or obtaining sufficient historical information through thorough environmental exploration is required, an Actor-Critic algorithm can be employed. After determining the appropriate solution based on the current needs, the optimal retail price sequence is obtained using this solution, thereby efficiently solving the price-based housing demand response management problem.

[0175] In this embodiment, residential appliances are categorized into dispatchable appliances, non-dispatchable appliances, and plug-in electric vehicles, and a residential appliance model is constructed. Social welfare is determined by combining the comprehensive electricity cost for residents and the profit of retailers. This social welfare is then used to balance the electricity cost for residents and the profit of retailers, representing the non-convex optimization problem of price-based residential demand response management. Based on the power grid transmission data, reinforcement learning solutions for the price-based residential demand response management non-convex optimization problem are constructed using three different algorithms: a Q-learning algorithm based on Q-tables, a Q-network algorithm combining deep learning and Q-learning, and an Actor-Critic algorithm. Based on these three reinforcement learning solutions, the optimal solution is determined, and the optimal retail price sequence is obtained. This approach fully considers the charging and discharging characteristics of electric vehicles. Furthermore, the three algorithms are used to model the price-based residential demand response management non-convex optimization problem, and the merits of the algorithms are compared. The optimal algorithm model in the current environment is adopted to obtain the corresponding optimal retail price sequence. In addition, by integrating the interests of residents and retailers, the maximum social welfare is achieved, thus providing an efficient solution for price-based residential demand response management.

[0176] In one embodiment, the effectiveness of the three model-free reinforcement learning algorithms described above can be verified through experiments. The algorithms were implemented using MATLAB R2014a on a desktop computer with an i5-12400F CPU@2.50GHz, 16GB of memory, and a 64-bit Windows 11 operating system.

[0177] The experiment considers the energy demand response management problem of 6 dispatchable appliances {d1,d2,d3,d4,d5,d6}, ​​4 electric vehicles {p1,p2,p3,p4}, and 5 non-dispatchable appliances {n1,n2,n3,n4,n5} throughout the day (24 time periods). The energy demand distribution of plug-in electric vehicles, non-dispatchable appliances, and dispatchable appliances is based on data from a gas-fired power company, while wholesale prices are determined by the energy market and data comes from the electricity company. To ensure the normal operation of the electricity market economy and protect the reasonable demand for normal electricity consumption by residents, we must coordinate the retail pricing strategies of electricity retailers and the electricity consumption strategies of users to strive for the maximization of social welfare. Tables 1 and 2 list the time-varying parameters and the appliance-related parameters, respectively. Note that the discrete retail price interval is 0.1, and according to the price parameters in Table 2, the retail price range is [2.4, 6.7]. Therefore, the number of discrete actions is 44 (i.e., M = 44).

[0178] Table 1 Time-varying parameters

[0179]

[0180] Table 2 Electrical Parameters

[0181]

[0182] First, based on the above data analysis, we examine the effectiveness of the Q-learning algorithm based on Q-tables. Figure 5 This shows a comparison of the energy demand and consumption of dispatchable appliances within a day. Figure 6 This chart compares the energy demand and consumption of plug-in electric vehicles (PHEVs) throughout the day. It shows that the energy demand of non-dispatchable appliances is significantly higher than that of dispatchable appliances and PHEVs, with peak demand occurring between 12:00-16:00 and 18:00-24:00. Furthermore, due to the sporadic usage pattern of PHEVs increasing energy demand elasticity, there is no electricity demand during certain periods of the day.

[0183] Table 3 shows the specific daily retail price planning obtained using the Q-table-based method. It can be seen that the retail prices strictly meet price constraints. The overall trend of the retail price planning fluctuates over the 24 time periods, influenced by social welfare and wholesale prices. When retail prices are too high, it is detrimental to social welfare relative to users. For example, during the 10:00-13:00 time period, due to the increase in retail prices (from $3 to $6.6) followed by a rapid decrease (from $6.6 to $4.2), d1 continuously reduced its energy consumption (from 8.9 kWh to 5.3 kWh).

[0184] Table 3: Optimal Retail Price Planning for All Appliances Across 24 Time Slots

[0185]

[0186] Factors that benefit customers (such as electricity well-being and the cost of plug-in electric vehicles) will subsequently lead residents to consume more electricity in the next time slot. This also impacts relative social welfare due to retailers' commercial interests. For example, if retail prices are too low during the 9:00 time slot, retailers might significantly increase prices {d2,p4,n1} during that time slot, while plug-in electric vehicles actively discharge during 2:00-6:00 and 20:00-24:00 to cope with the excessively high electricity prices.

[0187] As shown in Table 3, peak retail prices do not occur during peak demand periods (12:00-16:00 and 18:00-24:00), and the average retail price during peak demand periods ($77.24) is lower than the average price during other periods ($79.62). This is because the goal of maximizing social welfare ensures that pricing is beneficial to both retailers and customers, creating a virtuous cycle between retail prices and actual electricity consumption to maintain a relative balance in social welfare.

[0188] Figure 7 The convergence effect of the q-tables for four types of plug-in electric vehicles over 24 time periods is described. To maximize social welfare, retailers continuously change their electricity pricing strategies, gradually bringing the q-value to its maximum while taking into account stochastic electricity consumption patterns and the storage capacity limitations of plug-in electric vehicles.

[0189] Second, based on the above data analysis, the effectiveness of the Q-network algorithm combining deep learning and Q-learning is discussed. Deep learning and Q-learning are combined to construct a Q-network instead of a Q-table for evaluation, and two algorithms are proposed: online DQN and offline DDQN. The capacity of the experience buffer D is set to F=20, and M=15 sets of observations are sampled and removed from D in each iteration.

[0190] Figure 8 Represents the weights of the hidden / output layers. The convergence results across the 24 time slots of p3 demonstrate the effectiveness of the DQN-based algorithm. The DQN-based learning algorithm coordinates the energy consumption of each appliance and, based on the retail price with immediate feedback, makes decisions about the expected actual energy consumption using a greedy strategy, ultimately converging to the optimum.

[0191] Due to the complexity of the residential demand response management problem (e.g., the non-convexity of the optimization problem, algorithm-specific parameter settings), different learning algorithms often yield different scheduling strategies. Figure 9 The diagram shows the actual energy consumption and retail price planning for four plug-in electric vehicles over 24 time periods under the DDQN and DQN algorithms. It can be seen that although the retail prices and actual consumption obtained by the two algorithms differ, the relative actual consumption determines the relative retail price. Specifically, when the two algorithms discharge (Ei) in a certain time period... p,t <0) and charging (E) p,t When the value is greater than 0, the retail price of discharging is often lower than the retail price of charging.

[0192] Third, the effectiveness of the Actor-Critic algorithm is analyzed based on the data above. For the discrete action space, the Actor policy network uses the softmax function to select the best action for each tool based on the principle of maximum probability, as shown below.

[0193]

[0194] Figure 10 This represents the action probability output of an Actor network with four plug-in electric vehicles before and after training. Each component of the feature vector... The probability distribution of the 44 actions is represented by the softmax function. All action probabilities are randomly initialized before training, and the evaluation function TD... i,t During training, the objective function is used to select the optimal action, whose probability in the action set increases continuously while the probabilities of other actions decrease, eventually converging to 1.

[0195] Figure 11 The evolution of the action values ​​of the schedulable appliance d3 and the plug-in electric vehicle p2 is shown, which further demonstrates the effectiveness of the Actor-Critic algorithm, as can be seen from their eventual convergence.

[0196] Finally, the three algorithms mentioned above are compared. Figure 12The average social welfare level of the four proposed algorithms after 10 trials is shown. The following characteristics can be observed: Since Q-learning methods are value-based (q-value) algorithms, they tend to fall into overestimation by maximizing the q-value, resulting in curves with large amplitude, high frequency, and fast convergence. A drawback is that they may not converge to the optimal state for non-convex optimization objectives, thus requiring the exploration and learning of more potential solutions through a greedy strategy. In contrast, Actor-Critic utilizes both a value-based Critic network and a policy-based Actor network, allowing for thorough exploration of the environment and acquisition of broader historical information. A drawback is slow convergence speed; in fact, due to the online learning method, the Critic network may even fail to converge.

[0197] Furthermore, offline DDQN with experience replay functionality can effectively reduce the impact of overestimation on agent decision-making because its estimation of q-values ​​is unbiased. Experiments show that the oscillations in the curve are significantly smaller than those of online DQN and Q-table-based methods.

[0198] Table 4 Comparison of Average Retail Prices and Social Welfare

[0199]

[0200] Table 4 shows the total social welfare. A comparison of algorithms across all time periods shows that the Actor-Critic algorithm yields the highest social welfare, 13.04% higher than that of DQN, but its average daily retail price is indeed 34.47% lower than DQN. This indicates that the high social welfare combines corporate profits and user benefits, thus depicting the social welfare characteristics of electricity consumption. Furthermore, as the experiments demonstrate, the drawback of the discrete action space Q-table method is only the calculation and storage of q-values, while the experimental results for social welfare do not show any significant disadvantages compared to the neural network-based Q-learning algorithm.

[0201] In the above simulation experiment, a series of model-free reinforcement algorithms were used to coordinate the actual power consumption of electrical appliances on the environmental side with the electricity retail price on the intelligent agent side, providing a system solution with long-term decision-making capabilities for real-time demand response in smart grids, and the effectiveness of the above method was verified by simulation.

[0202] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The program can be stored in a computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. The storage medium can be a magnetic disk, optical disk, read-only memory (ROM), or random access memory (RAM), etc.

[0203] It is obvious to those skilled in the art that the modules or steps of the present invention described above can be implemented using general-purpose computing devices. They can be centralized on a single computing device or distributed across a network of multiple computing devices. Optionally, they can be implemented using computer-executable program code, thereby storing them in a computer storage medium (ROM / RAM, magnetic disk, optical disk) for execution by the computing device. In some cases, the steps shown or described can be performed in a different order than those described herein, or they can be fabricated as separate integrated circuit modules, or multiple modules or steps can be fabricated as a single integrated circuit module. Therefore, the present invention is not limited to any particular hardware and software combination.

[0204] The above description, in conjunction with specific embodiments, provides a further detailed explanation of the present invention. It should not be construed that the specific implementation of the present invention is limited to these descriptions. For those skilled in the art, various simple deductions or substitutions can be made without departing from the concept of the present invention, and all such deductions or substitutions should be considered within the scope of protection of the present invention.

Claims

1. A model-free reinforcement learning method for energy demand response management, characterized in that, Includes the following steps: Construct a residential appliance model, which includes dispatchable appliances, non-dispatchable appliances, and plug-in electric vehicles; Social welfare is determined by combining the comprehensive electricity cost of residents and the profit of retailers. Based on the social welfare, the comprehensive electricity cost of residents and the profit of retailers are balanced. Social welfare is expressed as a non-convex optimization problem of housing demand response management based on price. A retailer-resident electricity trading model is constructed based on a model-free reinforcement learning algorithm. The basic element of the model-free reinforcement learning algorithm includes a quintuple. <S,A,R,T t ,γ>, the corresponding retailer-resident electricity trading model is: State S = {s1,...,s} T }, R i,t Energy demand and E i,t-1 The actual electricity consumption is generated by the residence; Action A = {a1,...,a M Retail price ρ i,t Determined by the electricity retailer, M represents the discrete retail price range [ρ min ,ρ max The set number after ]; Reward R = {r1,...,r} T Social welfare F t (P); State transition function T t Related to retail price; The discount factor γ∈[0,1] represents the importance weight of future social welfare; Based on the power grid transmission data, reinforcement learning solutions for the non-convex optimization problem of housing demand response management based on price are constructed using Q-table-based Q-learning algorithm, Q-network algorithm combining deep learning and Q-learning, and Actor-Critic algorithm, respectively. Among them, a reinforcement learning solution for the price-based residential demand response management non-convex optimization problem is constructed using a Q-table-based Q-learning algorithm, including: The q-value is updated based on the Bellman equation and a greedy strategy. The update formula for the q-value is: In the formula, k represents the training index, and lr represents the learning rate; When the Q-table converges, the optimal retail price is obtained using a greedy strategy, as follows: The Q-value function is used to approximate the Q-table, that is: In the formula, α represents the weights of the Q-network; In supervised learning, the q-value of the next time slot is estimated as the label using the network with the current training index k. The target q-value can then be expressed as: The loss function of the Q-network can be expressed as: In the formula, the gradient descent method is used to iteratively update the weight α; The reinforcement learning solution for the non-convex optimization problem of housing demand response management based on price, constructed using a Q-network algorithm combining deep learning and Q-learning, includes: Using DDQN, two identical Q-network structures are used for each appliance: the current Q-network α k and target Q-network And respectively used for decision-making and q-function-based estimation; Based on formula (17), the target q-value is expressed as: In the formula, This represents the q-value of the target Q-network. Let q represent the estimated q-value of the current Q-network, then the loss function is expressed as: Using formulas (19) and (20), electricity retailers determine the current Q-network α. k The system interacts with the residence for F days, storing the observations in an experience replay buffer D. M sets of observations are extracted from D to train the current Q-network, and the target Q-network weights are... Then it is updated at a fixed period C; The reinforcement learning solution for the price-based residential demand response management non-convex optimization problem, constructed using the Actor-Critic algorithm, includes: Add the corresponding Actor policy network to the Q-network, and its policy distribution is as follows: In the formula, Pr represents the probability distribution, and T represents the total time slots; The outputs of the Actor-Critic network are represented as follows: or i,t =φ a (s i,t ,b i ) (22) J i,t =φ c (s i,t ,a i,t ,a i ) (23) In the formula, β i and α i These are the weights of the Actor neural network and the Critic neural network, respectively, φ. a and φ c For the activation function, the estimated action η i,t It is the output of the Actor network, the action-value function J. i,t If the output of the Critic network is in time slot t, then the timing difference error is: TD i,t =F t (P)+γJ i,t+1 -J i,t (24) In the formula, when the discount factor γ = 0, it means that future state values ​​are ignored; when γ = 1, it means that the learning algorithm gives fair consideration to the rewards of all time periods. The loss function of the Critic neural network is then defined as: Using the temporal difference error as the evaluation function of the Actor network, the update formula for the Actor-Critic network based on backpropagation is expressed as: In the formula, la and lc represent the learning rates of Actor and Critic, respectively; Based on the three reinforcement learning solutions to the nonconvex optimization problem of price-based residential demand response management, the optimal solution and the optimal retail price sequence are determined.

2. The model-free reinforcement learning method for energy demand response management according to claim 1, characterized in that, The construction of the residential appliance model specifically includes: Assume the set of dispatchable appliances is N. d ={1,...,D}, and the set of unschedulable appliances is N. n = {1,...,N}, and the set of plug-in electric vehicles is N. p If the set of all electrical appliances is {1,...,P}, then the set of all electrical appliances can be represented as N. z =N d ∪N n ∪N p ={1,…Z}.

3. The model-free reinforcement learning method for energy demand response management according to claim 2, characterized in that, Modeling the schedulable electrical appliances specifically includes: The schedulable electrical appliances d∈N d The actual electricity consumption is expressed as follows: In the formula, T represents the total time slots, and R... d,t It is the expected energy demand of dispatchable appliances, expressed in kWh, E d,t The actual electricity consumption of dispatchable appliances, expressed in kWh, ρ d,t The retail price of electricity determined by retailers, θ t The wholesale price at which retailers purchase energy from the energy market, and ρ d,t ≥θ t δ t It is the price elasticity coefficient, which is less than 0, representing the relationship between energy demand and electricity retail prices; The electricity consumption happiness index of residents is represented by the difference between expected energy demand and actual electricity consumption, i.e., the demand error. The formula is as follows: In the formula, C d,t Let E represent the electricity happiness function, which indicates the actual consumption of electricity by residents during time period t at a specific retail electricity price. d,t The closer to R d,t The stronger the sense of happiness; d1 and h d2 All are happiness coefficients related to electrical appliances; The usable range of the electricity happiness function is limited as follows: In the formula, This is the lower bound of the demand error. This is the upper bound of the demand error.

4. The model-free reinforcement learning method for energy demand response management according to claim 3, characterized in that, Modeling the unschedulable electrical appliances specifically includes: The unschedulable electrical appliances satisfy the constant relationship between energy demand and actual electricity consumption in all time periods, that is: R n,t =E n,t (4) In the formula, R n,t E represents the expected energy demand of non-dispatchable appliances. n,t This indicates the actual power consumption of undispatchable electrical appliances.

5. The model-free reinforcement learning method for energy demand response management according to claim 4, characterized in that, Modeling the plug-in electric vehicle specifically includes: Plug-in electric vehicles p∈N p The actual electricity consumption is expressed as follows: In the formula, R p,t For the anticipated energy demand of plug-in electric vehicles, E p,t This refers to the actual electricity consumption of plug-in electric vehicles, and in E p,t When <0, it indicates discharge, at E p,t When ρ > 0, it indicates charging; p,t For the retail price of electricity for electric vehicles, θ t For wholesale prices, δ t It is the price elasticity coefficient, which is less than 0, representing the relationship between energy demand and electricity retail prices; The electric well-being function C of the plug-in electric vehicle p,t Represented as: In the formula, h d1 and h d2 All are happiness coefficients related to plug-in electric vehicles; The plug-in electric vehicle has a corresponding rated power limit for each time slot, namely: In the formula, This indicates the rated power of a plug-in electric vehicle; Based on the charging and discharging characteristics and battery capacity of the plug-in electric vehicle, the rated power of the plug-in electric vehicle is limited as follows: In the formula, and These represent the minimum and maximum battery capacity, respectively. Represents the initial energy level, e p Indicates charging or discharging efficiency; The cost of battery degradation in the aforementioned plug-in electric vehicle is: You p,t =υ|E p,t | (9) In the formula, υ is the degradation coefficient; Among these restrictions, the retail price of electricity is subject to the following limitations: r min ≤ρ n,t ,r p,t ,r d,t ≤ρ max (10) In the formula, ρ min ρ represents the minimum retail price. max This represents the maximum retail price.

6. The model-free reinforcement learning method for energy demand response management according to claim 5, characterized in that, The process of determining social welfare by combining the comprehensive residential electricity cost and retailer profits, and balancing the comprehensive residential electricity cost and retailer profits based on the social welfare, specifically includes: If the energy consumption of residential appliances is obtained, then the comprehensive electricity cost for the residents is: In the formula, EC t This represents the total cost of all electrical appliances during time period t. For the electricity cost of plug-in electric vehicles, BD p,t This indicates the selling price of electricity generated during battery discharge. If we obtain the total energy consumption of residential electricity, the retail price of electricity, and the wholesale price of electricity, then the retailer's profit is: In the formula, EP t This represents the retailer's profit during time period t; The actual electricity consumption is determined based on retail prices. The goal is to maximize social welfare by balancing residential electricity costs and retailer profits. Maximizing social welfare refers to a weighted balance problem with a two-level optimization objective, namely, a price-based residential demand response management non-convex optimization problem, expressed as: In the formula, n,d,p∈N z ,t∈T,ω is the relative social value weight balancing commercial profits and residential energy consumption, andP is the vector of retail electricity prices consisting of residential appliances.

Citation Information

Patent Citations

  • Method for constructing electricity price optimization model considering electricity retailers and electricity dealers

    CN115062831A

  • Resident dynamic retail price optimization method under day-ahead wholesale electricity price fluctuation conduction

    CN115392963A