A demand-side response load analysis method based on reinforcement learning algorithm
Through the demand-side response load analysis method based on reinforcement learning algorithm, the user load transfer solution is optimized, and the problem that traditional power generation side load control is difficult to cope with peak power consumption is solved, achieving more effective load control and grid stability improvement.
Patent Information
- Application Number
- CN202111369919.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-11-18
- Publication Date
- 2025-06-03
- Estimated Expiration
- 2041-11-18
AI Technical Summary
Traditional power generation side load control is difficult to flexibly deal with peaks and impact loads of electricity, resulting in grid fluctuations and unstable operation.
The demand-side response load analysis method based on reinforcement learning algorithm is adopted, and the user load dynamic transfer problem is converted into discrete infinite Markov decision-making problem through the Q-learning algorithm, and the load transfer scheme is optimized to maximize the discount reward for user load transfer.
More effective load control is achieved, grid fluctuations are reduced, and grid operation is improved.
Smart Images

Figure CN114676949B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of automatic dispatching of power systems, and particularly to a method for analyzing demand-side response load based on a reinforcement learning algorithm. Background Art
[0002] With the increasingly diverse lifestyles, people's demand for electricity is growing, resulting in relatively weak load control on the traditional power generation side and an inability to flexibly cope with the grid fluctuations caused by peak electricity consumption and impact loads. With the application of modern advanced information and communication technologies in smart grid systems, demand response (DR) has become an effective method to improve grid reliability and reduce energy costs by adjusting flexible loads on the user side to quickly respond to supply-demand mismatches. In the case of load-side participation in demand response, there is an urgent need for an effective intelligent control method for load control. Summary of the Invention
[0003] The object of the present invention is to provide a method and device for analyzing demand-side response load based on a reinforcement learning algorithm, which analyzes demand-side load response based on the Q-learning theory, can provide decision-making references for users to participate in demand-side load response, thus more effectively realizing load control, reducing grid fluctuations, and improving the safety and reliability of grid operation. The technical solution adopted by the present invention is as follows.
[0004] On the one hand, the present invention provides a method for analyzing demand-side response load based on a reinforcement learning algorithm, including:
[0005] Obtaining adjustable load information on the power consumption side of the grid, power spot market information, and generator set power generation information on the power supply side;
[0006] Obtaining a pre-constructed demand response model based on user selection;
[0007] Based on the demand response model, converting the user load dynamic transfer problem into a discrete infinite Markov decision problem;
[0008] Based on the adjustable load information on the power consumption side of the grid, power spot market information, and generator set power generation information on the power supply side, using the Q-learning algorithm to solve the discrete infinite Markov decision problem, and taking the maximum discounted reward of user load transfer as the goal to solve the optimal load transfer plan for users to participate in load transfer.
[0009] Optionally, the adjustable load information on the power consumption side includes the critical load and the load that can be cut off of each power user, the power spot market information includes the retail electricity price of each time period, and the generator set power generation information on the power supply side includes the maximum power generation and the unit ramp rate.
[0010] Optionally, the objective function of the demand response model based on user selection is a user load transfer profit objective function F that takes into account load transfer benefits, discomfort costs, and energy supplier incentive benefits;
[0011] The solution constraints of the objective function include user demand constraints, supply and demand balance constraints, load transfer constraints and ramp rate constraints.
[0012] Optionally, the user load transfer profit objective function F is expressed as:
[0013] F=F tran -F sat +F inc
[0014] Among them, F tran represents the load transfer benefit of users in smart grid, F sat represents the user's discomfort cost, F inc represents the incentive benefits given by the energy supplier, and has:
[0015]
[0016] In the formula, For the nth user at t j Time shift to t k The load at the moment; t i ,t j ,t k The retail electricity price at the time; Indicates that the nth user at t j The amount of load removed at any given moment; Indicates that at t k The reward rate for participating in load transfer at any time; α, β, φ are related parameters, with values of (0-1), and n is the user number.
[0017] Optionally, the user demand constraint is expressed as:
[0018]
[0019]
[0020]
[0021] in, t j The electricity demand of the nth user at the moment, t j The actual power consumption of the nth user at the moment; Indicates t j The critical load of the nth user at time, Denote \(P_{n}(t)\) j as the curtailable load of the \(n\)th user at time \(t\);
[0022] The supply-demand balance constraint is expressed as:
[0023]
[0024] where \(E(t)\) max denotes the maximum power generation of the power supply side at time \(t\); j j
[0025] The load transfer constraint is expressed as:
[0026]
[0027] where denotes the curtailable load of the \(n\)th user at time \(t\); i
[0028] The ramp rate constraint is expressed as:
[0029]
[0030] where \(U\) and \(D\) are the upward and downward ramp rates of the power generation side units, respectively.
[0031] Optionally, based on the demand response model, converting the user load dynamic transfer problem into a discrete infinite Markov decision problem includes:
[0032] Simplifying and representing the demand response model based on user selection as:
[0033]
[0034] Determining the decision process elements of the discrete infinite Markov decision problem according to the demand response model includes: discrete time \(t\), the state where the agent is located action \(A\) reward \(r(s'|s, A)\), transition probability t t t and discount factor \(\gamma\);
[0035] where denotes the willingness of the user to transfer load at time \(t\), and \(t\) is the discrete time for performing load transfer; \(\pi(t)\) i denotes the retail electricity price at time \(t\); \(\lambda(t)\) t denotes the reward rate for participating in load transfer at time \(t\); t denotes the maximum power supply of the power supply side at time \(t\); Indicates the electricity consumption of user n at time t; Indicates that user n decides at time t i to transfer to time t j the load at that time, Indicates that user n decides at time t j to transfer to time t k the load at that time; Indicates the decision-making action behavior of all users regarding load transfer at time t; r(s t ′|s t ,A t ) represents the immediate reward obtained at the next time after making a decision in the state at time t; Indicates the state s at discrete time t t transfers to the state s at the next time t ′ of the selected action; γ represents the discount rate, which determines the importance of immediate rewards and future rewards.
[0036] Optionally, the transition probability is determined by a method based on the Boltzmann distribution according to the following formula:
[0037]
[0038] where P(a) represents the probability of choosing a, and τ refers to the temperature parameter. The larger the value of τ, the more the selection strategy tends to a pure random strategy.
[0039] Optionally, in the infinite Markov decision problem, the discounted reward R for user load transfer t is expressed as:
[0040] R t = r(s t+1 |s t ,A t ) + γr(s t+2 |s t+1 ,A t+1 ) + γ 2 r(s t+3 |s t+2 ,A t+2 )… + γ T-t-1 r(s T |s T-1 ,A T-1 )
[0041] In the formula, γ is the discount factor, and the immediate reward obtained at the next time t 1 after making a decision in the state at time t 2 is expressed according to the demand response model as:
[0042]
[0043] Optionally, the discount factor γ ∈ [0, 1], and the value range is 0.2 to 0.8.
[0044] Optionally, using the Q-learning algorithm to solve the discrete infinite Markov decision problem is to maximize the discounted reward of user load transfer. Within the set number of iterations, iterate the following Q-learning algorithm formula to obtain the optimal Q solution, that is, obtain the load transfer amount of each user at each future moment:
[0045]
[0046] Beneficial effects
[0047] Aiming at the problems of large power demand and uneven distribution in the smart grid environment, the present invention analyzes the influencing factors of user load transfer, converts the user load dynamic transfer problem into a discrete infinite Markov decision problem, and then uses the Q-learning theory to optimize the user load transfer scheme to obtain the optimal load transfer amount of grid users at each moment. At the same time, aiming at various constraint problems on the power generation side and user side in the smart grid, a reward mechanism is added during the optimization process, so as to obtain the best Q-learning strategy. The load transfer amount analyzed by the present invention can provide decision support for users to participate in scheduling, and can also provide load forecasting reference for grid power generation control, so as to more effectively achieve load control, reduce grid fluctuations, and improve the safety and reliability of grid operation. Brief description of the drawings
[0048] Figure 1 The figure shows the schematic flow principle of the demand-side response load analysis method based on the reinforcement learning algorithm of the present invention. Detailed implementation manners
[0049] The following is further described in conjunction with the drawings and specific embodiments. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0050] In this embodiment, the demand-side response load analysis method based on the reinforcement learning algorithm includes:
[0051] Obtain the adjustable load information on the power consumption side of the power grid, the power spot market information, and the power generation information of the power supply side units;
[0052] Obtain the demand response model pre-constructed based on user selection;
[0053] Based on the above demand response model, the user load dynamic transfer problem is transformed into a discrete infinite Markov decision problem;
[0054] Based on the adjustable load information on the power grid's electricity consumption side, the power spot market information, and the power generation information of the power supply side units, the Q-learning algorithm is used to solve the discrete infinite Markov decision problem. With the goal of maximizing the discounted reward of user load transfer, the optimal load transfer plan for users to participate in load transfer is obtained.
[0055] Reference Figure 1 , the above method process specifically involves the following content.
[0056] I. Construction of the demand response model
[0057] S1. With the load transfer revenue F tran of users in the smart grid, the discomfort cost F sat and the incentive revenue F inc given by the energy supplier, a transfer profit objective function F is established:
[0058]
[0059] Among them, is the load amount transferred by the nth user from time t j to time t k ; are the retail electricity prices at times t i , t j , and t k respectively; represents the load amount cut off by the nth user at time t j ; represents the reward rate for participating in load transfer at time t k ; α, β, and φ are discomfort function coefficients, with values between 0 and 1, and n is the user number.
[0060] S2. Determine the constraints of the transfer profit objective function F, including user demand constraints, supply-demand balance constraints, load transfer constraints, and ramp rate constraints.
[0061] User demand constraint:
[0062]
[0063]
[0064]
[0065] Among them, is the electricity demand of the nth user at time t j ; For t j The actual power consumption of the nth user at time t; Indicates t j The critical load of the nth user at time t, Indicates t j The load that can be shed by the nth user at time t.
[0066] Supply-demand balance constraint:
[0067]
[0068] Among them, E max (t j ) represents the maximum power generation on the power supply side at time t j .
[0069] Load transfer constraint:
[0070]
[0071] Among them, Indicates t i The load that can be shed by the nth user at time t.
[0072] Ramp rate constraint:
[0073]
[0074] Among them, U and D are the upward and downward ramp rates of the power generation side units respectively.
[0075] II. Converting the user load dynamic transfer problem into a discrete infinite Markov decision problem
[0076] Since the user's dynamic selection is a decision-making problem in a random environment, and the reward and transfer amount only depend on the energy demand, retail electricity price, and participation in regulation reward rate in the current time period, as well as the unit state in the previous time period, and are independent of the previous data, the user load transfer problem can be modeled as a discrete infinite Markov decision process (DIMDP model). In this embodiment, according to the influencing factors of user load transfer, the key elements in the DIMDP model mainly include: discrete time t, the state where the agent is located Action Reward r(s t ′|s t ,A t ), transition probability and discount factor γ.
[0077] The aforementioned demand response model can be simplified and described by the following formula:
[0078]
[0079] Among them, represents the willingness of the user to transfer the load at time t i , where t is the discrete time for implementing the load transfer; π t represents the retail electricity price at time t; λ t represents the reward rate for participating in the load transfer at time t; represents the maximum power supply of the power supply side at time t; represents the electricity consumption of user n at time t; represents that user n decides to transfer to t i at time t j the amount of load transferred to t represents that user n decides to transfer to t j at time t k the amount of load transferred to t represents the decision-making action behavior of all users regarding whether to transfer the load at time t; r(s t ′|s t ,A t ) represents the immediate reward obtained from the next moment after making a decision in the state at time t; represents the state s at discrete time t t transfers to the state s at the next moment t ′ the probability of the selected action; γ represents the discount rate, which determines the importance of immediate rewards and future rewards.
[0080] DIMDP forms an infinite time-step sequence:
[0081] If the time length T is defined, then after running DIMDP once, the total reward at all times for one iteration can be calculated, which can be expressed as:
[0082]
[0083] Among them,
[0084]
[0085] The future return calculated from time t can be expressed as:
[0086]
[0087] In order to make the influence of the more distant future on the current smaller, in this embodiment, a discount coefficient is added, and the future return is expressed as a discounted return, which is expressed as:
[0088] R t = r(s t+1 |s t ,A t) + γr(s t+2 |s t+1 , A t+1 ) + γ 2 r(s t+3 |s t+2 , A t+2 )… + γ T-t-1 r(s T |s T-1 , A T-1 )
[0089] Among them, γ ∈ [0, 1]. Taking 0 emphasizes immediate rewards, taking 1 emphasizes future rewards, and taking values in the range of 0.2 - 0.8 is appropriate.
[0090] The transition probability in the DIMDP model can be determined by a method based on the Boltzmann distribution:
[0091]
[0092] Among them, P(a) corresponds to the probability of choosing a, and τ refers to the temperature parameter. The larger the value of τ, the more the selection strategy tends to a pure random strategy.
[0093] Due to the information technology of the smart grid, the load data of users can be clearly obtained. Based on this, the load side in the load transfer problem is discretized, and the discrete load vector for each user is expressed as:
[0094]
[0095] The above formula corresponds to the load side of a certain user n, which will generate a huge amount of data. Traditional methods are difficult to process and cannot guarantee timeliness. Therefore, the present invention proposes to use the reinforcement learning method to solve the user load transfer decision problem.
[0096] III. Iterative solution of the Q-learning algorithm
[0097] When the present invention is applied, for the demand-side load response load analysis, the adjustable load information on the power consumption side that needs to be obtained includes the critical load and the load that can be cut off of each power user, the power spot market information includes the retail electricity price of each time period, and the power generation information of the power supply side units includes the maximum power generation and the unit ramp rate.
[0098] First, in the initial stage, the present invention initializes the state value S 0 and the action value A 0 of the decision-making agent, indicating the situation where there is no demand response and all users do not participate in load transfer.
[0099] Next, upon receiving the retail electricity price π tand the reward rate λ for participating in load transfer in the current period t After receiving the information, the agent's decision-making behavior starts, and the action A will also generate a reward r. Then, the Q-table is updated through the following formula. Finally, after a certain number of iterations, the update of the Q-table is completed.:
[0100]
[0101] Since the user behavior is model-free, it only needs to meet several constraint conditions of the following demand response model:
[0102]
[0103] And the above constraint conditions have been implicitly considered in the reward value r(s t ′|s t ,A t ) received by the agent. For example, at discrete time t 1 as follows:
[0104]
[0105] Considering the discounted reward:
[0106] R t =r(s t+1 |s t ,A t )+γr(s t+2 |s t+1 ,A t+1 )+γ 2 r(s t+3 |s t+2 ,A t+2 )…+γ T-t-1 r(s T |s T-1 ,A T-1 )
[0107] According to the Q-learning algorithm:
[0108]
[0109] Combining the reward value received by the agent and the discounted reward, iterating the above formula within a certain number of iterations can obtain the optimal solution Q, and then obtain the actions and states of each agent at each future moment, so that the load transfer amounts of each user at each future moment can be calculated according to the aforementioned DIMDP model.
[0110] The specific iterative update steps include:
[0111] 1. At the beginning of the episode, first randomly select the first state s 1 ;
[0112] 2. Then, select action A in state s through the ε-greedy strategy 1 to obtain the next state s 1 and get the reward R 2 . At this time, update the Q-table through Q(s, A) = Q(s, A) + α[R + γQ(s', A') - Q(s, A)]; 2
[0113] 3. After the update, select the A value at different times s through the Q-table to obtain the optimal load transfer scheme.
[0114] Based on the Q-learning algorithm in the above embodiment, in practical applications, it can provide data support for the optimization of the demand response scheme, guide users to better participate in the load regulation of the demand side response, thereby more effectively reducing the power grid fluctuations and improving the safety and reliability of the power grid operation.
[0115] Those skilled in the art should understand that the embodiments of the present application can be provided as methods, systems, or computer program products. Therefore, the present application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0116] The present application is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or block in the flowchart and / or block diagram can be implemented by computer program instructions, and the combination of processes and / or blocks in the flowchart and / or block diagram can also be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in one process Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0117] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including instruction means, and the instruction means implements the functions specified in one process Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0118] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus, so that a series of operation steps are performed on the computer or other programmable apparatus to produce a computer-implemented process, thereby the instructions executed on the computer or other programmable apparatus provide steps for realizing the functions specified in one process or a plurality of processes and / or blocks Figure 1 one process or a plurality of processes and / or blocks Figure 1 steps for realizing the functions specified in one block or a plurality of blocks.
[0119] The embodiments of the present invention have been described above in conjunction with the accompanying drawings. However, the present invention is not limited to the above specific embodiments. The above specific embodiments are merely illustrative rather than restrictive. Under the inspiration of the present invention, those of ordinary skill in the art can also make many forms without departing from the spirit of the present invention and the scope protected by the claims. All of these fall within the protection scope of the present invention.
Claims
1. A method for analyzing demand - side response load based on reinforcement learning algorithm, characterized in that, it includes: Obtain the adjustable load information on the power grid's electricity consumption side, power spot market information, and power generation information of power supply - side units; Obtain a pre - constructed demand response model based on user selection; Based on the demand response model, convert the user load dynamic transfer problem into a discrete infinite Markov decision problem; Based on the adjustable load information on the power grid's electricity consumption side, power spot market information, and power generation information of power supply - side units, use the Q - learning algorithm to solve the discrete infinite Markov decision problem, with the goal of maximizing the discounted reward of user load transfer, and solve to obtain the optimal load transfer plan for users to participate in load transfer; where the adjustable load information on the electricity consumption side includes the critical load and the load that can be cut off of each power user, the power spot market information includes the retail electricity price of each time period, and the power generation information of power supply - side units includes the maximum power generation and the unit ramp rate; The objective function of the demand response model based on user selection is the user load transfer profit objective function considering the load transfer revenue, discomfort cost, and energy supplier incentive revenue. The solution constraints of the objective function include user demand constraints, supply-demand balance constraints, load transfer constraints, and ramp rate constraints. The user load transfer profit objective function is expressed as: , Among them, represents the load transfer benefit of users in the smart grid, represents the discomfort cost of users, represents the incentive benefit given by the energy supplier, and there is: , Wherein, is the load transferred by the -th user at the moment to ; , are respectively the retail electricity prices at , ; represents the load shed by the -th user at ; represents the reward rate for participating in load transfer at ; , , are related parameters, is the user number; The user demand constraint is expressed as: , , , Among them, represents the th user's load transferred to at the moment, is the th user's electricity demand at the moment, and is the th user's actual electricity consumption at the moment; represents the th user's critical load at and represents the th user's cuttable load at the moment; The supply - demand balance constraint is expressed as: , Among them, represents the maximum power generation of the power supply side at a certain moment; The load transfer constraint is expressed as: , Among them, denotes the reducible load of the th user at the moment; The ramp rate constraint is expressed as: , Among them, and are the upward and downward ramp rates of the power generation side units respectively, represents the actual power consumption of the th user at the previous moment of the Based on the demand response model, converting the user load dynamic transfer problem into a discrete infinite Markov decision problem includes: Simplify and represent the demand response model based on user selection as: , According to the demand response model, the decision-making process elements for determining the discrete infinite Markov decision problem include: discrete time , the state of the agent , actions , rewards , transition probabilities and discount factors ; wherein, represents the user's load transfer willingness at moment, is the discrete time for implementing load transfer; , respectively represent the load transfer willingness of the th user at , moments; represents the retail electricity price at moment; represents the reward rate for participating in load transfer at moment; represents the maximum power supply of the power supply side at moment; represents the electricity consumption of user at moment; represents the amount of load transferred by user at moment; represents the decision-making action behavior of all users regarding whether to transfer the load at moment; represents the immediate reward obtained from the next moment after making a decision under the state at moment; represents the state at discrete time the probability of the selected action for transferring to the next moment state; represents the discount rate, which determines the importance of immediate rewards and future rewards.
2. According to the method described in claim 1, characterized in that, The transition probability is determined by a Boltzmann distribution-based method according to the following formula: , Among them, represents the probability of the selection action , refers to the temperature parameter, represents the utility value of selecting the action in state s, represents the utility value of selecting the th optional action in state.
3. According to the method described in claim 1, characterized in that, In the infinite Markov decision problem, the discounted reward of user load transfer is expressed as: , In the formula, is the discount factor, After making a decision in the state at time The immediate reward obtained is expressed according to the demand response model as: , In the formula, represents the th user's action selected at time, and represents the reward rate for participating in load transfer at 4. According to the method described in claim 3, characterized in that, The discount factor has a value range of 0.2 to 0.
8.
5. According to the method described in claim 3, characterized in that, Using the Q - learning algorithm to solve the discrete infinite Markov decision problem means that, with the goal of maximizing the discounted reward of user load transfer, iterate the following Q - learning algorithm formula within the set number of iterations to obtain the optimal Q - solution, that is, obtain the load transfer volume of each user at each future moment: , Wherein, represents the set of user actions at different times, and is the next state of Considering the retail electricity price in the state of the selected action is the utility value of Considering the retail electricity price in the state of the selected action is the utility value of
Citation Information
Patent Citations
Real-time demand response method and device
CN112132350A
Heterogeneous flexible load real-time regulation and control method and device based on deep reinforcement learning
CN112488531A