Power system source-load forward scheduling method and device based on deep reinforcement learning

By adopting a power system source-load look-ahead scheduling method based on deep reinforcement learning, the gap in the existing technology of power system look-ahead optimization scheduling is filled, realizing fast, reliable and intelligent decision-making in power system scheduling, and adapting to the demand-side response of complex power systems.

CN113902176BActive Publication Date: 2025-11-28TSINGHUA UNIVERSITY +3
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202111112177.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-09-18
Publication Date
2025-11-28
Estimated Expiration
2041-09-18

AI Technical Summary

Technical Problem

Existing reinforcement learning techniques have not yet been applied to forward-looking optimization and dispatching of power systems. They are difficult to adapt to the uncertainties brought about by the large number of participants in complex power systems and the increasing penetration rate of new energy sources, which increases the difficulty of optimizing the operation of power systems.

Method used

This paper proposes a power system source-load forward scheduling method based on deep reinforcement learning. By constructing a power system forward scheduling model, designing the state space, action space and reward function, and applying deep reinforcement learning algorithm to improve it, the method obtains basic data on the economic operation of the power system to achieve forward scheduling of demand-side response.

Benefits of technology

It improves the decision-making speed, reliability, and intelligence level of power system dispatch, adapts to the economic optimization dispatch of smart grids with multi-stakeholder participation, and provides a flexible and efficient data-driven solution.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113902176B_ABST
    Figure CN113902176B_ABST
Patent Text Reader

Abstract

The application provides a power system source-load forward scheduling method and device based on deep reinforcement learning, wherein the method comprises the following steps: acquiring power system economic operation basic data, constructing a power system source-load forward scheduling model according to the power system economic operation basic data, so as to construct a power system forward scheduling model containing demand side response; based on the power system forward scheduling model, designing a state space, an action space and a reward function, so as to design a time sequence decision mechanism of the power system economic dispatching problem; according to the time sequence decision mechanism, applying a deep reinforcement learning algorithm to the power system forward scheduling model, and improving and applying the deep reinforcement learning algorithm, so as to obtain a forward scheduling strategy based on deep reinforcement learning. The application provides a solution for the economic optimal dispatching of the smart grid with sufficient interaction between supply and demand, a large number of subjects participating and uncertainty being improved, and improves the decision speed, reliability, automation and intelligent level of the power system dispatching.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of power system optimal scheduling and reinforcement learning, and particularly relates to a power system source-load forward scheduling method and device based on deep reinforcement learning. BACKGROUND

[0002] With the gradual advancement of the construction of new power systems in China, the traditional power grid is gradually developing into a complex power system with a large number of participants, and the strengthening of source-load interaction significantly increases the number of participants in the power system operation. In addition, the increasing penetration rate of new energy over the years also brings a certain degree of uncertainty to the operation of the power system, increasing the difficulty of the optimal operation of the power system. The traditional manual day-ahead scheduling method is difficult to adapt to this new change, and a more flexible and efficient data-driven method provides a feasible solution to the operation of future smart grids, such as reinforcement learning algorithms.

[0003] Existing research has applied reinforcement learning technology to some directions of smart grid operation and management. In the field of smart microgrids, research has applied reinforcement learning algorithms to energy storage management strategies for smart microgrids. Some literature has applied reinforcement learning algorithms to energy storage management in electric-thermal integrated energy systems containing renewable energy, forming a medium and long-term sustainable automated energy management strategy. In the field of demand side response, research has also applied reinforcement learning algorithms to the management and pricing strategies of demand side response participants. Some research has applied reinforcement learning algorithms to price-based demand side response pricing, and the pricing strategy generated by the agent can improve system robustness and reduce the cost of load service providers. In the field of power system scheduling, demand side response is also applied to the generation of intelligent scheduling strategies with high real-time performance. Some literature has applied reinforcement learning algorithms to multi-objective optimal scheduling of power systems containing renewable energy to minimize system operation cost and maximize renewable energy consumption.

[0004] Existing research on the application of reinforcement learning to related problems in the power system mainly focuses on the above three aspects, and there is no literature on the application of reinforcement learning to forward optimal scheduling of the power system. SUMMARY

[0005] The present application aims to at least partially solve one of the technical problems in the related art.

[0006] To this end, the purpose of the present application is to fill the gap in the application of reinforcement learning to forward optimal scheduling of the power system, and to propose a power system source-load forward scheduling method based on deep reinforcement learning. The present application applies deep reinforcement learning to forward optimal scheduling of the power system and considers demand side response, thereby providing a solution to the economic optimal scheduling of smart grids with full interaction between supply and demand, a large number of participants, and increasing uncertainty.

[0007] The second object of the present application is to provide a power system source-load forward scheduling device based on deep reinforcement learning.

[0008] To achieve the above object, the first aspect of the present application provides a power system source-load forward scheduling method based on deep reinforcement learning, comprising:

[0009] acquiring power system economic operation basic data, and constructing a power system source-load forward scheduling model according to the power system economic operation basic data, so as to construct a power system forward scheduling model containing demand side response;

[0010] designing a state space, an action space and a reward function based on the power system forward scheduling model, so as to design a timing decision mechanism of the power system economic dispatching problem;

[0011] applying a deep reinforcement learning algorithm to the power system forward scheduling model according to the timing decision mechanism, and improving and applying the deep reinforcement learning algorithm, so as to obtain a forward scheduling strategy based on deep reinforcement learning.

[0012] In addition, the power system source-load forward scheduling method based on deep reinforcement learning according to the above embodiments of the present application can have the following additional technical features:

[0013] Further, in an embodiment of the present application, the power system economic operation basic data comprises:

[0014] a unit output upper and lower limit, a unit climbing increase and decrease rate upper and lower limit, a unit cost function, a demand side response load upper limit and a demand side response price function.

[0015] Further, in an embodiment of the present application, the power system source-load forward scheduling model is constructed according to the power system economic operation basic data, comprising:

[0016] 1) establishing a power system operation constraint condition, the expression is as follows:

[0017]

[0018]

[0019]

[0020]

[0021] wherein (1) is a system power balance constraint; in the formula is the output of the generator unit i at the time period t, N g is the number of adjustable units; is the load of bus j in the grid at time period t, N b is the number of buses in the grid; is the load curtailment of demand side response subject k at time period t, N dr is the total number of demand side response subjects;

[0022] Equation (2) is the generator unit output constraint; in which is the upper and lower limit of the output of generator unit i;

[0023] Equation (3) is the increase and decrease output rate constraint of each unit; in which are the upper limits of the increase and decrease output of unit i in adjacent time periods; in which Δt is the unit time interval;

[0024] Equation (4) is the load curtailment constraint of each demand side response subject; in which αd r, i is the maximum load curtailment ratio of demand side response subject i, B(i) represents the system bus number where the demand side response subject i is located, is the maximum load curtailment amount thereof at time t;

[0025] Among the above four constraints, time t represents any time of the look-ahead window 0, 1,..., T-1;

[0026] 2) Determine the economic dispatch target function of the power system, the expression is as follows:

[0027]

[0028]

[0029]

[0030]

[0031]

[0032] Among them, (5) is the target function; the total operating cost in the look-ahead window is minimum, including the total operating cost (6) of the generator unit and the total cost (7) of the demand side response;

[0033] Equation (6) is the total operating cost of the generator unit, which is obtained by summing the cost function (8) of each generator; Equation (8) is the cost function of each generator, which adopts the form of a quadratic function, a g,i , b g,i and c g,i are the coefficients thereof;

[0034] Equation (7) is the total cost of the demand side response, which is obtained by summing the cost function (9) of each demand side response subject; Equation (9) is the cost function of each demand side response subject, which adopts the form of Ki +1 segment function form, is the slope of each segment, is the intercept of each segment, is the segment point.

[0035] Further, in an embodiment of the present application, the design state space is expressed as follows:

[0036]

[0037]

[0038]

[0039]

[0040]

[0041]

[0042] wherein (10) is the state vector definition, the state quantity includes the generator set output at the last time (11), the bus load in the look-ahead window (12) and the current time t (15);

[0043] Equation (11) is the state vector of the generator set output at time t, including the output state value (13) of all N g generators;

[0044] Equation (12) is the state vector of the bus load at time t, including the load state value (14) of all N b busbars;

[0045] Equations (13), (14) and (15) all use the following normalization function to normalize the named value according to the corresponding upper and lower bounds:

[0046]

[0047] wherein x is the named value, is the normalization result, L is the lower bound of x, and U is the upper bound;

[0048] In equation (14), and are the maximum and minimum values of the load of bus j in the entire training period 0, 1, …, T train -1, and the and the upper and lower limits are only related to the load condition at the training time, and the and the upper and lower limit values still need to be used in testing and application; in equation (15), Ttrain is the total training length.

[0049] Further, in one embodiment of the present application, the design action space is expressed as follows:

[0050]

[0051]

[0052]

[0053]

[0054]

[0055] where (17) is the action vector definition, the action quantity contains the generator output (18) and the demand side response load shedding (19) of all time points within the look-ahead window;

[0056] (18) is the generator output action vector at time t, containing the output action value (20) of all N g -1 generators except the balancing unit, the output of the balancing unit is calculated according to the system power balance constraint (1);

[0057] (19) is the demand side response subject load shedding action vector at time t, containing the load shedding action value (21) of all N dr subjects;

[0058] Both (20) and (21) use the normalization function (16) to normalize the named value according to the corresponding upper and lower bounds.

[0059] Further, in one embodiment of the present application, the reward function is designed as follows:

[0060]

[0061]

[0062]

[0063] R t = (1 - I (M t , M t+1 ,..., M t+T-1 )) R (25)

[0064]

[0065]

[0066] wherein (22) is the reward function definition, including the total operation cost (23) in the look-ahead window, the penalty term (24) and the reward term (25);

[0067] (23) is the system operation cost at time t, including the generator operation cost and the demand side response cost;

[0068] (24) is the penalty term, wherein M g and M r are the penalty coefficients of the unit output constraint (2) and the ramp constraint (3), and are defined by (26) and (27), respectively, which are the out-of-limit values of the unit output constraint and the ramp constraint;

[0069] (25) is the reward term, wherein I(·) is a logic function: if M t ,M t+1 ,…,M t+T-1 are all 0, there is no out-of-limit situation in all time points in the look-ahead window, then I = 0, R t = R is a positive reward value; if there is an out-of-limit situation in the look-ahead window, then I = 1, R t = 0 without reward.

[0070] Further, in an embodiment of the present application, the deep reinforcement learning algorithm is improved and applied, including: pre-training an agent, training an agent, testing and applying an agent, wherein the deep reinforcement learning algorithm adopts a deep deterministic policy gradient algorithm.

[0071] Further, in an embodiment of the present application, the pre-trained agent comprises:

[0072] preparing pre-training data, converting real historical scheduling data according to the state space, the action space and the reward function definition for agent training; and,

[0073] pre-training the action and evaluation networks respectively, and initializing the expert network using the same parameters.

[0074] Further, in an embodiment of the present application, the trained agent comprises:

[0075] allowing the agent to interact with the environment in the time series decision-making process, and storing all interaction experiences in an experience replay pool, and additionally storing the experiences without out-of-limit into a separate experience replay pool; and,

[0076] every certain number of decisions, randomly extracting experience samples from the two experience replay pools to update the agent network, and updating the expert network parameters.

[0077] The power system source-load forward scheduling method based on deep reinforcement learning of the embodiment of the application, by acquiring power system economic operation basic data, constructing a power system source-load forward scheduling model according to the power system economic operation basic data, to construct a power system forward scheduling model containing demand side response; and based on the power system forward scheduling model, designing a state space, an action space and a reward function, to design a timing decision mechanism of the power system economic dispatching problem; according to the timing decision mechanism, applying a deep reinforcement learning algorithm to the power system forward scheduling model, and improving and applying the deep reinforcement learning algorithm, to obtain a forward scheduling strategy based on deep reinforcement learning. The application of deep reinforcement learning to the power system forward optimization scheduling is considered, and the demand side response can serve the multi-agent participated intelligent power grid forward optimization scheduling, which is beneficial to improving the decision speed, reliability, automation and intelligent level of power system scheduling.

[0078] To achieve the above purpose, the second aspect of the embodiment of the application provides a power system source-load forward scheduling device based on deep reinforcement learning, comprising:

[0079] The construction module is configured to acquire power system economic operation basic data, construct a power system source-load forward scheduling model according to the power system economic operation basic data, to construct a power system forward scheduling model containing demand side response; and based on the power system forward scheduling model, design a state space, an action space and a reward function, to design a timing decision mechanism of the power system economic dispatching problem;

[0080] The optimization module is configured to apply a deep reinforcement learning algorithm to the power system forward scheduling model according to the timing decision mechanism, and improve and apply the deep reinforcement learning algorithm, to obtain a forward scheduling strategy based on deep reinforcement learning.

[0081] The power system source-load forward scheduling device based on deep reinforcement learning of the embodiment of the application, by the construction module, is configured to acquire power system economic operation basic data, construct a power system source-load forward scheduling model according to the power system economic operation basic data, to construct a power system forward scheduling model containing demand side response; and based on the power system forward scheduling model, design a state space, an action space and a reward function, to design a timing decision mechanism of the power system economic dispatching problem; the optimization module is configured to apply a deep reinforcement learning algorithm to the power system forward scheduling model according to the timing decision mechanism, and improve and apply the deep reinforcement learning algorithm, to obtain a forward scheduling strategy based on deep reinforcement learning. The application of deep reinforcement learning to the power system forward optimization scheduling is considered, and the demand side response can serve the multi-agent participated intelligent power grid forward optimization scheduling, which is beneficial to improving the decision speed, reliability, automation and intelligent level of power system scheduling.

[0082] Additional aspects and advantages of the present application will be apparent from the following description, taken in conjunction with the accompanying drawings, wherein: BRIEF DESCRIPTION OF DRAWINGS

[0083] The above and / or additional aspects and advantages of the present application will become apparent and be readily appreciated from the following description, taken in conjunction with the accompanying drawings, in which:

[0084] Figure 1 A flow chart of a deep reinforcement learning based power system source-load look-ahead scheduling method according to an embodiment of the present application;

[0085] Figure 2 A flow chart of a deep reinforcement learning based power system source-load look-ahead scheduling method according to an embodiment of the present application;

[0086] Figure 3 A structural schematic diagram of a deep reinforcement learning based power system source-load look-ahead scheduling device according to an embodiment of the present application. DETAILED DESCRIPTION

[0087] Embodiments of the present application are described in detail below with reference to the accompanying drawings, in which like or similar elements are denoted by the same or similar reference signs, and examples of the embodiments are shown in the drawings. The embodiments described below by reference to the drawings are exemplary and are intended to explain the present application, and are not to be understood as limiting the present application.

[0088] A deep reinforcement learning based power system source-load look-ahead scheduling method and device according to an embodiment of the present application are described below with reference to the accompanying drawings.

[0089] Figure 1 A flow chart of a large-scale data classification method based on hypergraph structure provided by an embodiment of the present application.

[0090] As shown in the figure, the large-scale data classification method based on hypergraph structure includes the following steps: Figure 1

[0091] Step S1, obtaining power system economic operation basic data, constructing a power system source-load look-ahead scheduling model according to the power system economic operation basic data, to construct a power system look-ahead scheduling model containing demand side response.

[0092] Specifically, as shown in the figure: Figure 2

[0093] 1) Constructing a power system look-ahead scheduling model containing demand side response, including 2 steps: obtaining power system economic operation basic data, constructing a power system source-load look-ahead scheduling model; ​​

[0094] 1-1) Obtain the economic operation basic data of the power system:

[0095] The economic operation basic data of the power system includes the upper and lower limits of unit output, the upper and lower limits of unit ramping rate, the unit cost function, the upper limit of demand side response load, and the demand side response price function.

[0096] 1-2) Construct a power system source-load forward scheduling model:

[0097] 1-2-1) Establish the power system operation constraints, the expression is as follows:

[0098]

[0099]

[0100]

[0101]

[0102] wherein (1) is the system power balance constraint; in the formula is the output of the generator unit i at time period t, N g is the number of dispatchable units; is the load of bus j in the power grid at time period t, N b is the number of buses in the power grid; is the load reduced by the demand side response subject k at time period t, N dr is the total number of demand side response subjects;

[0103] Equation (2) is the generator unit output constraint; in the formula is the upper and lower limits of the output of the generator unit i;

[0104] Equation (3) is the increase and decrease output rate constraint of each unit; in the formula are respectively the upper limits of the increase and decrease output of unit i in adjacent time periods; in the formula Δt is the unit time interval;

[0105] Equation (4) is the load reduction constraint of each demand side response subject; in the formula dr,i is the maximum load reduction rate of the demand side response subject i, B(i) represents the system bus number where the demand side response subject i is located, that is, the maximum load reduction amount thereof at time t;

[0106] In the above four constraints, unless otherwise specified, time t represents any time of the forward window 0, 1,..., T-1;

[0107] 1-2-2) Determine the economic dispatching objective function of the power system, the expression is as follows:

[0108]

[0109]

[0110]

[0111]

[0112]

[0113] where (5) is the objective function, i.e., the minimum total operation cost in the look-ahead window, including the total operation cost of generators (6) and the total cost of demand side response (7);

[0114] (6) is the total operation cost of generators, which is obtained by summing the cost function (8) of each generator; (8) is the cost function of each generator, which is in the form of a quadratic function, a g,i , b g,i and c g,i are its coefficients;

[0115] (7) is the total cost of demand side response, which is obtained by summing the cost function (9) of each demand side response subject; (9) is the cost function of each demand side response subject, which is in the form of a piecewise function of K i +1 segment, is the slope of each segment, is the intercept of each segment, is the segmentation point.

[0116] Step S2, based on the look-ahead scheduling model of the power system, design the state space, action space and reward function to design the time sequence decision mechanism of the economic dispatching problem of the power system.

[0117] Specifically, as Figure 2 shown:

[0118] 2) Design the time sequence decision mechanism of the economic dispatching problem of the power system, including 3 steps: design the state space, design the action space, and design the reward function;

[0119] 2-1) Design the state space, the expression is as follows:

[0120]

[0121]

[0122]

[0123]

[0124]

[0125]

[0126] where (10) is the state vector definition, the state quantities include the generator output at the last time (11), the bus load in the look-ahead window (12) and the current time t (15);

[0127] Equation (11) is the state vector of the generator output at time t, which includes the output state value (13) of all N g generators;

[0128] Equation (12) is the state vector of the bus load at time t, which includes the load state value (14) of all N b busbars;

[0129] Equations (13), (14) and (15) all use the following normalization function to normalize the named value according to the corresponding upper and lower limits:

[0130]

[0131] where x is the named value, is the normalized result, L is the lower limit of x, and U is the upper limit;

[0132] In equation (14), and are the maximum and minimum values of the load of bus j in the entire training period 0, 1, …, T train -1, which are only related to the load condition at the training time, and the upper and lower limits still need to be used in testing and application; in equation (15), T train is the total training time;

[0133] 2-2) Design the action space, the expression is as follows:

[0134]

[0135]

[0136]

[0137]

[0138]

[0139] where (17) is the action vector definition, the action quantities include the generator output (18) and the demand side response load reduction (19) at all times in the look-ahead window;

[0140] (18) is the generation units' output action vector at time t, containing all N g -1 generator's output action value (20), the output of balancing units is not given by the algorithm, but is calculated according to the system power balance constraint (1);

[0141] (19) is the demand side response subject's load reduction action vector at time t, containing all N dr subject's load reduction action value (21);

[0142] Both formula (20) and formula (21) use the normalization function of formula (16) to normalize the named value according to the corresponding upper and lower bounds;

[0143] 2-3) Design the reward function, the expression is as follows:

[0144]

[0145]

[0146]

[0147] R t =(1-I(M t ,M t+1 ,...,M t+T-1 ))R (25)

[0148]

[0149]

[0150] Wherein (22) is the definition of the reward function, containing the total operation cost (23) in the look-ahead window, the penalty term (24) and the reward term (25); Since the optimization goal of the agent is to maximize the reward, the operation cost term in the formula has a negative sign to minimize the total cost;

[0151] (23) is the system operation cost at time t, containing the generation unit operation cost and the demand side response cost;

[0152] (24) is the penalty term, wherein M g and M r are the penalty coefficients of the unit output constraint (2) and the ramp constraint (3), and are defined by formula (26) and formula (27), which are the out-of-bound values of the unit output constraint and the ramp constraint, respectively;

[0153] (25) is the reward term, wherein I(·) is a logic function: if Mt,M t+1 ,…,M t+T-1I = 0, R = 0 if all the time steps in the look-ahead window are in-bounds, i.e. no out-of-bounds situation occurs in the look-ahead window, then I = 0, R t = R is a positive reward value; if there is an out-of-bounds situation in the look-ahead window, then I = 1, R t = 0 no reward.

[0154] Step S3, according to the time sequence decision mechanism, the deep reinforcement learning algorithm is applied to the power system look-ahead scheduling model, and the deep reinforcement learning algorithm is improved and applied, and the look-ahead scheduling strategy based on the deep reinforcement learning is obtained.

[0155] Specifically, as Figure 2 shown:

[0156] 3) Application and improvement of deep reinforcement learning algorithm, including 3 steps: pre-training, training, testing and application; this patent adopts deep deterministic policy gradient algorithm (DDPG) as the deep reinforcement learning algorithm, and the following steps are for this algorithm;

[0157] 3-1) Pre-training of agent: before formally training the agent, the network of the agent needs to be pre-trained using historical data to initialize its parameters and speed up the convergence of formal training:

[0158] 3-1-1) Prepare pre-training data, convert the real historical scheduling data according to the state space, action space and reward function definition in step 2) for agent training;

[0159] 3-1-2) Use gradient descent method and other methods to pre-train the action and evaluation networks, and use the same parameters to initialize the expert network;

[0160] 3-2) Training of agent: according to the definition of DDPG algorithm, the agent is placed in the environment and allowed to learn experience in interaction with the environment;

[0161] 3-2-1) Allow the agent to interact with the environment in the time sequence decision process, and store all interaction experiences in the experience replay pool, and store the experiences that do not exceed the limit into a separate experience replay pool;

[0162] 3-2-2) Every certain number of decisions, randomly sample experience samples from the two experience replay pools to update the agent network, and update the expert network parameters;

[0163] 3-3) Test and apply the agent: after the agent is trained, it only needs to be put into the environment again to interact with it, and the decision of each step can be collected, that is, the look-ahead scheduling strategy based on deep reinforcement learning is obtained.

[0164] The proposed power system source-load forward scheduling method based on deep reinforcement learning of the embodiment of the present application obtains power system economic operation basic data, constructs a power system source-load forward scheduling model according to the power system economic operation basic data, to construct a power system forward scheduling model containing demand side response; and based on the power system forward scheduling model, designs a state space, an action space and a reward function, to design a timing decision mechanism of the power system economic dispatching problem; according to the timing decision mechanism, applies a deep reinforcement learning algorithm to the power system forward scheduling model, and improves and applies the deep reinforcement learning algorithm, to obtain a forward scheduling strategy based on deep reinforcement learning. The present application applies deep reinforcement learning to power system forward optimization scheduling, and considers demand side response, and can serve the multi-agent participated intelligent power grid forward optimization scheduling, and is beneficial to improving the decision speed, reliability, automation and intelligent level of power system scheduling.

[0165] Figure 3 The structure diagram of the power system source-load forward scheduling device based on deep reinforcement learning according to one embodiment of the present application.

[0166] As shown in Figure 3 , the device 10 comprises a construction module 100 and an optimization module 200.

[0167] The construction module 100 is used for obtaining power system economic operation basic data, constructing a power system source-load forward scheduling model according to the power system economic operation basic data, to construct a power system forward scheduling model containing demand side response; and based on the power system forward scheduling model, designing a state space, an action space and a reward function, to design a timing decision mechanism of the power system economic dispatching problem;

[0168] The optimization module 200 is used for applying a deep reinforcement learning algorithm to the power system forward scheduling model according to the timing decision mechanism, and improving and applying the deep reinforcement learning algorithm, to obtain a forward scheduling strategy based on deep reinforcement learning.

[0169] The power system source-load forward scheduling device based on deep reinforcement learning according to the embodiment of the present application comprises a construction module, an optimization module and a deep reinforcement learning algorithm.

[0170] In addition, the terms "first", "second", "third", etc. are used herein only to describe various circumstances, and cannot be understood as indicating or implying relative importance or implicitly indicating the number of the technical features indicated. Therefore, the features defined with "first", "second", etc. can explicitly or implicitly include at least one of the features. In the description of the present application, the meaning of "a plurality of" is at least two, for example, two, three, etc., unless otherwise explicitly and specifically limited.

[0171] In the description of the present application, the description of the terms "one embodiment", "some embodiments", "an example", "a specific example", or "some examples" means that the specific features, structures, materials or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present application. In the present application, the illustrative description of the above terms is not necessarily directed to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in any one or more embodiments or examples in a suitable manner. In addition, the person skilled in the art can combine and combine the different embodiments or examples described in the present application and the features of the different embodiments or examples without contradiction.

[0172] Although the embodiments of the present application have been shown and described above, it can be understood that the above embodiments are exemplary and cannot be understood as limiting the present application, and the person skilled in the art can make changes, modifications, replacements and variations to the above embodiments within the scope of the present application.

Claims

1. A source-load look-ahead scheduling method for power systems based on deep reinforcement learning, characterized in that, The method includes the following steps: Acquire basic data on the economic operation of the power system, and construct a power system source-load forward scheduling model based on the aforementioned basic data to build a power system forward scheduling model incorporating demand-side response; and, Based on the aforementioned power system forward scheduling model, a state space, action space, and reward function are designed to formulate a time-series decision-making mechanism for the economic scheduling problem of the power system. Based on the aforementioned time-series decision-making mechanism, a deep reinforcement learning algorithm is applied to the power system forward scheduling model, and the deep reinforcement learning algorithm is improved and applied to obtain a forward scheduling strategy based on deep reinforcement learning. The basic data for the economic operation of the power system include: Unit output upper and lower limits, unit ramp-up and deceleration rate upper and lower limits, unit cost function, demand-side response load upper limit, and demand-side response price function; The construction of the power system source-load forward scheduling model based on the basic data of power system economic operation includes: 1) Establish the power system operation constraints, expressed as follows: Where (1) is the system power balance constraint; in the formula N is the output of generator unit i in time period t. g This represents the number of dispatchable units. N is the load of bus j in the power grid during time period t. b This refers to the number of busbars in the power grid. N represents the load reduction by demand-side response subject k during time period t. dr This represents the total number of entities responding to demand. Equation (2) represents the output constraint of the generator set; where These are the upper and lower limits of the output of generator set i; Equation (3) represents the constraint on the rate of increase and decrease in output for each unit; where These represent the upper limit of output increase and decrease for unit i in adjacent time periods, respectively; where Δt is the unit time interval. Equation (4) represents the load reduction constraints for each demand-side response entity; where α dr,i Let B(i) be the maximum load reduction ratio for demand-side response subject i, and B(i) be the system bus number where demand-side response subject i is located. It is its maximum load reduction at time t; In the above four constraints, time t represents any time in the look-ahead window 0, 1, ..., T-1; 2) Determine the objective function for economic dispatch of the power system, as shown in the following expression: (5) is the objective function, which is to minimize the total operating cost within the look-ahead window, including the total operating cost of generator sets (6) and the total demand-side response cost (7). Equation (6) represents the total operating cost of the generator set, obtained by summing the cost functions (8) of each generator; Equation (8) represents the cost function of each generator, which takes the form of a quadratic function, a g,i b g,i With c g,i Its coefficient; Equation (7) represents the total cost of demand-side response, obtained by summing the cost functions (9) of each demand-side response entity; Equation (9) represents the cost function of each demand-side response entity, using K... i +1 segment piecewise function form, The slope of each segment, For each segment intercept, These are the segmentation points; The design state space is expressed as follows: Where (10) is the definition of the state vector, and the state variables include the generator output at the previous moment (11), the bus load in the look-ahead window (12), and the current moment t (15); Equation (11) is the generator output state vector at time t, which includes all N g Output status value of the generator (13); Equation (12) is the bus load state vector at time t, which includes all N b Load status value of busbar (14); Equations (13), (14), and (15) all use the following normalization function to normalize the named values ​​according to their corresponding upper and lower bounds: Where x is a named value. For the normalized result, L is the lower bound of x, and U is the upper bound; In equation (14), and For the bus j throughout the training period 0, 1, ..., T train The maximum and minimum load values ​​in -1, the With the The upper and lower limits are only related to the load conditions at the training time; the actual values ​​still need to be used during testing and application. With the The upper and lower limits; T in equation (15) train Total training duration; Design the action space, expressed as follows: (17) is the definition of the action vector, and the action quantity includes the generator output (18) and the demand-side response load reduction (19) at all times within the look-ahead window; Equation (18) is the generator output action vector at time t, which includes all N except for the balancing unit. g -1 The output action value of the generator (20) is calculated based on the system power balance constraint (1); Equation (19) is the load reduction action vector of the demand-side response entity at time t, which includes all N dr The load reduction action value of each main body (21); Both equations (20) and (21) use the normalization function of equation (16) to normalize the named values ​​according to the corresponding upper and lower bounds; Design a reward function, expressed as follows: R t =(1-I(M t ,M t+1 ,…,M t+T-1 ))R (25) (22) is the definition of the reward function, which includes minimizing the total running cost within the look-ahead window (23), the penalty (24), and the reward (25); Equation (23) represents the system operating cost at time t, which includes the generator set operating cost and the demand-side response cost; Equation (24) is the penalty, where M g With M r The penalty coefficients for the unit output constraint (2) and the ramp constraint (3) are given. and Defined by equations (26) and (27), these are the out-of-bounds values ​​of the unit output constraint and the ramp constraint, respectively; Equation (25) is the reward term, where I(·) is a logical function: if M t M t+1 ,…,M t+T-1 If all values ​​are 0, meaning there are no out-of-bounds situations at any time within the look-ahead window, then I = 0, R t =R is a positive reward value; if there is an out-of-limit situation within the lookahead window, then I=1, R t =0 No reward.

2. The power system source-load look-ahead scheduling method based on deep reinforcement learning according to claim 1, characterized in that, The deep reinforcement learning algorithm is improved and applied, including: a pre-trained agent, a training agent, and a testing and application agent, wherein the deep reinforcement learning algorithm adopts a deep deterministic policy gradient algorithm.

3. The power system source-load look-ahead scheduling method based on deep reinforcement learning according to claim 2, characterized in that, The pre-trained agent includes: Prepare pre-training data by transforming real historical scheduling data according to the defined state space, action space, and reward function for the agent's training; and, The action and evaluation networks were pre-trained separately, and the expert network was initialized using the same parameters.

4. The power system source-load look-ahead scheduling method based on deep reinforcement learning according to claim 2, wherein the training agent comprises: The agent interacts with the environment during the time-series decision-making process and stores all interaction experiences in the experience replay pool. Experiences that do not exceed the limits are stored separately in a separate experience replay pool. as well as, After a certain number of decisions, experience samples are randomly drawn from two experience replay pools to update the agent network and the expert network parameters.

5. A power system source-load look-ahead scheduling device based on deep reinforcement learning using the method described in claim 1, characterized in that, include: The module is used to acquire basic data on the economic operation of the power system and to construct a power system source-load forward scheduling model based on the basic data on the economic operation of the power system, so as to construct a power system forward scheduling model including demand-side response. Based on the aforementioned power system forward scheduling model, a state space, action space, and reward function are designed to formulate a time-series decision-making mechanism for the economic scheduling problem of the power system. The optimization module is used to apply the deep reinforcement learning algorithm to the power system forward scheduling model according to the time-series decision mechanism, and to improve and apply the deep reinforcement learning algorithm to obtain a forward scheduling strategy based on deep reinforcement learning.