Micro-grid economic dispatching method and system based on deep reinforcement learning
Through the microgrid economic scheduling method based on deep reinforcement learning, combined with multi-time scale optimization strategies and deep neural network modeling, the economic and reliability problems caused by uncertainty in wind and light power generation in the microgrid are solved, and the economic optimal scheduling of the microgrid is achieved.
Patent Information
- Application Number
- CN202410169567.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-02-06
- Publication Date
- 2025-08-08
AI Technical Summary
The prior art has affected economics, reliability and safety in microgrids due to the modeling dependence on distributed power supplies, making it difficult to effectively deal with the uncertainty of new energy power generation such as wind and light.
The economic scheduling method of microgrid based on deep reinforcement learning is adopted. By establishing an economic optimization scheduling model in the recent scheduling stage, using deep neural networks to model Markov decision-making process in the intraday scheduling stage, combining multi-time scale optimization strategies, taking into account the prediction errors of scenery prediction data and important loads, the economic optimal scheduling of the microgrid is achieved.
It improves the accuracy and solution efficiency of the microgrid economic scheduling, effectively solves the dependence problem on the model, and realizes the optimal economic scheduling of the microgrid.
Smart Images

Figure CN120454069A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of microgrid dispatching technology, and in particular to a microgrid economic dispatching method and system based on deep reinforcement learning. Background Art
[0002] The global shortage of fossil energy and environmental challenges are driving the rapid development and application of distributed power sources, primarily based on renewable energy. However, due to the inherent characteristics of distributed power sources, such as randomness, volatility, and intermittency, they can impact the larger power grid when connected, significantly impacting the grid's economic viability, reliability, and security. Most current mainstream algorithms rely on modeling the components of microgrids, and the accuracy of these models directly impacts the accuracy of the overall algorithm's performance.
[0003] With the rapid development of artificial intelligence, reinforcement learning has entered the public eye. Deep reinforcement learning, which combines it with neural networks, can be applied to microgrid economic dispatch. It can not only effectively solve the problem of model dependence, but also greatly improve the efficiency of solution. Therefore, it is very necessary to study the microgrid economic dispatch method based on deep reinforcement learning. Summary of the Invention
[0004] The purpose of the present invention is to address the uncertainty problem caused by the volatility of renewable energy power generation such as wind and solar power, and to provide a microgrid economic dispatch method and system based on deep reinforcement learning to achieve economic optimal dispatch of microgrids.
[0005] To achieve the above objectives, the present invention provides a microgrid economic dispatch method based on deep reinforcement learning, comprising:
[0006] In the day-ahead dispatch phase, based on the day-ahead forecast data of the uncontrollable units, an economic optimization dispatch model for the microgrid is established and solved to obtain the day-ahead dispatch plan for each device in the microgrid.
[0007] In the intraday pre-scheduling stage, based on the day-ahead forecast data of the uncontrollable units and the economic optimization scheduling model, a deep neural network is used to perform Markov decision process modeling to obtain an intraday scheduling model;
[0008] In the intraday scheduling stage, the intraday ultra-short-term forecast data and the day-ahead scheduling plan data of the uncontrollable units are input into the intraday scheduling model to obtain output data, and the output data is used as the scheduling basis for the controllable units;
[0009] Among them, the microgrid includes uncontrollable units and controllable units.
[0010] Furthermore, in the day-ahead dispatch phase, based on the day-ahead forecast data of the uncontrollable units, an economic optimization dispatch model for the microgrid is established and solved to obtain the day-ahead dispatch plan, including:
[0011] Based on the day-ahead forecast data of uncontrollable units, the economic dispatch model of the microgrid is established by considering the microgrid operation cost and the demand response of interruptible loads.
[0012] Based on the economic dispatch model of the microgrid, a mixed integer programming method is used to solve the scheduling plan of each device in the microgrid during the day-ahead scheduling phase, namely the day-ahead scheduling plan;
[0013] Wherein, the microgrid is an AC / DC hybrid microgrid;
[0014] In the day-ahead scheduling phase, the uncontrollable units include wind turbines, photovoltaics and important loads, and the important loads include AC important loads and DC important loads; the controllable units include lithium batteries, fuel cells and interruptible loads.
[0015] Furthermore, based on the day-ahead forecast data of uncontrollable units, and taking into account the microgrid operation cost and demand response of interruptible loads, an economic optimization scheduling model for the microgrid is established, including:
[0016] Based on the day-ahead forecast data of wind turbines, photovoltaics, AC important loads, and DC important loads, an economic optimization scheduling model for the microgrid is established according to the operation and maintenance costs of each device in the microgrid, real-time electricity prices, interruptible loads, and interruption compensation costs.
[0017] The economic optimization scheduling model established in the day-ahead scheduling stage includes: the objective function of minimizing the operating cost of the microgrid and the operating constraints of each device in the microgrid;
[0018] The operating constraints of each device in the microgrid include: operating constraints of fuel cells, operating constraints of lithium batteries, and equality or inequality constraints of each device that need to be met during the actual operation of the microgrid;
[0019] The equality or inequality constraints of each device include: AC area power balance constraint, DC area power balance constraint, converter interaction power constraint, tie line interaction power constraint and interruptible load constraint;
[0020] The interactive power of the interruptible load and the tie line is carried out according to the day-ahead dispatch plan during the operation of the microgrid;
[0021] The objective function for minimizing the operating cost of the microgrid is:
[0022]
[0023] Where F1 represents the microgrid operation cost during the day-ahead dispatch phase, C CV (P CV (t)) represents the converter operation cost of the microgrid in time period t, C grid (P grid (t)) represents the amount of electricity purchased and sold by the microgrid in time period t, C Fi (P Fi (t)) represents the operating cost of the i-th fuel cell in the microgrid during time period t, n represents the total number of fuel cells in the microgrid, and i represents the i-th fuel cell; C Lj (P Lj (t)) represents the operating cost of the jth lithium battery in the microgrid during time period t, m represents the total number of lithium batteries in the microgrid, and j represents the jth lithium battery; C 1k (P 1k (t)) represents the interruption compensation amount of the kth interruptible load in the microgrid in time period t, h represents the total number of interruptible loads in the microgrid, and k represents the kth interruptible load;
[0024] The fuel cell operating cost is:
[0025]
[0026] Where C FC represents the gas price, P Fi (t) represents the charging and discharging power of the i-th fuel cell in time period t, η FC Indicates the efficiency of the fuel cell, L HVFC Indicates the lower calorific value of the gas, K MF represents the fuel cell maintenance cost coefficient, Δt represents the operation and maintenance time;
[0027] The operating cost of the lithium battery is:
[0028]
[0029] Where C inv Indicates the initial investment cost of lithium batteries, P Lj (t) represents the charge and discharge power of the jth lithium battery in time period t, N life (t) represents the operating life of the lithium battery in time period t, E LB Indicates the rated capacity of lithium battery, K ML Indicates the maintenance cost coefficient of lithium batteries;
[0030] The lithium battery operating life is:
[0031] N life (t) = -3278D od (t) 4 -5Dod (t) 3 +12823D od (t) 2 -14122D od (t)+5112,
[0032] Where D od (t) represents the discharge depth of the lithium battery in time period t, N life (t) indicates that the depth of discharge of the lithium battery in time period t is D od Cycle life under (t);
[0033] The operating cost of the converter is:
[0034]
[0035] Where, P CV (t) represents the converter interaction power in time period t, m CV-loss represents the converter loss cost coefficient converted to the converter operating power, g CV -l oss represents the loss cost coefficient of the converter, η CV Indicates the circulating efficiency of the converter.
[0036] Furthermore, in the intraday pre-scheduling stage, based on the day-ahead forecast data of the uncontrollable units and the economic optimization scheduling model, a deep neural network is used to perform Markov decision process modeling to obtain an intraday scheduling model, including:
[0037] In the intraday pre-dispatching stage, the economic optimization dispatching model is used to obtain the economic optimization dispatching cost of the microgrid; wherein the controllable units of the microgrid in the intraday pre-dispatching stage include: lithium batteries and fuel cells;
[0038] The deep neural network in the DDPG algorithm is used to define the scheduling of controllable units as an action space, the day-ahead forecast data of uncontrollable units as a state space, and the economic optimization scheduling cost of the microgrid is defined as a benefit function. The economic scheduling strategy of the microgrid is modeled using the Markov decision process, and the intraday scheduling model is obtained through microgrid training.
[0039] The Markov decision process includes M = (S, A, P, R, λ), where S represents the state space, A represents the action space, P represents the state transition probability, R represents the profit function, and λ represents the loss factor;
[0040] In the intraday pre-dispatching stage, the whole day is divided into 96 time periods with 15 minutes as the unit period; that is, in the intraday pre-dispatching stage, the Markov decision process is completed with 15 minutes as a cycle, and it takes 96 cycles to complete the economic optimization dispatch of the microgrid in the intraday pre-dispatching stage.
[0041] Furthermore, the state space represents the input of the intraday scheduling model, and the state space is:
[0042] S=(t,P S-PV (t),P S-WT (t),P S-lac (t),P S-ldc (t),
[0043] ΔP S-PV (t),ΔP S-WT (t),ΔP S-lac (t),ΔP S-ldc (t),
[0044] S OC (t),P grid (t),P CV (t))
[0045] Where S represents the state space, t represents the current time period, and P S-PV (t) represents the simulated PV power during the pre-dispatch phase within time period t, P S-WT (t) represents the simulated wind turbine power in the pre-dispatch stage within the time period t, P S-lac (t) represents the simulated AC important load power in the pre-dispatch stage within the time period t, P S-ldc (t) represents the simulated DC important load power in the pre-dispatch stage within the time period t, ΔP S-PV (t) represents the difference between the simulated PV power in the pre-dispatch phase and the day-ahead predicted PV power in period t, ΔP S-WT (t) represents the difference between the wind turbine power simulated in the pre-dispatch phase and the wind turbine power predicted on the day before in period t; ΔP S-lac (t) represents the difference between the simulated AC important load power in the pre-dispatch stage within the time period t and the AC important load power predicted the day before, ΔP S-ldc (t) represents the difference between the simulated DC important load power during the pre-dispatch phase within the time period t and the day-ahead predicted DC important load power;
[0046] The action space represents the output of the intraday scheduling model, and the action space is:
[0047] A=(P S-F1 (t),…,P S-Fi (t),…P S-Fn (t),P S-L1 (t),…,P S-Lj (t),…P S-Lm (t)),
[0048] Where A represents the action space, P S-Fi (t) represents the charge and discharge power of the i-th fuel cell in the pre-scheduling stage within the time period t, P S-Fn (t) represents the charge and discharge power of the nth fuel cell in the pre-scheduling stage during the day in time period t, n represents the total number of fuel cells in the microgrid, and i represents the i-th fuel cell; P S-Lj (t) represents the charge and discharge power of the jth lithium battery in the pre-scheduling stage during the period t, P S-Lm (t) represents the charge and discharge power of the jth lithium battery in the pre-dispatching stage during the day in time period t, m represents the total number of lithium batteries in the microgrid, and j represents the jth lithium battery.
[0049] Furthermore, the DDPG algorithm is based on an actor-evaluator framework, and the deep neural network includes a policy network and a value network;
[0050] The policy network represents the mapping function from the current state to the action set, and the reward is maximized by optimizing the parameters and configuration of the policy network;
[0051] The value network updates the policy network by outputting the action value function Q(S,A) in the current state; where S represents the state space and A represents the action space;
[0052] The value network updates the policy network by outputting the action value function Q(S,A) in the current state, including:
[0053] In the policy network for the kth state S k According to the strategy μ k Given action A k After that, the state and action sequence is obtained as follows: (S1, A1),..., (S k ,A k ),...,(S N ,A N ),
[0054] Where S k With S N Represent the kth state and the Nth state respectively, A k and A N They represent the kth action and the Nth action respectively; k = 1, 2, ..., N, N represents the total number of microgrid economic optimization dispatch cycles in the intraday pre-dispatch stage, and the total number of states and actions is the same as the total number of dispatch cycles, that is, N = 96;
[0055] The value network outputs the action value function Q(S,A) in the current state, and obtains the state action value function Q μ (S k ,A k), and iterate through the Bellman equation to obtain:
[0056] Q μ (S k ,A k )=E[R(S k ,A k )+γQ μ (S k+1 ,μ(S k+1 ))],
[0057] Where Q μ (S k ,A k ) represents the action value function Q(S,A) in the kth state, E represents the expectation, R represents the profit function, λ represents the loss factor, λ∈[0,1]; μ represents the current strategy;
[0058] From any given S k ∈S according to μ * =arg max μ J(μ) is the update strategy of the policy network, which obtains the loss-benefit model and updates the policy network according to the loss-benefit model; where μ * represents the updated strategy;
[0059] The impairment loss model is:
[0060]
[0061] Where J(μ) represents the current loss-benefit, τ represents a trajectory of the loss-benefit model in the reinforcement learning process, τ = (S1, A1, S2, ...); N represents the total number of scheduling cycles, λ k represents the kth power of the loss factor λ, R k represents the profit function at the kth moment.
[0062] Furthermore, in the intraday scheduling stage, the intraday ultra-short-term forecast data and the day-ahead scheduling plan data of the uncontrollable units are input into the intraday scheduling model to obtain output data, and the output data is used as the scheduling basis for the controllable units, including:
[0063] During the intraday scheduling phase, the intraday ultra-short-term forecast data of the uncontrollable unit is considered to be consistent with the actual operation data. The difference between the day-ahead forecast data and the actual operation data of the uncontrollable unit is:
[0064]
[0065] Where, ΔP D-PV(t) represents the difference between the intraday ultra-short-term predicted photovoltaic power and the day-ahead predicted photovoltaic power in period t during the intraday scheduling phase, ΔP D-WT (t) represents the difference between the ultra-short-term forecast wind turbine power during the intraday scheduling phase and the day-ahead forecast wind turbine power during the t period, ΔP D-lac (t) represents the difference between the ultra-short-term forecast of the important AC load power during the intraday dispatch phase and the day-ahead forecast of the important AC load power during the t period; ΔP D-ldc (t) represents the difference between the intraday ultra-short-term forecast DC important load power during the intraday dispatch phase and the day-ahead forecast DC important load power in period t;
[0066] According to the difference between the day-ahead forecast data and the actual operation data of the uncontrollable unit, the input data of the intraday scheduling model is obtained as follows:
[0067] N D-input (t)=[t,P D-PV (t),P D-WT (t),P D-lac (t),P D-ldc (t),
[0068] ΔP D-PV (t),ΔP D-WT (t),ΔP D-lac (t),ΔP D-ldc (t),
[0069] S OC (t),P grid (t),P CV (t)]
[0070] Where N D-input represents the input data of the intraday scheduling model, P D-PV (t) represents the ultra-short-term predicted photovoltaic power within the period t, P D-WT (t) represents the ultra-short-term forecast wind turbine power within the period t, P D-ldc (t) represents the daily ultra-short-term forecast DC load power within time period t, S OC (t) represents the state of charge of the lithium battery in time period t, P grid (t) represents the tie line interaction power in time period t, P CV (t) represents the converter interaction power in time period t;
[0071] The output data obtained by inputting the input data into the intraday scheduling model is:
[0072] N D-output =(P D-F1 (t),…,P D-Fi(t),…P D-Fn (t),P D-L1 (t),…,P D-Lj (t),…P D-Lm (t)),
[0073] Where N D-output Represents the output data of the intraday scheduling model, P D-Fi (t) represents the charge and discharge power of the i-th fuel cell in the daily scheduling phase within time period t, P D-Fn (t) represents the charge and discharge power of the nth fuel cell in the daily scheduling phase within time period t, n represents the total number of fuel cells in the microgrid, and i represents the i-th fuel cell; P D-Lj (t) represents the charge and discharge power of the jth lithium battery in the intraday scheduling phase during time period t, P D-Lm (t) represents the charge and discharge power of the mth lithium battery in time period t during the intraday scheduling phase, m represents the total number of lithium batteries in the microgrid, and j represents the jth lithium battery.
[0074] Based on the same inventive concept, the present invention also provides a microgrid economic dispatch system based on deep reinforcement learning, comprising:
[0075] The first unit is used to establish and solve the economic optimization scheduling model of the microgrid based on the day-ahead forecast data of the uncontrollable units during the day-ahead scheduling phase, and obtain the day-ahead scheduling plan for each device in the microgrid;
[0076] The second unit is used to perform Markov decision process modeling using a deep neural network in the intraday pre-scheduling stage based on the day-ahead forecast data of the uncontrollable units and the economic optimization scheduling model to obtain an intraday scheduling model;
[0077] The third unit is configured to input the intraday ultra-short-term forecast data and the day-ahead scheduling plan data of the uncontrollable units into the intraday scheduling model to obtain output data during the intraday scheduling phase, and use the output data as a scheduling basis for the controllable units;
[0078] Among them, the microgrid includes uncontrollable units and controllable units.
[0079] Based on the same inventive concept, an embodiment of the present invention also provides an electronic device, including: a memory and a processor; the processor is used to read and execute the computer program stored in the memory to implement the aforementioned microgrid economic dispatch method based on deep reinforcement learning.
[0080] Based on the same inventive concept, an embodiment of the present invention further provides a computer storage medium, wherein the computer storage medium stores computer executable instructions, and when the computer executable instructions are executed, the aforementioned microgrid economic dispatch method based on deep reinforcement learning is implemented.
[0081] The technical effects and advantages of the present invention are as follows: 1. In response to the uncertainty problem caused by the volatility of renewable energy power generation such as wind and solar power, the present invention combines multi-time-scale optimization strategies with artificial intelligence to design a microgrid economic dispatch strategy based on a deep reinforcement learning algorithm. This method first considers the wind and solar power forecast data and the total operating cost of the microgrid in the day-ahead phase to establish a microgrid economic optimization dispatch model. Then, based on deep reinforcement learning, the microgrid optimization dispatch process is modeled as a Markov decision process in the intraday dispatch phase, and the prediction errors of wind, solar, and important loads are considered in the state space to cope with the uncertainty of renewable energy and achieve the economic optimal dispatch of the microgrid.
[0082] 2. Among the current mainstream algorithms, most rely on modeling of components in the microgrid, and the accuracy of the modeling directly affects the accuracy of the entire algorithm's operating results. The present invention combines deep reinforcement learning with neural networks and applies it to microgrid economic scheduling, which can not only effectively solve the problem of dependence on the model, but also greatly improve the efficiency of the solution.
[0083] Other features and advantages of the present invention will be described in the following description, and in part will become apparent from the description, or will be understood by practicing the present invention. The purpose and other advantages of the present invention can be realized and obtained by the structures pointed out in the description, claims and drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0084] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0085] Figure 1 This is a flow chart of a microgrid economic dispatch method based on deep reinforcement learning according to an embodiment of the present invention;
[0086] Figure 2 Schematic diagram of the microgrid structure in an embodiment of the present invention;
[0087] Figure 3 Schematic diagram of wind power forecast data in the day-ahead scheduling phase according to an embodiment of the present invention;
[0088] Figure 4Schematic diagram of photovoltaic prediction data in the day-ahead scheduling phase according to an embodiment of the present invention;
[0089] Figure 5 This is a schematic diagram of important load forecast data in the day-ahead scheduling phase in an embodiment of the present invention;
[0090] Figure 6 A schematic diagram of a technical route of a microgrid economic dispatch method according to an embodiment of the present invention;
[0091] Figure 7 This is a schematic diagram of the DDPG algorithm principle in an embodiment of the present invention;
[0092] Figure 8 Schematic diagram showing the comparison of the lithium battery before and after scheduling during the intraday scheduling phase in an embodiment of the present invention;
[0093] Figure 9 Schematic diagram showing a comparison of a fuel cell before and after scheduling during the intraday scheduling phase in an embodiment of the present invention;
[0094] Figure 10 Schematic diagram showing the comparison of the state of charge of a lithium battery before and after scheduling during the intraday scheduling phase in an embodiment of the present invention;
[0095] Figure 11 This is a schematic diagram of the structure of a microgrid economic dispatch system based on deep reinforcement learning according to an embodiment of the present invention;
[0096] Figure 12 The figure is a schematic structural diagram of an electronic device according to an embodiment of the present invention. DETAILED DESCRIPTION
[0097] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0098] In order to solve the shortcomings of the existing technology, the present invention discloses a microgrid economic dispatch method based on deep reinforcement learning, such as Figure 1 As shown, the following steps are included:
[0099] Step S1: In the day-ahead scheduling phase, based on the day-ahead forecast data of the uncontrollable units, an economic optimization scheduling model of the microgrid is established and solved to obtain the day-ahead scheduling plan for each device in the microgrid;
[0100] Step S2: In the intraday pre-scheduling stage, based on the day-ahead forecast data of the uncontrollable units and the economic optimization scheduling model, a deep neural network is used to perform Markov decision process modeling to obtain an intraday scheduling model;
[0101] Step S3: In the intraday scheduling stage, the intraday ultra-short-term forecast data and the day-ahead scheduling plan data of the uncontrollable units are input into the intraday scheduling model to obtain output data, and the output data is used as the scheduling basis for the controllable units;
[0102] Among them, the microgrid includes uncontrollable units and controllable units.
[0103] In some specific embodiments, step S1: in the day-ahead scheduling phase, based on the day-ahead forecast data of the uncontrollable units, an economic optimization scheduling model of the microgrid is established and solved to obtain a day-ahead scheduling plan for each device in the microgrid, including:
[0104] Step S101: Based on the day-ahead forecast data of the uncontrollable units, an economic dispatch model of the microgrid is established taking into account the operating cost of the microgrid and the demand response of the interruptible load.
[0105] Step S102: Based on the economic dispatch model of the microgrid, a mixed integer programming method is used to solve the problem, and a dispatch plan for each device in the microgrid during the day-ahead dispatch phase is obtained, namely, the day-ahead dispatch plan.
[0106] Among them, Figure 2 As shown, the microgrid is an AC / DC hybrid microgrid. In the day-ahead scheduling phase, the uncontrollable units include wind turbines, photovoltaics, and important loads, and the important loads include AC important loads and DC important loads. The controllable units include lithium batteries, fuel cells, and interruptible loads. The operating parameters of each device in the microgrid are shown in Table 1:
[0107] Table 1 Operating parameters of each device in the AC / DC hybrid microgrid
[0108]
[0109]
[0110] like Figure 3 、 Figure 4 and Figure 5 As shown, the day-ahead forecast data of the uncontrollable units include: the day-ahead forecast data of wind turbines, photovoltaics, AC important loads and DC important loads.
[0111] In some specific embodiments, step S101: establishing an economic dispatch model for a microgrid based on day-ahead forecast data of uncontrollable units, taking into account the microgrid operating costs and demand response of interruptible loads, includes:
[0112] Based on the day-ahead forecast data of wind turbines, photovoltaics, AC important loads, and DC important loads, an economic optimization scheduling model for the microgrid is established according to the operation and maintenance costs of each device in the microgrid, real-time electricity prices, interruptible loads, and interruption compensation costs.
[0113] In the day-ahead dispatching stage, the economic optimization dispatching model established includes: the objective function of minimizing the operating cost of the microgrid and the operating constraints of each device in the microgrid;
[0114] The objective function for minimizing the operating cost of the microgrid takes into account variables such as the life cycle of the lithium battery, the operating cost of the lithium battery (including maintenance cost), the operating cost of the fuel cell (including maintenance cost), the converter cost, the converter power loss, the converter power purchase and sales cost, and the interruptible load compensation cost. Therefore, the objective function for minimizing the operating cost of the microgrid is:
[0115]
[0116] Where F1 represents the microgrid operation cost during the day-ahead dispatch phase, C CV (P CV (t)) represents the converter operation cost of the microgrid in time period t, C grid (P grid (t)) represents the amount of electricity purchased and sold by the microgrid in time period t, C Fi (P Fi (t)) represents the operating cost of the i-th fuel cell in the microgrid during time period t, n represents the total number of fuel cells in the microgrid, and i represents the i-th fuel cell; C Lj (P Lj (t)) represents the operating cost of the jth lithium battery in the microgrid during time period t, m represents the total number of lithium batteries in the microgrid, and j represents the jth lithium battery; C 1k (P 1k (t)) represents the interruption compensation amount of the kth interruptible load in the microgrid in time period t, h represents the total number of interruptible loads in the microgrid, and k represents the kth interruptible load;
[0117] The fuel cell operating cost is:
[0118]
[0119] Where C FC represents the gas price, P Fi (t) represents the charge and discharge power of the i-th fuel cell in time period t, η FC Indicates the efficiency of the fuel cell, L HVFC Indicates the lower calorific value of the gas, K MFrepresents the fuel cell maintenance cost coefficient, Δt represents the operation and maintenance time;
[0120] The operating cost of the lithium battery is:
[0121]
[0122] Where C inv Indicates the initial investment cost of lithium batteries, P Lj (t) represents the charging and discharging power of the lithium battery in time period t, N life (t) represents the operating life of the lithium battery in time period t, E LB Indicates the rated capacity of lithium battery, K ML Indicates the maintenance cost coefficient of lithium batteries;
[0123] The lithium battery operating life is:
[0124] N life (t) = -3278D od (t) 4 -5D od (t) 3 +12823D od (t) 2 -14122D od (t)+5112,
[0125] Where D od (t) represents the discharge depth of the lithium battery in time period t, N life (t) indicates that the depth of discharge of the lithium battery in time period t is D od Cycle life under (t);
[0126] The operating cost of the converter is:
[0127]
[0128] Where, P CV (t) represents the converter interaction power in time period t, m CV-loss represents the converter loss cost coefficient converted to the converter operating power, g CV-loss represents the loss cost coefficient of the converter, η CV Indicates the circulating efficiency of the converter.
[0129] The operating constraints of each device in the microgrid include: the operating constraints of the fuel cell, the operating constraints of the lithium battery, and the equality or inequality constraints of each device that need to be met during the actual operation of the microgrid;
[0130] The equality or inequality constraints of each device include: AC area power balance constraint, DC area power balance constraint, converter interaction power constraint, tie line interaction power constraint and interruptible load constraint;
[0131] In order to reduce the impact on the distribution network, the load and tie-line interactive power can be interrupted during the operation of the microgrid according to the day-ahead dispatch plan.
[0132] The fuel cell operation constraints are:
[0133]
[0134] Where, P Fi (t) represents the output power of the fuel cell in time period t, Indicates the lower limit of the fuel cell's output power, Indicates the upper limit of the fuel cell's output power;
[0135] The operating constraints of lithium batteries are:
[0136]
[0137] Where S OC (t) represents the state of charge of the lithium battery in time period t, E LB (t) represents the remaining capacity of the lithium battery in time period t, E LB Indicates the rated capacity of the lithium battery. Indicates the minimum charge capacity of the lithium battery. Indicates the maximum charge capacity of the lithium battery, P Lj (t) represents the output power of the lithium battery in time period t, γ represents the charging and discharging efficiency of the lithium battery, E LB (0) represents the initial capacity of the lithium battery, E LB (24) represents the final capacity of the lithium battery, Indicates the lower limit of lithium battery output power. Indicates the upper limit of lithium battery output power;
[0138] The AC area power balance constraint is:
[0139] P WT (t)+P grid (t)+P CV (t) = P lac (t),
[0140] Where, P WT (t) represents the power generated by the wind turbine in the period t predicted by the day before, P grid (t) represents the tie line power, P CV (t) represents the converter interaction power in time period t, Plac (t) represents the power of the AC important load predicted in time period t;
[0141] The power balance constraint in the DC region is:
[0142]
[0143] Where, P PV (t) represents the power generated by photovoltaic power in time period t, P Fi (t) represents the fuel cell power in time period t, n represents the total number of fuel cells in the microgrid, and i represents the i-th fuel cell; P CV (t) represents the converter interaction power in time period t, P Lj (t) represents the output power of the lithium battery in time period t, m represents the total number of lithium batteries in the microgrid, and j represents the jth lithium battery; P ldc (t) represents the power of the DC important load predicted in time period t, P lk (t) represents the power of the interruptible load predicted in time period t, h represents the total number of interruptible loads in the microgrid, and k represents the kth interruptible load;
[0144] The converter interaction power constraint is:
[0145]
[0146] Where, P CV represents the converter interaction power, and Respectively represent the upper and lower limits of the converter interaction power;
[0147] The tie line interaction power constraint is:
[0148]
[0149] Where, P grid represents the tie line interaction power, and Respectively represent the upper and lower limits of the tie line interaction power;
[0150] The interruptible load constraint is:
[0151]
[0152] Where, I 1k (t) represents the operating status of the kth interruptible load in time period t, T 1k represents the maximum interruptible duration of the kth interruptible load in a day, h represents the total number of interruptible loads in the microgrid, and k represents the kth interruptible load.
[0153] In some specific embodiments, step S2: in the intraday pre-scheduling stage, based on the day-ahead forecast data of the uncontrollable units and the economic optimization scheduling model, a deep neural network is used to model the Markov decision process, and the environment (microgrid) and the intelligent agent are continuously interacted to train and obtain the intraday scheduling model, including:
[0154] In the intraday pre-dispatching stage, the economic optimization dispatching model is used to obtain the economic optimization dispatching cost of the microgrid; wherein the controllable units of the microgrid in the intraday pre-dispatching stage include: lithium batteries and fuel cells;
[0155] The deep neural network in the DDPG algorithm is used to define the scheduling of controllable units as action space, the day-ahead forecast data of uncontrollable units as state space, and the economic optimization scheduling cost of the microgrid as the benefit function. The Markov decision process is used to model the economic scheduling strategy of the microgrid, and the intraday scheduling model is obtained through microgrid training.
[0156] The Markov decision process includes M = (S, A, P, R, λ), where S represents the state space, A represents the action space, P represents the state transition probability, R represents the profit function, and λ represents the loss factor;
[0157] In the intraday pre-dispatching stage, the whole day is divided into 96 time periods with 15 minutes as the unit period; that is, in the intraday pre-dispatching stage, the Markov decision process is completed with 15 minutes as a cycle, and it takes 96 cycles to complete the economic optimization dispatch of the microgrid in the intraday pre-dispatching stage.
[0158] The state space represents the input of the intraday scheduling model, and the state space is:
[0159] S=(t,P S-PV (t),P S-WT (t),P S-lac (t),P S-ldc (t),
[0160] ΔP S-PV (t),ΔP S-WT (t),ΔP S-lac (t),ΔP S-ldc (t),
[0161] S OC (t),P grid (t),P CV (t))
[0162] Where S represents the state space, t represents the current time period, and P S-PV (t) represents the simulated PV power during the pre-dispatch phase within time period t, P S-WT(t) represents the simulated wind turbine power in the pre-dispatch stage within the time period t, P S-lac (t) represents the simulated AC important load power in the pre-dispatch stage within the time period t, P S-ldc (t) represents the simulated DC important load power in the pre-dispatch stage within the time period t, ΔP S-PV (t) represents the difference between the simulated PV power in the pre-dispatch phase and the day-ahead predicted PV power in period t, ΔP S-WT (t) represents the difference between the wind turbine power simulated in the pre-dispatch phase and the wind turbine power predicted on the day before in period t; ΔP S-lac (t) represents the difference between the simulated AC important load power in the pre-dispatch stage within the time period t and the AC important load power predicted the day before, ΔP S-ldc (t) represents the difference between the simulated DC important load power in the intraday pre-dispatch stage and the day-ahead predicted DC important load power in time period t.
[0163] The action space represents the output of the intraday dispatch model. Its physical meaning is the set of controlled variables of the agent during decision-making and control, corresponding to the controllable units of the microgrid economic dispatch. The action space is:
[0164] A=(P S-F1 (t),…,P S-Fi (t),…P S-Fn (t),P S-L1 (t),…,P S-Lj (t),…P S-Lm (t)),
[0165] Where A represents the action space, P S-Fi (t) represents the charge and discharge power of the i-th fuel cell in the pre-scheduling stage within the time period t, P S-Fn (t) represents the charge and discharge power of the nth fuel cell in the pre-scheduling stage during the day in time period t, n represents the total number of fuel cells in the microgrid, and i represents the i-th fuel cell; P S-Lj (t) represents the charge and discharge power of the jth lithium battery in the pre-scheduling stage during the period t, P S-Lm (t) represents the charge and discharge power of the jth lithium battery in the pre-dispatching stage during the day in time period t, m represents the total number of lithium batteries in the microgrid, and j represents the jth lithium battery.
[0166] Since the optimization goal of the microgrid in the present invention is to consider the operation and maintenance costs, electricity purchase and sales costs, and interruptible load compensation costs of each device in the microgrid, the main body of the profit function is the economic optimization scheduling cost of the microgrid; after obtaining the economic optimization scheduling cost of the microgrid according to the economic optimization scheduling model formula, the opposite of the economic optimization scheduling cost of the microgrid is used as the profit value.
[0167] The state transition probability is unknown, and the present invention uses a model-free DDPG algorithm to solve it.
[0168] The DDPG algorithm is based on the actor-evaluator framework, which enables the agent and the environment to continuously interact to solve the model.
[0169] The deep neural network includes a policy network and a value network. The policy network represents the mapping function from the current state to the action set, and the reward is maximized by optimizing the parameters and configuration of the policy network. The update of the weight parameters of the policy network relies on the value network. The value network outputs the action value function Q(S,A) under the current state, and then updates the policy network according to gradient descent. Where S represents the state space and A represents the action space.
[0170] Among them, Figure 6 As shown, the DDPG algorithm uses the critic neural network Q(S,A|θ Q ) to approximate the action-value function Q(S,A), and use the actor neural network μ(S|θ μ ) to approximate the strategy μ; where θ Q and θ μ Represents the current Actor and current Critic network parameters respectively;
[0171] In addition, the DDPG algorithm also adds a target network with the same structure as the corresponding network, namely the target actor network μ'(S|θ μ' ) and the target critic network Q'(S,A|θ Q' ), the setting of the target network can effectively improve the convergence of the algorithm; where θ Q' and θ μ' Represents the network parameters of Target Actor and Target Critic.
[0172] The target network is updated using soft update, which is as follows:
[0173]
[0174] Where θ Q',l+1 represents the updated target critic network parameters, θ Q,l represents the current critic network parameters before update, θ Q',l represents the target critic network parameters before update, θ μ',l+1 represents the updated target actor network parameters, θ μ,l represents the current actor network parameters before update, θ μ',lRepresents the target actor network parameters before update, l represents the current iteration number, τ' represents the learning rate of the target network parameters, τ'<<1.
[0175] For the Critic network, its purpose is to accurately evaluate the state-action value function Q(S,A), so its loss function can be expressed as:
[0176]
[0177] Where L represents the loss function, α represents the index of the currently selected Batch-size sample, Z represents the number of steps to calculate the cumulative reward, and δ α represents the timing difference error, y α Represents the state-action value function value estimated by the Target Critic network, Q(S α ,A α |θ Q ,) represents the state-action value function value of the αth sample calculated by the current evaluator network;
[0178] Among them, y α =R α +λQ'(S α+1 ,μ'(S α+1 |θ μ' )|θ Q' ),
[0179] Where R α represents the immediate reward corresponding to the selected αth sample, λ represents the loss factor, and Table Q'(S α+1 ,μ'(S α+1 |θ μ' )|θ Q' ) represents the state-action value function value of the α+1th sample calculated by the target evaluator network.
[0180] In some specific embodiments, the value network outputs the action-value function Q(S, A) in the current state, and then updates the policy network according to gradient descent, including:
[0181] Since it takes 96 cycles to complete the economic optimization dispatch of the microgrid in the pre-dispatching stage within the day, the strategy network is used for the kth state S k According to the strategy μ k Given action A k After that, the state and action sequence is obtained as follows: (S1, A1),..., (S k ,A k ),...,(S N ,A N ),
[0182] Where Sk With S N Represent the kth state and the Nth state respectively, A k and A N They represent the kth action and the Nth action respectively, k = 1, 2, ..., N, N represents the total number of microgrid economic optimization dispatch cycles in the intraday pre-dispatch stage, and the total number of states and actions is the same as the total number of dispatch cycles, that is, N = 96.
[0183] The value network outputs the action value function Q(S,A) in the current state, and obtains the state action value function Q μ (S k ,A k ), and iterate through the Bellman equation to obtain:
[0184] Q μ (S k ,A k )=E[R(S k ,A k )+γQ μ (S k+1 ,μ(S k+1 ))],
[0185] Where Q μ (S k ,A k ) represents the action value function Q(S,A) in the kth state, E represents the expectation, R represents the profit function, λ represents the loss factor, λ∈[0,1]; μ represents the current strategy;
[0186] From any given S k ∈S according to μ * =arg max μ J(μ) is the update strategy of the policy network, which obtains the loss-benefit model and updates the policy network according to the loss-benefit model; where μ * Indicates the updated policy.
[0187] The impairment loss model is:
[0188]
[0189] Where J(μ) represents the current loss-benefit, τ represents a trajectory of the loss-benefit model in the reinforcement learning process, τ = (S1, A1, S2, ...); N represents the total number of scheduling cycles, λ k R represents the k-th power of the impairment factor λ, which is an important parameter for measuring current and future returns. λ∈[0,1], when λ approaches 1, it means that future returns are more important, and when λ approaches 0, it means that current returns are more important. krepresents the profit function at the kth moment.
[0190] In some specific embodiments, such as Figure 7 As shown, step S3: in the intraday scheduling stage, the intraday ultra-short-term forecast data of the uncontrollable unit and the day-ahead scheduling plan data of the uncontrollable unit are input into the intraday scheduling model to obtain output data, and the output data is used as the scheduling basis of the controllable unit, including:
[0191] During the intraday scheduling phase, the intraday ultra-short-term forecast data of the uncontrollable unit is considered to be consistent with the actual operation data. The difference between the day-ahead forecast data and the actual operation data of the uncontrollable unit is:
[0192]
[0193] Where ΔP D-PV (t) represents the difference between the intraday ultra-short-term predicted photovoltaic power and the day-ahead predicted photovoltaic power in period t during the intraday scheduling phase, ΔP D-WT (t) represents the difference between the ultra-short-term forecast wind turbine power during the intraday scheduling phase and the day-ahead forecast wind turbine power during the t period, ΔP D-lac (t) represents the difference between the ultra-short-term forecast of the important AC load power during the intraday dispatch phase and the day-ahead forecast of the important AC load power during the t period; ΔP D-ldc (t) represents the difference between the intraday ultra-short-term forecast DC important load power during the intraday dispatch phase and the day-ahead forecast DC important load power in period t;
[0194] According to the difference between the day-ahead forecast data and the actual operation data of the uncontrollable unit, the input data of the intraday scheduling model is obtained as follows:
[0195] N D-input (t)=[t,P D-PV (t),P D-WT (t),P D-lac (t),P D-ldc (t),
[0196] ΔP D-PV (t),ΔP D-WT (t),ΔP D-lac (t),ΔP D-ldc (t),
[0197] S OC (t),P grid (t),P CV (t)]
[0198] Where N D-input represents the input data of the intraday scheduling model, P D-PV(t) represents the ultra-short-term predicted photovoltaic power within the period t, P D-WT (t) represents the ultra-short-term forecast wind turbine power within the period t, P D-ldc (t) represents the daily ultra-short-term forecast DC load power within time period t, S OC (t) represents the state of charge of the lithium battery in time period t, P grid (t) represents the tie line interaction power in time period t, P CV (t) represents the converter interaction power in time period t;
[0199] The input data is input into the output data (power of lithium battery and fuel cell) obtained in the intraday scheduling model; the output data is:
[0200] N D-output =(P D-F1 (t),…,P D-Fi (t),…P D-Fn (t),P D-L1 (t),…,P D-Lj (t),…P D-Lm (t)),
[0201] Where N D-output Represents the output data of the intraday scheduling model, P D-Fi (t) represents the charge and discharge power of the i-th fuel cell in the daily scheduling phase within time period t, P D-Fn (t) represents the charge and discharge power of the nth fuel cell in the daily scheduling phase within time period t, n represents the total number of fuel cells in the microgrid, and i represents the i-th fuel cell; P D-Lj (t) represents the charge and discharge power of the jth lithium battery in the intraday scheduling phase during time period t, P D-Lm (t) represents the charge and discharge power of the mth lithium battery in time period t during the intraday scheduling phase, m represents the total number of lithium batteries in the microgrid, and j represents the jth lithium battery.
[0202] like Figure 8 、 Figure 9 and Figure 10 As shown, there are comparison diagrams of lithium batteries, fuel cells and the state of charge of lithium batteries before and after scheduling through the intraday scheduling model.
[0203] Based on the same inventive concept, the present invention also discloses a microgrid economic dispatching system based on deep reinforcement learning, such as Figure 11 Shown, including:
[0204] The first unit is used to establish and solve the economic optimization scheduling model of the microgrid based on the day-ahead forecast data of the uncontrollable units during the day-ahead scheduling phase, and obtain the day-ahead scheduling plan for each device in the microgrid;
[0205] The second unit is used to perform Markov decision process modeling using a deep neural network in the intraday pre-scheduling stage based on the day-ahead forecast data of the uncontrollable units and the economic optimization scheduling model to obtain an intraday scheduling model;
[0206] The third unit is configured to input the intraday ultra-short-term forecast data and the day-ahead scheduling plan data of the uncontrollable units into the intraday scheduling model to obtain output data during the intraday scheduling phase, and use the output data as a scheduling basis for the controllable units;
[0207] Among them, the microgrid includes uncontrollable units and controllable units.
[0208] Regarding the system in the above embodiment, the specific manner in which each unit module performs operations has been described in detail in the embodiment of the method, and will not be elaborated here.
[0209] Based on the same inventive concept, an embodiment of the present invention further provides an electronic device, the structure of which is as follows: Figure 12 As shown, it includes: a memory and a processor, and the processor is used to read and execute the computer program stored in the memory to implement the aforementioned microgrid economic dispatch method based on deep reinforcement learning.
[0210] Based on the same inventive concept, an embodiment of the present invention further provides a computer storage medium, wherein the computer storage medium stores computer executable instructions, and when the computer executable instructions are executed, the aforementioned microgrid economic dispatch method based on deep reinforcement learning is implemented.
[0211] Finally, it should be noted that the above is only a preferred embodiment of the present invention and is not intended to limit the present invention. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art can still modify the technical solutions described in the aforementioned embodiments or make equivalent substitutions for some of the technical features therein. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.
Claims
1. A microgrid economic dispatch method based on deep reinforcement learning, characterized in that: include: In the day-ahead dispatch phase, based on the day-ahead forecast data of the uncontrollable units, an economic optimization dispatch model for the microgrid is established and solved to obtain the day-ahead dispatch plan for each device in the microgrid. In the intraday pre-scheduling stage, based on the day-ahead forecast data of the uncontrollable units and the economic optimization scheduling model, a deep neural network is used to perform Markov decision process modeling to obtain an intraday scheduling model; In the intraday scheduling stage, the intraday ultra-short-term forecast data and the day-ahead scheduling plan data of the uncontrollable units are input into the intraday scheduling model to obtain output data, and the output data is used as the scheduling basis for the controllable units; Among them, the microgrid includes uncontrollable units and controllable units.
2. A microgrid economic dispatch method based on deep reinforcement learning according to claim 1, characterized in that: In the day-ahead dispatch phase, based on the day-ahead forecast data of the uncontrollable units, an economic optimization dispatch model for the microgrid is established and solved to obtain the day-ahead dispatch plan, including: Based on the day-ahead forecast data of uncontrollable units, the economic dispatch model of the microgrid is established by considering the microgrid operation cost and the demand response of interruptible loads. Based on the economic dispatch model of the microgrid, a mixed integer programming method is used to solve the scheduling plan of each device in the microgrid during the day-ahead scheduling phase, namely the day-ahead scheduling plan; Wherein, the microgrid is an AC / DC hybrid microgrid; In the day-ahead scheduling phase, the uncontrollable units include wind turbines, photovoltaics and important loads, and the important loads include AC important loads and DC important loads; the controllable units include lithium batteries, fuel cells and interruptible loads.
3. A microgrid economic dispatch method based on deep reinforcement learning according to claim 2, characterized in that: Based on the day-ahead forecast data of uncontrollable units, and taking into account the microgrid operation cost and demand response of interruptible loads, an economic optimization dispatch model for the microgrid is established, including: Based on the day-ahead forecast data of wind turbines, photovoltaics, AC important loads, and DC important loads, an economic optimization scheduling model for the microgrid is established according to the operation and maintenance costs of each device in the microgrid, real-time electricity prices, interruptible loads, and interruption compensation costs. The economic optimization scheduling model established in the day-ahead scheduling stage includes: the objective function of minimizing the operating cost of the microgrid and the operating constraints of each device in the microgrid; The operating constraints of each device in the microgrid include: operating constraints of fuel cells, operating constraints of lithium batteries, and equality or inequality constraints of each device that need to be met during the actual operation of the microgrid; The equality or inequality constraints of each device include: AC area power balance constraint, DC area power balance constraint, converter interaction power constraint, tie line interaction power constraint and interruptible load constraint; The interactive power of the interruptible load and the tie line is carried out according to the day-ahead dispatch plan during the operation of the microgrid; The objective function for minimizing the operating cost of the microgrid is: Where F1 represents the microgrid operation cost during the day-ahead dispatch phase, C CV (P CV (t)) represents the converter operation cost of the microgrid in time period t, C grid (P grid (t)) represents the amount of electricity purchased and sold by the microgrid in time period t, C Fi (P Fi (t)) represents the operating cost of the i-th fuel cell in the microgrid during time period t, n represents the total number of fuel cells in the microgrid, and i represents the i-th fuel cell; C Lj (P Lj (t)) represents the operating cost of the jth lithium battery in the microgrid during time period t, m represents the total number of lithium batteries in the microgrid, and j represents the jth lithium battery; C 1k (P 1k (t)) represents the interruption compensation amount of the kth interruptible load in the microgrid in time period t, h represents the total number of interruptible loads in the microgrid, and k represents the kth interruptible load; The fuel cell operating cost is: Where C FC represents the gas price, P Fi (t) represents the charging and discharging power of the i-th fuel cell in time period t, η FC Indicates the efficiency of the fuel cell, L HVFC Indicates the lower calorific value of the gas, K MF represents the fuel cell maintenance cost coefficient, Δt represents the operation and maintenance time; The operating cost of the lithium battery is: Where C inv Indicates the initial investment cost of lithium batteries, P Lj (t) represents the charge and discharge power of the jth lithium battery in time period t, N life (t) represents the operating life of the lithium battery in time period t, E LB Indicates the rated capacity of lithium battery, K ML Indicates the maintenance cost coefficient of lithium batteries; The lithium battery operating life is: N life (t)=-3278D od (t) 4 -5D od (t) 3 +12823D od (t) 2 -14122D od (t)+5112, Where D od (t) represents the discharge depth of the lithium battery in time period t, N life (t) indicates that the depth of discharge of the lithium battery in time period t is D od Cycle life under (t); The operating cost of the converter is: Where, P CV (t) represents the converter interaction power in time period t, m CV-loss represents the converter loss cost coefficient converted to the converter operating power, g CV-loss represents the loss cost coefficient of the converter, η CV Indicates the circulating efficiency of the converter.
4. A microgrid economic dispatch method based on deep reinforcement learning according to claim 2, characterized in that: In the intraday pre-scheduling stage, based on the day-ahead forecast data of the uncontrollable units and the economic optimization scheduling model, a deep neural network is used to perform Markov decision process modeling to obtain an intraday scheduling model, including: In the intraday pre-dispatching stage, the economic optimization dispatching model is used to obtain the economic optimization dispatching cost of the microgrid; wherein the controllable units of the microgrid in the intraday pre-dispatching stage include: lithium batteries and fuel cells; The deep neural network in the DDPG algorithm is used to define the scheduling of controllable units as an action space, the day-ahead forecast data of uncontrollable units as a state space, and the economic optimization scheduling cost of the microgrid is defined as a benefit function. The economic scheduling strategy of the microgrid is modeled using the Markov decision process, and the intraday scheduling model is obtained through microgrid training. The Markov decision process includes M = (S, A, P, R, λ), where S represents the state space, A represents the action space, P represents the state transition probability, R represents the profit function, and λ represents the loss factor; In the intraday pre-dispatching stage, the whole day is divided into 96 time periods with 15 minutes as the unit period; that is, in the intraday pre-dispatching stage, the Markov decision process is completed with 15 minutes as a cycle, and it takes 96 cycles to complete the economic optimization dispatch of the microgrid in the intraday pre-dispatching stage.
5. A microgrid economic dispatch method based on deep reinforcement learning according to claim 4, characterized in that: The state space represents the input of the intraday scheduling model, and the state space is: S=(t,P S-PV (t),P S-WT (t),P S-lac (t),P S-ldc (t), ΔP S-PV (t),ΔP S-WT (t),ΔP S-lac (t),ΔP S-ldc (t), S OC (t),P grid (t),P CV (t)) Where S represents the state space, t represents the current time period, and P S-PV (t) represents the simulated PV power during the pre-dispatch phase within time period t, P S-WT (t) represents the simulated wind turbine power in the pre-dispatch stage within the time period t, P S-lac (t) represents the simulated AC important load power in the pre-dispatch stage within the time period t, P S-ldc (t) represents the simulated DC important load power in the pre-dispatch stage within the time period t, ΔP S-PV (t) represents the difference between the simulated PV power in the pre-dispatch phase and the day-ahead predicted PV power in period t, ΔP S-WT (t) represents the difference between the wind turbine power simulated in the pre-dispatch phase and the wind turbine power predicted on the day before in period t; ΔP S-lac (t) represents the difference between the simulated AC important load power in the pre-dispatch stage within the time period t and the AC important load power predicted the day before, ΔP S-ldc (t) represents the difference between the simulated DC important load power during the pre-dispatch phase within the time period t and the day-ahead predicted DC important load power; The action space represents the output of the intraday scheduling model, and the action space is: A=(P S-F1 (t),…,P S-Fi (t),…P S-Fn (t),P S-L1 (t),…,P S-Lj (t),…P S-Lm (t)), Where A represents the action space, P S-Fi (t) represents the charge and discharge power of the i-th fuel cell in the pre-scheduling stage within the time period t, P S-Fn (t) represents the charge and discharge power of the nth fuel cell in the pre-scheduling stage during the day in time period t, n represents the total number of fuel cells in the microgrid, and i represents the i-th fuel cell; P S-Lj (t) represents the charge and discharge power of the jth lithium battery in the pre-scheduling stage during the period t, P S-Lm (t) represents the charge and discharge power of the jth lithium battery in the pre-dispatching stage during the day in time period t, m represents the total number of lithium batteries in the microgrid, and j represents the jth lithium battery.
6. A microgrid economic dispatch method based on deep reinforcement learning according to claim 4 or 5, characterized in that: The DDPG algorithm is based on the actor-evaluator framework, and the deep neural network includes a policy network and a value network; The policy network represents the mapping function from the current state to the action set, and the reward is maximized by optimizing the parameters and configuration of the policy network; The value network updates the policy network by outputting the action value function Q(S,A) in the current state; where S represents the state space and A represents the action space; The value network updates the policy network by outputting the action value function Q(S,A) in the current state, including: In the policy network for the kth state S k According to the strategy μ k Given action A k After that, the state and action sequence is obtained as follows: (S1, A1),..., (S k ,A k ),...,(S N ,A N ), Where S k With S N Represent the kth state and the Nth state respectively, A k and A N They represent the kth action and the Nth action respectively; k = 1, 2, ..., N, N represents the total number of microgrid economic optimization dispatch cycles in the intraday pre-dispatch stage, and the total number of states and actions is the same as the total number of dispatch cycles, that is, N = 96; The value network outputs the action value function Q(S,A) in the current state, and obtains the state action value function Q μ (S k ,A k ), and iterate through the Bellman equation to obtain: Q μ (S k ,A k )=E[R(S k ,A k )+γQ μ (S k+1 ,μ(S k+1 ))], Where Q μ (S k ,A k ) represents the action value function Q(S,A) in the kth state, E represents the expectation, R represents the profit function, λ represents the loss factor, λ∈[0,1]; μ represents the current strategy; From any given S k ∈S according to μ * =arg max μ J(μ) is the update strategy of the policy network, which obtains the loss-benefit model and updates the policy network according to the loss-benefit model; where μ * represents the updated strategy; The impairment loss model is: Where J(μ) represents the current loss-benefit, τ represents a trajectory of the loss-benefit model in the reinforcement learning process, τ = (S1, A1, S2, ...); N represents the total number of scheduling cycles, λ k represents the kth power of the loss factor λ, R k represents the profit function at the kth moment.
7. A microgrid economic dispatch method based on deep reinforcement learning according to claim 1, characterized in that: In the intraday scheduling phase, the intraday ultra-short-term forecast data and the day-ahead scheduling plan data of the uncontrollable units are input into the intraday scheduling model to obtain output data, and the output data is used as the scheduling basis for the controllable units, including: During the intraday scheduling phase, the intraday ultra-short-term forecast data of the uncontrollable unit is considered to be consistent with the actual operation data. The difference between the day-ahead forecast data and the actual operation data of the uncontrollable unit is: Where, ΔP D-PV (t) represents the difference between the intraday ultra-short-term predicted photovoltaic power and the day-ahead predicted photovoltaic power in period t during the intraday scheduling phase, ΔP D-WT (t) represents the difference between the ultra-short-term forecast wind turbine power during the intraday scheduling phase and the day-ahead forecast wind turbine power during the t period, ΔP D-lac (t) represents the difference between the ultra-short-term forecast of the important AC load power during the intraday dispatch phase and the day-ahead forecast of the important AC load power during the t period; ΔP D-ldc (y) represents the difference between the intraday ultra-short-term forecast DC important load power during the intraday dispatch phase and the day-ahead forecast DC important load power in period t; According to the difference between the day-ahead forecast data and the actual operation data of the uncontrollable unit, the input data of the intraday scheduling model is obtained as follows: N D-input (t)=[t,P D-PV (t),P D-WT (t),P D-lac (t),P D-ldc (t), ΔP D-PV (t),ΔP D-WT (t),ΔP D-lac (t),ΔP D-ldc (t), S OC (t),P grid (t),P CV (t)] Where N D-input represents the input data of the intraday scheduling model, P D-PV (t) represents the ultra-short-term predicted photovoltaic power within the period t, P D-WT (t) represents the ultra-short-term forecast wind turbine power within the period t, P D-ldc (t) represents the daily ultra-short-term forecast DC load power within time period t, S OC (t) represents the state of charge of the lithium battery in time period t, P grid (t) represents the tie line interaction power in time period t, P CV (t) represents the converter interaction power in time period t; The output data obtained by inputting the input data into the intraday scheduling model is: N D-output =(P D-F1 (t),…,P D-Fi (t),…P D-Fn (t),P D-L1 (t),…,P D-Lj (t),…P D-Lm (t)), Where N D-outDut Represents the output data of the intraday scheduling model, P D-Fi (t) represents the charge and discharge power of the i-th fuel cell in the daily scheduling phase within time period t, P D-Fn (t) represents the charge and discharge power of the nth fuel cell in the daily scheduling phase within time period t, n represents the total number of fuel cells in the microgrid, and i represents the i-th fuel cell; P D-Lj (t) represents the charge and discharge power of the jth lithium battery in the intraday scheduling phase during time period t, P D-Lm (t) represents the charge and discharge power of the mth lithium battery in time period t during the intraday scheduling phase, m represents the total number of lithium batteries in the microgrid, and j represents the jth lithium battery.
8. A microgrid economic dispatch system based on deep reinforcement learning, characterized in that: include: The first unit is used to establish and solve the economic optimization scheduling model of the microgrid based on the day-ahead forecast data of the uncontrollable units during the day-ahead scheduling phase, and obtain the day-ahead scheduling plan for each device in the microgrid; The second unit is used to perform Markov decision process modeling using a deep neural network in the intraday pre-scheduling stage based on the day-ahead forecast data of the uncontrollable units and the economic optimization scheduling model to obtain an intraday scheduling model; The third unit is configured to input the intraday ultra-short-term forecast data and the day-ahead scheduling plan data of the uncontrollable units into the intraday scheduling model to obtain output data during the intraday scheduling phase, and use the output data as a scheduling basis for the controllable units; Among them, the microgrid includes uncontrollable units and controllable units.
9. An electronic device, characterized in that: include: Memory, processor; The processor is used to read and execute the computer program stored in the memory to implement the microgrid economic dispatch method based on deep reinforcement learning as described in any one of claims 1-7.
10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer-executable instructions, which, when executed, implement the microgrid economic dispatch method based on deep reinforcement learning described in any one of claims 1 to 7.