Deep reinforcement learning power distribution network-microgrid optimal scheduling method based on dichotomy

Through a deep reinforcement learning algorithm based on dichotomy, the distribution grid-micro grid system model is built, and the agent is trained to perform optimal scheduling, solving the problem of resource integration under the traditional scheduling mode, and achieving low-cost and efficient system collaborative operation.

CN120511643APending Publication Date: 2025-08-19STATE GRID JIANGSU ELECTRIC POWER CO LTD NANJING POWER SUPPLY COMPANY +2
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510537511.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-27
Publication Date
2025-08-19

AI Technical Summary

Technical Problem

The traditional centralized management model is difficult to effectively integrate and utilize small and micro resources in distributed ways, resulting in high scheduling costs and slow response speed, making it difficult to achieve efficient coordinated operation of distribution network-micro grid systems.

Method used

A deep reinforcement learning algorithm based on dichotomy is adopted to build a distribution network-micro grid system model, and train an agent to achieve optimal scheduling, and efficient search of the action space combined with dichotomy is used to optimize the coordinated operation of the distribution network-micro grid.

Benefits of technology

It reduces the operating costs of the distribution network system, improves the comprehensive benefits of the microgrid, and improves the collaborative operation efficiency of the distribution network-microgrid system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120511643A_ABST
    Figure CN120511643A_ABST
Patent Text Reader

Abstract

The invention discloses a deep reinforcement learning power distribution network-micro-grid optimal scheduling method based on a dichotomy, and the method comprises the following steps: constructing a power distribution network-micro-grid system model which comprises a first scheduling model and a second scheduling model; training an intelligent agent by using a deep reinforcement learning algorithm, and searching an action space of the deep reinforcement learning algorithm by using a dichotomy during training; and performing real-time optimal scheduling on the power distribution network-micro-grid system model by using the trained intelligent agent. Compared with the existing method, the method provided by the invention has the advantages that the operation cost of the power distribution network system is reduced, the comprehensive benefit of the micro-grid is improved, and the cooperative operation efficiency of the power distribution network-micro-grid system is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of distribution network-microgrid system scheduling methods, and specifically to a distribution network-microgrid optimal scheduling method based on deep reinforcement learning of dichotomy. Background Art

[0002] Building a new power system dominated by renewable energy has become an inevitable trend. However, the intermittent and volatile nature of renewable energy poses significant challenges to grid resilience. Urban power grids face challenges such as supply risk warning, renewable energy absorption, peak voltage drop, and congestion management. The massive influx of distributed power sources, energy storage devices, and electric vehicles connected to the distribution network has spawned a large number of small and micro resources. These resources hold enormous potential, but their high interaction costs and difficulty in participating in regulation make them difficult to effectively participate in grid dispatch. Traditional centralized management models, such as virtual power plants, face challenges with high dispatch costs and slow response times when faced with a massive influx of small and micro resources, making it difficult to effectively integrate and utilize these resources. Summary of the Invention

[0003] The present invention provides a deep reinforcement learning distribution network-microgrid optimal scheduling method based on dichotomy to solve the optimal scheduling problem of the distribution network-microgrid system, improve the operating efficiency of the distribution-microgrid system and the benefits of each entity.

[0004] In order to achieve the above object, the technical solution adopted by the present invention is:

[0005] The optimal dispatching method of distribution network-microgrid based on deep reinforcement learning of bisection method includes the following steps:

[0006] Constructing a distribution network-microgrid system model, wherein the distribution network-microgrid system model includes a first scheduling model and a second scheduling model, wherein the first scheduling model includes a distribution network system scheduling model and a distribution network-microgrid node power price model, and the second scheduling model includes a microgrid system scheduling model and a microgrid energy storage and power flow constraint variable model;

[0007] The deep reinforcement learning algorithm is used to train the intelligent agent with the state of the distribution network-microgrid system as input, and the dichotomy method is used to search the action space of the deep reinforcement learning algorithm. Through training, the intelligent agent outputs the optimal scheduling action, and a trained intelligent agent is obtained;

[0008] The trained intelligent agent is used to perform real-time optimal scheduling of the constructed distribution network-microgrid system model.

[0009] Furthermore, the distribution network system dispatching model in the first dispatching model constructs an objective function with the goal of minimizing the sum of the distribution network's power generation cost, the upper-level power grid's power purchase cost, the distribution network's energy storage system's charging and discharging cost, the load shedding penalty cost, and the distribution-micro interaction cost, and uses the generator climbing constraint, the distribution network's energy storage charge state constraint, the load shedding constraint, the branch flow constraint, the node voltage constraint, the line current constraint, and the node power constraint as the constraint conditions of the objective function.

[0010] Furthermore, the distribution network-microgrid node power price model in the first scheduling model is constructed through Lagrangian function transformation based on the objective function and constraints.

[0011] Furthermore, in step 1, the microgrid system scheduling model in the second scheduling model constructs an objective function with the goal of maximizing the difference between the distribution network-microgrid interaction benefit and the microgrid's energy storage charging and discharging cost and microgrid constraint penalty cost, and uses the microgrid's energy storage charge state constraint, node voltage constraint, line current constraint, node power constraint, and photovoltaic inverter operation constraint as constraints of the objective function.

[0012] Furthermore, in step 1, the microgrid energy storage and power flow constraint variable model in the second scheduling model includes a microgrid energy storage state of charge variable model and a branch power flow variable model.

[0013] Furthermore, the process of training the agent is as follows:

[0014] (2.1) Construct the state space, action space, and single-step reward value of the distribution network-microgrid system required for deep reinforcement learning algorithm training of intelligent agents. The state is used as the input of the reinforcement learning policy network at time t, and the corresponding action at time t is output. The reward at time t is calculated by the environment, and the state at time t+1 is updated based on the state and action at time t, where:

[0015] The state space is: {time t, distribution network active power pricing, distribution network reactive power pricing, distribution-microgrid boundary node voltage, microgrid load node active power, microgrid load node reactive power, microgrid photovoltaic node active power}

[0016] The action space is: {microgrid energy storage node active power, microgrid photovoltaic node reactive power};

[0017] The single-step reward value is the comprehensive benefit of the microgrid at time t;

[0018] (2.2) Initialize the time step t0 and set the total training steps Total step and the current training step Trainstep to 0;

[0019] (2.3) Initialize the deep reinforcement learning model;

[0020] (2.4) Obtain the system state information at the current time t from the constructed state space;

[0021] (2.5) Based on the current state information, a binary search method is used to select the corresponding action from the constructed action space, which represents the scheduling strategy for the microgrid;

[0022] (2.6) Perform the microgrid optimal power flow calculation to obtain the interaction power between the distribution network and the microgrid;

[0023] (2.7) Set the time counter tDN to 0 and initialize the reward value Reward to 0;

[0024] (2.8) Perform optimal power flow calculation on the distribution network and obtain the interactive electricity price between the distribution network and the microgrid;

[0025] (2.9) Calculate the comprehensive benefit of the microgrid at time t, add it to the reward value Reward, and increase the time counter tDN by 1;

[0026] (2.10) Determine whether the time counter tDN is less than 3: If yes, proceed to step (2.11); if not, refresh the distribution network status and return to step (2.4);

[0027] (2.11) Determine the value of tDN: If the time counter tDN is less than 3, read the active power of the PV node, the active power and reactive power of the load node, and the reactive power of the energy storage node based on the active power pricing and reactive power pricing of the distribution network, the voltage of the distribution-microgrid boundary node, and the state of charge information of the microgrid energy storage system. Then, increase the total number of training steps Trainstep by 1.

[0028] (2.12) Determine whether the total number of training steps Train step is less than the maximum number of training steps max Train step: If yes, return to step (2.4); if not, increase the total number of training steps Total step by 1;

[0029] (2.13) Determine whether the total number of training steps Total step is less than the maximum number of training steps max Total step: If yes, return to step (2.4); if not, the process ends.

[0030] Furthermore, in step (2.5), the binary search method is used to search and select the corresponding action from the constructed action space. The process is as follows:

[0031] (a) Initialization: Set the initial interval as (low, high), the interval error as ε, the maximum number of iterations as max_iterations, and initialize the number of iterations iterations to 0;

[0032] (b) Loop condition: When the number of iterations iterations is less than the maximum number of iterations max_iterations or the interval length (high - low) is greater than the interval error ε, execute the steps inside the loop;

[0033] (c) Calculate the intermediate value: Calculate the action intermediate value mid = (low + high) / 2, the action intermediate left - offset value left = mid - ε, and the action intermediate right - offset value right = mid + ε;

[0034] (d) Calculate the reward values: Calculate the reward values reward_mid, reward_left, and reward_right corresponding to mid, left, and right respectively;

[0035] (e) Compare the reward values and update the interval: If reward_left < reward_mid < reward_right or reward_mid < reward_left < reward_right, update the lower bound of the interval low = mid, and keep the upper bound unchanged; If reward_left < reward_right < reward_mid or reward_right < reward_left < reward_mid, update the lower bound of the interval low = left, and the upper bound high = right; Otherwise, update the upper bound of the interval high = mid, and keep the lower bound unchanged;

[0036] (f) Update the number of iterations: Increment the number of iterations iterations by 1;

[0037] (g) Return the result: After the loop ends, return the final interval (low, high).

[0038] A distribution network - micro - grid optimal scheduling search device includes a model construction module, a training module, and a scheduling module; the model construction module constructs the distribution network - micro - grid system model; the training module trains the agent with the state of the distribution network - micro - grid system as the input; the scheduling module obtains the trained agent from the training module, obtains the distribution network - micro - grid system model from the model construction module, and performs real - time optimal scheduling on the distribution network - micro - grid system model through the agent, thereby implementing the above - mentioned distribution network - micro - grid optimal scheduling method based on the dichotomy - based deep reinforcement learning.

[0039] A storage medium stores program instructions that can be read and executed. When the program instructions are read and executed, the above-mentioned deep reinforcement learning distribution network-microgrid optimal scheduling method based on dichotomy is executed.

[0040] An electronic device includes a processor and a memory, wherein the memory stores program instructions. When the program instructions are read and executed by the processor, the program instructions execute the above-mentioned deep reinforcement learning distribution network-microgrid optimal scheduling method based on bisection.

[0041] Compared with the prior art, the present invention has the following advantages:

[0042] Compared to traditional mathematical modeling solutions, this invention employs a deep reinforcement learning algorithm to train intelligent agents and constructs a scheduling model encompassing both the distribution network and microgrid systems. This approach, combined with a bisection approach to efficiently search the action space, yields an optimal scheduling strategy for the coordinated operation of the distribution network and microgrid. Results demonstrate that this approach reduces distribution network operating costs, increases the overall benefits of the microgrid, and improves the efficiency of the coordinated operation of the distribution network and microgrid systems compared to existing methods. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] Figure 1 It is a flow chart of a method according to an embodiment of the present invention.

[0044] Figure 2 This is a diagram of the distribution network configuration in an embodiment of the present invention.

[0045] Figure 3 1 is a microgrid configuration diagram according to an embodiment of the present invention.

[0046] Figure 4 is an interaction diagram between deep reinforcement learning and the matching-micro system in an embodiment of the present invention.

[0047] Figure 5 This is a flowchart of deep reinforcement learning in an embodiment of the present invention.

[0048] Figure 6 1 is a flow chart of the dichotomy method in an embodiment of the present invention. DETAILED DESCRIPTION

[0049] The present invention will be further described below with reference to the accompanying drawings and examples.

[0050] like Figure 1 As shown, this embodiment discloses a deep reinforcement learning distribution network-microgrid optimal scheduling method based on bisection, including the following steps:

[0051] Step 1: Construct a distribution network-microgrid system model. The distribution network-microgrid system model includes a first scheduling model and a second scheduling model. The first scheduling model includes a distribution network system scheduling model and a distribution network-microgrid node power price model. The second scheduling model includes a microgrid system scheduling model and a microgrid energy storage and power flow constraint variable model. The distribution network-microgrid system model in this embodiment is specifically described as follows:

[0052] (1) In the distribution network-microgrid system model of this embodiment, the distribution network system scheduling model in the first scheduling model, the distribution network topology diagram is as follows Figure 2 The distribution network model constructs an objective function based on the goal of minimizing the sum of the distribution network's power generation cost, the upstream power grid's power purchase cost, the distribution network's energy storage system's charging and discharging cost, the load shedding penalty cost, and the distribution-microgrid interaction cost. The objective function is shown in the following formula:

[0053] minF DN =F4+F5+F6+F7+F1(A-1)

[0054] In formula (A-1): F DN , F4, F5, F6, F7, and F1 are the distribution network operation cost, distribution network power generation cost, upper-level power grid power purchase cost, distribution network energy storage system charging and discharging cost, load shedding penalty cost, and distribution-micro-interaction cost, respectively.

[0055] Among them, the power generation cost F4 of the distribution network is as follows:

[0056]

[0057] In formula (A-2): t represents the time period; g represents the generator group number; a g 、b g 、c g are the quadratic, linear and constant cost coefficients of the generator set g respectively; P g,t is the output power of generator set g at time t.

[0058] The power purchase cost F5 of the upper-level power grid is as follows:

[0059]

[0060] In formula (A-3), sub represents the root node where the distribution network is connected to the upper power grid; c sub,t is the electricity purchase price of the root node at time t; P sub,t is the power purchased by the root node at time t.

[0061] The energy storage charging and discharging cost F6 of the distribution network is as follows:

[0062]

[0063] In formula (A-4): DNe represents the energy storage system of the distribution network; c DNe The unit power charging and discharging cost of the energy storage system in the distribution network; and are the discharge power and charging power of the energy storage system belonging to the distribution network at time t respectively.

[0064] The load shedding penalty cost F7 is as follows:

[0065]

[0066] In formula (A-5), cut represents load shedding; i represents node number; c cut is the unit penalty cost coefficient for load shedding; is the load shedding power of node i at time t.

[0067] The cost of micro-interaction F1 is as follows:

[0068]

[0069] In formula (A-6), MG represents the microgrid; m represents the node number where the microgrid is located; a represents the active power; is the active power interacting between microgrid node m and distribution network at time t; is the pricing of unit active power by the distribution network for microgrid node m when it is connected to the grid at time t.

[0070] The distribution network system dispatch model in the first dispatch model of this embodiment uses generator ramp constraints, distribution network energy storage charge state constraints, load shedding constraints, branch flow constraints, node voltage constraints, line current constraints, and node power constraints as the constraint conditions of the objective function. Among them:

[0071] The generator ramp constraint is as follows:

[0072] P g,t -P g,t-1 ≤RU g (A-7)

[0073] P g,t -P g,t-1 ≤RD g (A-8)

[0074] In formula (A-7) and (A-8): P g,t-1 is the output power of generator set g at time t-1; RU g and RD g They are respectively increasing the upper limit of output power and reducing the lower limit of output power of generator set g per unit time.

[0075] The state of charge constraint of the energy storage in the distribution network is as follows:

[0076]

[0077] SOC DNe,min ≤SOC DNe,t ≤SOC DNe,max (A-10)

[0078] In formulas (A-9) and (A-10), DN represents the distribution network; SOC DNe,t and SOC DNe,t+1 are the charge states of the energy storage systems in the distribution network at time t and time t+1 respectively; and They are the operating state coefficients of the energy storage system of the distribution network at time t, and when the energy storage system of the distribution network is charging and When the energy storage system of the distribution network discharges and η ch and η dis are the charging and discharging efficiencies of the energy storage system in the distribution network; Δt DN Calculate time intervals for distribution networks; SOC DNe,max and SOC DNe,min They are the upper and lower limits of the state of charge of the energy storage system DNe respectively.

[0079] The load shedding constraint is as follows:

[0080]

[0081] In formula (A-11): P i,t is the active power of node i at time t; N I A collection of system load nodes.

[0082] The branch power flow constraint is as follows:

[0083]

[0084] In formulas (A-12), (A-13), (A-14), and (A-15), j and k represent node numbers; m represents the node number where the microgrid is located; ij and ki represent the branches between nodes i and j and between nodes k and i, respectively; F DN (i) is the set of end nodes of the branch with node i as the head node in the distribution network; T DN (i) is the set of head-end nodes of the branch with node i as the terminal node in the distribution network; N DN is the set of distribution network nodes; B DN It is a collection of distribution network branches; is the active power of node i after load shedding at time t; Q i,t is the reactive power of node i at time t; P ij,t and Q ij,t are the active power and reactive power of branch ij at time t respectively; P ki,t and Q ki,t are the active power and reactive power of branch ki at time t; I ki,t and I ij,t are the squares of the currents of branch ki and branch ij at time t; R ki and R ij are the resistances of branch ki and branch ij respectively; X ki and X ij are the reactances of branch ki and branch ij respectively; V i,t and V j,t are the squares of the voltages at nodes i and j at time t, respectively; MG represents the microgrid; P and Q are active power and reactive power, respectively; and are the Lagrange multipliers of the active power and reactive power of microgrid node m at time t, respectively.

[0085] The node voltage constraint is as follows:

[0086]

[0087] In formula (A-16): v i,t is the voltage amplitude of node i at time t; v i,min and v i,max are the minimum and maximum voltage amplitudes of node i, respectively.

[0088] The line current constraint is as follows:

[0089]

[0090] In formula (A-17): ij,max is the square of the maximum current amplitude of branch ij.

[0091] The node power constraint is as follows:

[0092] 0≤P g,t ≤P g,max (A-18)

[0093] P sub,min ≤P sub,t ≤P sub,max (A-19)

[0094]

[0095] In formula (A-18), (A-19), (A-20), (A-21), (A-22), (A-23), (A-24): P g,max 、P sub,max 、 P i,max and are the upper limits of active power of generator group g, root node, energy storage system discharge, charging, node i, and microgrid node m, respectively; P sub,min is the lower limit of active power of the root node; is the reactive power upper limit of microgrid node m; and are the active power and reactive power of the interaction between the microgrid node m and the distribution network at time t, respectively.

[0096] In the first scheduling model of this embodiment, a distribution network-microgrid node power price model based on dual variables is constructed based on the objective function and the constraints through Lagrangian function transformation. The process is as follows:

[0097] (a) The simplified distribution network system model is shown as follows:

[0098]

[0099] In formula (A-25): x is the constraint variable; F DN (x) is the operating cost of the distribution network system; h(x) is the equality constraint of the distribution network system; g(x) is the inequality constraint of the distribution network system; b max and b min are the upper and lower limit constraint vectors of the distribution network system respectively.

[0100] (b) The Lagrangian function transformation of the simplified model of the distribution network system is shown as follows:

[0101]

[0102] In formula (A-26), L is the Lagrangian function; λ is the dual multiplier vector of the distribution network system equation; and are the upper and lower limit multiplier vectors of the inequality constraints of the distribution network system, respectively.

[0103] (c) The power price of the distribution network-microgrid node is constructed according to the Lagrangian function transformation formula as shown below:

[0104]

[0105] In formulas (A-27) and (A-28), a and r represent active power and reactive power; and are the pricing of unit active power and reactive power of the distribution network under the condition of grid connection of microgrid node m at time t, respectively.

[0106] (2) In the distribution network-microgrid system model of this embodiment, the microgrid system scheduling model in the second scheduling model, the microgrid topology diagram is as follows Figure 3 The microgrid model constructs an objective function based on the goal of maximizing the difference between the distribution network-microgrid interaction benefit and the microgrid's energy storage charging and discharging cost and the microgrid constraint penalty cost, as shown in the following formula:

[0107] maxF MG =F1-F2-F3(A-29)

[0108] In formula (A-29): F MG is the comprehensive benefit of the microgrid, i.e., the objective function; F1, F2, and F3 are the distribution-microgrid interaction benefit, the microgrid's energy storage charging and discharging cost, and the microgrid constraint penalty cost, respectively.

[0109] Among them, the pair-micro interaction benefit F1 is shown in the above formula (A-6);

[0110] The energy storage charging and discharging cost F2 of the microgrid is as follows:

[0111]

[0112] In formula (A-30), MGe is the energy storage system of the microgrid; c MGe The unit power charging and discharging cost of the energy storage system of the microgrid; and are the discharge power and charging power of the energy storage system of the microgrid at time t respectively.

[0113] The microgrid constraint penalty cost F3 is as follows:

[0114] F3=α soft f(x)+α hard f(x) (A-31)

[0115]

[0116] In formula (A-31), (A-32): α soft and α hard are the microgrid critical constraint penalty coefficient and the dangerous constraint penalty coefficient respectively; f(x) is the microgrid constraint penalty function; x is the constraint variable; x max and x min are the upper and lower limits of the microgrid constraint variables respectively; λ soft is the critical constraint interval partition coefficient; γ is the set of all constraints of the microgrid.

[0117] The distribution network system dispatch model in the second dispatch model of this embodiment uses the microgrid's energy storage charge state constraint, node voltage constraint, line current constraint, node power constraint, and photovoltaic inverter operation constraint as the constraint conditions of the objective function. Among them:

[0118] The state of charge constraint of the energy storage in the microgrid is as follows:

[0119] SOC MGe,min ≤SOC MGe,t ≤SOC MGe,max (A-33)

[0120] In formula (A-33): SOC MGe,t is the state of charge of the energy storage system of the microgrid at time t; SOC MGe,max and SOC MGe,min They are the upper and lower limits of the state of charge of the energy storage system of the microgrid.

[0121] The node voltage constraint is as follows:

[0122]

[0123] In formula (A-34): N MG is the set of microgrid nodes.

[0124] The line current constraint is as follows:

[0125]

[0126] In formula (A-35): B MG is a collection of microgrid branches.

[0127] The node power constraint is as follows:

[0128]

[0129]

[0130] In formula (A-36), (A-37), (A-38): and are the upper limits of discharge and charging active power of the energy storage system belonging to the microgrid, respectively.

[0131] The operating constraints of the PV inverter are as follows:

[0132]

[0133] In formulas (A-39) and (A-40): and They are the upper limit of reactive power generated by the microgrid photovoltaic inverter and the upper limit of the apparent power of the photovoltaic node; and are the active power and reactive power generated by the microgrid photovoltaic inverter at time t respectively.

[0134] In the second scheduling model of this embodiment, the constructed microgrid energy storage and power flow constraint variable model includes the microgrid energy storage charge state variable model and the branch power flow variable model, wherein:

[0135] The variable model of the energy storage state of charge of the microgrid is as follows:

[0136]

[0137] In formula (A-41): SOC MGe,t+1 is the state of charge of the energy storage system of the microgrid at time t+1;

[0138] μ k is the self-discharge coefficient of the energy storage system of the microgrid; and They are the operating state coefficients of the energy storage system of the microgrid at time t, and when the energy storage system of the microgrid is charging and When the energy storage system of the microgrid discharges and η ch and η dis are the charging and discharging efficiencies of the energy storage system of the microgrid; Δt MG Calculate time intervals for the microgrid.

[0139] The branch power flow variable model is shown as follows:

[0140]

[0141] In formula (A-42), (A-43), (A-44): F MG (i) is the set of end nodes of the branch with node i as the head node in the microgrid; T MG (i) is the set of head-end nodes of the branch with node i as the terminal node in the microgrid.

[0142] Step 2: Use the deep reinforcement learning algorithm to train an intelligent agent. The state of the distribution network-microgrid system is used as input during training, and the binary search method is used to search the action space of the deep reinforcement learning algorithm during training. Through training, the intelligent agent outputs the optimal scheduling action.

[0143] In this embodiment, Figure 4 、 Figure 5As shown in the figure, the state is used as the input of the reinforcement learning policy network at time t, and the output is the corresponding action at time t. The reward at time t is calculated by the environment, and the state at time t+1 is updated based on the state and action at time t. The process of training an agent using a deep reinforcement learning algorithm is as follows:

[0144] (2.1) Construct the state space, action space, and single-step reward value of the distribution network-microgrid system required for deep reinforcement learning algorithm training of intelligent agents, where:

[0145] The state space is: {time t, distribution network active power pricing, distribution network reactive power pricing, distribution-microgrid boundary node voltage, microgrid load node active power, microgrid load node reactive power, microgrid photovoltaic node active power}

[0146] The action space is: {microgrid energy storage node active power, microgrid photovoltaic node reactive power};

[0147] The single-step reward value is the comprehensive benefit of the microgrid at time t;

[0148] (2.2) Initialize the time step t0 and set the total training steps Total step and the current training step Trainstep to 0;

[0149] (2.3) Initialize the deep reinforcement learning model;

[0150] (2.4) Obtain the system state information at the current time t from the constructed state space;

[0151] (2.5) Based on the current state information, a binary search method is used to select the corresponding action from the constructed action space, which represents the scheduling strategy for the microgrid;

[0152] (2.6) Perform the microgrid optimal power flow calculation to obtain the interaction power between the distribution network and the microgrid;

[0153] (2.7) Set the time counter tDN to 0 and initialize the reward value Reward to 0;

[0154] (2.8) Perform optimal power flow calculation on the distribution network and obtain the interactive electricity price between the distribution network and the microgrid;

[0155] (2.9) Calculate the comprehensive benefit of the microgrid at time t, add it to the reward value Reward, and increase the time counter tDN by 1;

[0156] (2.10) Determine whether the time counter tDN is less than 3: If yes, proceed to step (2.11); if not, refresh the distribution network status and return to step (2.4);

[0157] (2.11) Determine the value of tDN: If the time counter tDN is less than 3, then read the active power of the PV node, the active power and reactive power of the load node, and the reactive power of the energy storage node based on the active power pricing and reactive power pricing of the distribution network, the voltage of the distribution-microgrid boundary node, and the state of charge information of the microgrid energy storage system, and increase the total number of training steps Train step by 1;

[0158] (2.12) Determine whether the total number of training steps Train step is less than the maximum number of training steps max Train step: If yes, return to step (2.4); if not, increase the total number of training steps Total step by 1;

[0159] (2.13) Determine whether the total number of training steps Total step is less than the maximum number of training steps max Total step: If yes, return to step (2.4); if not, the process ends.

[0160] Among them Figure 6 As shown in step (2.5), the process of using the binary search method to search and select the corresponding action from the constructed action space is as follows:

[0161] (a) Initialization: Set the initial interval to (low, high), the interval error to ε, the maximum number of iterations to max_iterations, and initialize the number of iterations to 0;

[0162] (b) Loop condition: When the number of iterations is less than the maximum number of iterations max_iterations or the interval length (high-low) is greater than the interval error ε, the steps in the loop are executed;

[0163] (c) Calculate the middle value: calculate the middle value of the action mid = (low + high) / 2, the left deviation value of the middle action left = mid - ε, and the right deviation value of the middle action right = mid + ε;

[0164] (d) Calculate the reward value: calculate the reward value reward_mid, reward_left and reward_right corresponding to mid, left and right respectively;

[0165] (e) Compare the reward values and update the interval: If reward_left < reward_mid < reward_right or reward_mid < reward_left < reward_right, update the lower bound of the interval low = mid, and keep the upper bound of the interval unchanged; If reward_left < reward_right < reward_mid or reward_right < reward_left < reward_mid, update the lower bound of the interval low = left, and the upper bound of the interval high = right; Otherwise, update the upper bound of the interval high = mid, and keep the lower bound of the interval unchanged;

[0166] (f) Update the iteration count: Increment the iteration count iterations by 1;

[0167] (g) Return the result: After the loop ends, return the final interval (low, high).

[0168] Step 3. Use the agent trained in Step 2 to perform real-time optimal scheduling on the distribution network - microgrid system model constructed in Step 1. Specifically, solve the distribution network - microgrid system model constructed in Step 1 as shown in Formulas (A - 1)-(A - 44) through the trained agent to obtain the following parameters:

[0169] The distribution network - microgrid system model obtains the minimum distribution network operating cost F DN and the maximum microgrid comprehensive revenue F MG ; Obtain the distribution network side generation cost F4, the upstream grid power purchase cost F5, the charge - discharge cost F6 of the energy storage system belonging to the distribution network, the load shedding cost F7, and the distribution - micro interaction cost F1, and obtain the distribution network side generator output P g,t the state of charge SOC of the energy storage system belonging to the distribution network DNe,t the load shedding amount the active power after load shedding the active power P of the node i,t and the reactive power Q i,t the square of the node voltage V i,t the square of the line current I ij,t the active power price of the distribution - micro node and the reactive power price Obtain the distribution - micro interaction revenue F1, the charge - discharge cost F2 of the energy storage system belonging to the microgrid, and the microgrid constraint penalty cost F3 on the microgrid side, and obtain the state of charge SOC of the energy storage system on the microgrid side MGe,t the square of the node voltage V i,t the square of the line current I ij,t the active power of the PV inverter and the reactive power Node active power P i,t and reactive power Q i,t .

[0170] Based on the above parameters obtained by solution, real-time optimal scheduling of the distribution network-microgrid system model can be achieved.

[0171] This embodiment also discloses a distribution network-microgrid optimal scheduling search device, comprising a model construction module, a training module, and a scheduling module. The model construction module constructs the distribution network-microgrid system model; the training module trains the intelligent agent using the distribution network-microgrid system state as input; the scheduling module obtains the trained intelligent agent from the training module and the distribution network-microgrid system model from the model construction module, and uses the intelligent agent to perform real-time optimal scheduling of the distribution network-microgrid system model. Thus, steps 1-3 of the aforementioned dichotomy-based deep reinforcement learning distribution network-microgrid optimal scheduling method are implemented through the model construction module, the training module, and the scheduling module.

[0172] This embodiment also discloses a storage medium storing program instructions that can be read and executed. When the program instructions are read and executed, steps 1 to 3 of the above-mentioned deep reinforcement learning distribution network-microgrid optimal scheduling method based on bisection are executed.

[0173] This embodiment also discloses an electronic device, including a processor and a memory, wherein program instructions are stored in the memory. When the program instructions in the memory are read and executed by the processor, steps 1 to 3 of the above-mentioned deep reinforcement learning distribution network-microgrid optimal scheduling method based on bisection are executed.

[0174] The preferred embodiments of the present invention are described in detail above with reference to the accompanying drawings. The embodiments described in the present invention are merely descriptions of the preferred embodiments of the present invention and do not limit the concept and scope of the present invention. The various specific technical features described in the above specific embodiments can be combined in any suitable manner unless there is any contradiction. Such combinations should also be regarded as the contents disclosed in this disclosure as long as they do not violate the concept of the present invention. In order to avoid unnecessary repetition, the present invention will not further describe various possible combinations.

[0175] The present invention is not limited to the specific details of the above-mentioned embodiments. Within the scope of the technical concept of the present invention and without departing from the design concept of the present invention, various modifications and improvements made to the technical solution of the present invention by those skilled in the art should fall within the scope of protection of the present invention. The technical contents for which protection is sought in the present invention have been fully recorded in the claims.

Claims

1. A deep reinforcement learning distribution network-microgrid optimal scheduling method based on bisection method, characterized by: The following steps are involved: Constructing a distribution network-microgrid system model, wherein the distribution network-microgrid system model includes a first scheduling model and a second scheduling model, wherein the first scheduling model includes a distribution network system scheduling model and a distribution network-microgrid node power price model, and the second scheduling model includes a microgrid system scheduling model and a microgrid energy storage and power flow constraint variable model; The deep reinforcement learning algorithm is used to train the intelligent agent with the state of the distribution network-microgrid system as input, and the dichotomy method is used to search the action space of the deep reinforcement learning algorithm. Through training, the intelligent agent outputs the optimal scheduling action, and a trained intelligent agent is obtained; The trained intelligent agent is used to perform real-time optimal scheduling of the constructed distribution network-microgrid system model.

2. The distribution network-microgrid optimal scheduling method based on deep reinforcement learning of bisection method according to claim 1 is characterized in that: The distribution network system dispatching model in the first dispatching model constructs an objective function with the goal of minimizing the sum of the distribution network's power generation cost, the upper-level power grid's power purchase cost, the distribution network's energy storage system's charging and discharging cost, the load shedding penalty cost, and the distribution-micro-interaction cost, and uses the generator climbing constraint, the distribution network's energy storage charge state constraint, the load shedding constraint, the branch flow constraint, the node voltage constraint, the line current constraint, and the node power constraint as the constraint conditions of the objective function.

3. The distribution network-microgrid optimal scheduling method based on deep reinforcement learning of bisection method according to claim 2 is characterized in that: The distribution network-microgrid node power price model in the first scheduling model is constructed through Lagrangian function transformation based on the objective function and constraint conditions.

4. The distribution network-microgrid optimal scheduling method based on deep reinforcement learning of bisection method according to claim 1 is characterized in that: In step 1, the microgrid system scheduling model in the second scheduling model constructs an objective function with the goal of maximizing the difference between the distribution network-microgrid interaction benefit and the microgrid's energy storage charging and discharging cost and microgrid constraint penalty cost, and uses the microgrid's energy storage charge state constraint, node voltage constraint, line current constraint, node power constraint, and photovoltaic inverter operation constraint as the constraint conditions of the objective function.

5. The distribution network-microgrid optimal scheduling method based on deep reinforcement learning of bisection method according to claim 1 is characterized in that: In step 1, the microgrid energy storage and power flow constraint variable models in the second scheduling model include the microgrid energy storage charge state variable model and the branch power flow variable model.

6. The distribution network-microgrid optimal scheduling method based on deep reinforcement learning of bisection method according to claim 1 is characterized in that: The process of training an agent is as follows: (2.1) Construct the state space, action space, and single-step reward value of the distribution network-microgrid system required for deep reinforcement learning algorithm training of intelligent agents. The state serves as the input of the reinforcement learning policy network at time t, and the corresponding action at time t is output. The reward at time t is calculated by the environment, and the state at time t+1 is updated based on the state and action at time t, where: The state space is: { time t , distribution network active power pricing, distribution network reactive power pricing, distribution - micro boundary node voltage, microgrid load node active power, microgrid load node reactive power, microgrid photovoltaic node active power} The action space is: {microgrid energy storage node active power, microgrid photovoltaic node reactive power}; The single-step reward value is the comprehensive benefit of the microgrid at time t; (2.2) Initialize the time step t0 and set the total training steps Total step and the current training step Trainstep to 0; (2.3) Initialize the deep reinforcement learning model; (2.4) Obtain the system state information at the current time t from the constructed state space; (2.5) Based on the current state information, a binary search method is used to select the corresponding action from the constructed action space, which represents the scheduling strategy for the microgrid; (2.6) Perform the optimal power flow calculation for the microgrid and obtain the interaction power between the distribution network and the microgrid; (2.7) Set the time counter tDN to 0 and initialize the reward value Reward to 0; (2.8) Perform optimal power flow calculation on the distribution network and obtain the interactive electricity price between the distribution network and the microgrid; (2.9) Calculate the comprehensive benefit of the microgrid at time t, add it to the reward value Reward, and increase the time counter tDN by 1; (2.10) Determine whether the time counter tDN is less than 3: If yes, proceed to step (2.11); if not, refresh the distribution network status and return to step (2.4); (2.11) Determine the value of tDN: If the time counter tDN < 3, read the active power of the PV node, the active power and reactive power of the load node, and the reactive power of the energy storage node based on the active power pricing and reactive power pricing of the distribution network, the voltage of the distribution-microgrid boundary node, and the charge state information of the microgrid energy storage system. Then, increase the total number of training steps Trainstep by 1. (2.12) Determine whether the total number of training steps Train step is less than the maximum number of training steps max Train step: If yes, return to step (2.4); if not, increase the total number of training steps Total step by 1; (2.13) Determine whether the total number of training steps Total step is less than the maximum number of training steps max Total step: If yes, return to step (2.4); if not, the process ends.

7. The distribution network-microgrid optimal scheduling method based on deep reinforcement learning of bisection method according to claim 6 is characterized in that: In step (2.5), the binary search method is used to search and select the corresponding action from the constructed action space. The process is as follows: (a) Initialization: Set the initial interval to (low, high), the interval error to ε, the maximum number of iterations to max_iterations, and initialize the number of iterations to 0; (b) Loop condition: When the number of iterations is less than the maximum number of iterations max_iterations or the interval length (high - low) is greater than the interval error ε, the steps in the loop are executed; (c) Calculate the middle value: calculate the middle value of the action mid = (low + high) / 2, the left deviation value of the middle of the action left = mid - ε, and the right deviation value of the middle of the action right = mid + ε; (d) Calculate the reward value: Calculate the reward values reward_mid, reward_left and reward_right corresponding to mid, left and right respectively; (e) Compare the reward values and update the interval: If reward_left < reward_mid < reward_right or reward_mid < reward_left < reward_right, update the lower bound of the interval low = mid, and keep the upper bound of the interval unchanged; If reward_left < reward_right < reward_mid or reward_right < reward_left < reward_mid, update the lower bound of the interval low = left, and the upper bound of the interval high = right; Otherwise, update the upper bound of the interval high = mid, and keep the lower bound of the interval unchanged; (f) Update the iteration count: Increment the iteration count iterations by 1; (g) Return the result: After the loop ends, return the final interval (low, high).

8. A distribution network-microgrid optimal scheduling search device, characterized in that: It includes a model construction module, a training module, and a scheduling module; the model construction module constructs the distribution network - microgrid system model; the training module trains the agent with the state of the distribution network - microgrid system as the input; the scheduling module obtains the trained agent from the training module, obtains the distribution network - microgrid system model from the model construction module, and performs real - time optimal scheduling on the distribution network - microgrid system model through the agent, thereby implementing the optimal scheduling method for the distribution network - microgrid based on dichotomy - based deep reinforcement learning according to any one of claims 1 - 7.

9. A storage medium storing program instructions that can be read and executed, characterized in that: When the program instructions are read and run, execute the optimal scheduling method for the distribution network - microgrid based on dichotomy - based deep reinforcement learning according to any one of claims 1 - 7.

10. An electronic device comprising a processor and a memory, wherein the memory stores program instructions, characterized in that: When the program instructions are read and run by the processor, execute the optimal scheduling method for the distribution network - microgrid based on dichotomy - based deep reinforcement learning according to any one of claims 1 - 7.

Citation Information

Cited By

  • Method for comprehensively evaluating carrying capacity of power distribution network on electric vehicle and energy storage load

    CN121352388A