Two-stage edge cloud cooperative scheduling method for park power distribution network containing micro-grid group
By adopting a two-stage edge-cloud collaborative scheduling method of deep reinforcement learning in the active distribution network, combined with SAC and LSTM, the problems of high-dimensional state interaction and uncertainty fluctuations of the micronet group are solved, efficient unified coordinated scheduling of the micronet group is realized, and the operation efficiency and safety of the distribution network are improved.
Patent Information
- Application Number
- CN202510663512.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-22
- Publication Date
- 2025-08-29
AI Technical Summary
In active distribution networks containing micronet groups, traditional optimization methods are insufficient in the face of high-dimensional state interaction, uncertainty fluctuations and multi-time scale scheduling characteristics, and it is difficult to effectively coordinate the consumption and scheduling of distributed energy.
A two-stage edge-cloud collaborative scheduling method based on deep reinforcement learning is adopted, combined with soft strategy gradient algorithm (SAC) and long and short-term memory network (LSTM), a cloud and edge agent are built to realize unified coordinated scheduling of micronet groups, and the controllable resource configuration of microgrids and distribution networks is optimized through recent and intraday scheduling strategies.
It improves the operating efficiency and safety of the active distribution network, reduces the dependence on the external power grid, achieves a balance between economy and voltage safety, and improves the continuity and stability of the scheduling process.
Smart Images

Figure CN120566409A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of distribution network dispatching, and specifically relates to a two-stage edge-cloud collaborative dispatching method for a campus distribution network containing a microgrid cluster. Background Art
[0002] By integrating distributed power sources, loads, and energy storage systems into microgrids and integrating them with active distribution networks (ADNs), the absorption of distributed energy can be effectively improved. However, with the highly coupled "source-grid-load-storage" system, traditional optimization methods face challenges in efficiency and adaptability due to the high-dimensional state interactions, uncertain fluctuations, and multi-timescale scheduling characteristics of microgrid clusters. Deep reinforcement learning (DRL), with its ability to autonomously learn policies and map state to action, has become an important approach to solving power system scheduling problems. Furthermore, the edge-cloud collaborative scheduling architecture, by establishing a layered "cloud-edge" collaborative mechanism within the power system, combines global control with local rapid response, providing key support for building efficient and intelligent ADNs. Therefore, combining edge-cloud collaboration with DRL is an effective means of addressing the degradation of distribution network scheduling performance caused by source-load fluctuations.
[0003] At present, research on ADN scheduling mainly focuses on the optimal configuration of source and load resources, and the economic efficiency, safety, and stability of operation. In existing technologies, the coordination problem of new energy and load in the distribution network is solved by combining demand response with hydrogen and ammonia. Jiang Tao et al. proposed a day-ahead and real-time optimization scheduling method for transmission-distribution coordination based on vehicle-grid interaction to improve the economy and flexibility of the system. Zheng Shunwei et al. constructed a robust collaborative optimization model with a min-max-min structure for the optimal operation of ADN containing multiple microgrids, achieving the optimal operating cost. However, with the diversification of ADN scheduling resources, scheduling problems have become more complex, and traditional optimization methods face huge challenges in solving efficiency and adaptability.
[0004] In response to the problems of high uncertainty and high real-time requirements in active distribution network scheduling, deep reinforcement learning uses deep neural networks (DNN) to model state-action mapping, effectively alleviating the problems of traditional optimization methods in the scheduling of multiple types of resources in distribution networks, and has good adaptability and strategy generalization capabilities. Wang Ke et al. proposed a dynamic economic scheduling method for ADN based on deep reinforcement learning. By constructing an adaptive state space and PPO intelligent agent, the scheduling economy and real-time performance of the distribution network were improved. Wei Zhinong et al. constructed a two-stage scheduling framework for hydrogen energy transportation and multi-microgrid scenarios, which improved the resilience and scheduling flexibility of the distribution system. Although the DRL method has shown good performance in ADN scheduling, the scheduling method is still immature in the multi-time scale optimization of ADN for microgrid groups. Summary of the Invention
[0005] To address the increased difficulty in optimizing resource scheduling caused by microgrid clusters accessing active distribution networks, this paper proposes a two-stage edge-cloud collaborative scheduling method for campus distribution networks containing microgrid clusters. This method achieves unified and coordinated scheduling of the controllable resources of the microgrid cluster and the ADN. The intraday scheduling strategy of the present invention not only considers the local operating characteristics of the microgrid, but also takes into account the overall operational safety requirements of the distribution network, thereby improving the node voltage margin of ADN operation.
[0006] In order to achieve the above object, the present invention is implemented through the following technical solutions: A two-stage edge-cloud collaborative scheduling method for a campus distribution network containing a microgrid cluster is performed according to the following steps;
[0007] Step S1: establishing an active distribution network dispatching model with the goal of minimizing the total operating cost of the active distribution network. The active distribution network dispatching model includes a day-ahead dispatching model of the active distribution network including a microgrid cluster and an intraday dispatching model of the microgrid cluster.
[0008] Step S2: A day-ahead dispatch strategy based on the soft policy gradient (SAC) algorithm is constructed. The cloud-based intelligent agent takes minimizing the operating cost of the active distribution network as its goal. The policy network is used to train the model based on the input active distribution network status information. The cloud-based intelligent agent uses the execution results as feedback to update the experience pool and optimize the SAC model.
[0009] Step S3: Construct an intraday dispatch strategy that combines the long short-term memory network and the soft policy gradient algorithm, with the goal of minimizing the mean square error of the current microgrid operating cost and the day-ahead action deviation. The long short-term memory network is embedded in the actor strategy network front end of the soft policy gradient algorithm framework to extract the local state sequence of the microgrid and generate dispatch actions.
[0010] Furthermore, the objective function of the active distribution network day-ahead dispatch model in step S1 is as follows:
[0011]
[0012] Among them, C ADN is the total operating cost of the active distribution network, C pcc is the electricity purchase cost, C mg is the total operating cost of the microgrid, C loss is the network loss cost, T is the total dispatching time;
[0013] Power purchase cost C pcc : Cost of purchasing electricity from the active distribution network to the upper level
[0014]
[0015] Where T is the total scheduling time, is the time-of-use electricity price at time t, P cc (t) is the power purchased from the upper grid at time t, and Δt is the time interval;
[0016] The total operating cost of the microgrid C mg :
[0017]
[0018] Where C m is the comprehensive operation and maintenance cost of the microgrid, C curt is the cost of curtailing wind and solar power, C G is the power generation cost of the diesel engine, m mg is the number of microgrids;
[0019]
[0020] Where a i 、b i 、c i is the fuel cost coefficient of diesel engine power generation, λ i , ρ i is the additional fuel cost coefficient due to valve point effect, P Gi is the power of the i-th microgrid generator, is the minimum power of the i-th microgrid generator, and N is the number of generators;
[0021] Comprehensive operation and maintenance cost C m for;
[0022]
[0023] Where, P pv,w is the output power of the w-th photovoltaic PV, P ess,h is the output power of the hth energy storage system ESS, P wt,vis the output power of the vth wind turbine WT, P mt,y is the output power of the yth micro fuel engine MT; c pv 、c ess 、c wt 、c mt are the operation and maintenance cost coefficients of PV, ESS, WT, and MT respectively, N pv is the number of photovoltaic cells, N ess is the number of energy storage systems, Nwt is the number of wind turbines, Nmt is the number of micro fuel engines;
[0024] Cost of curtailing wind and solar power C curt for:
[0025] C curt =c dis,t ΔP dis,t Δt Formula (6)
[0026] Where c dis,t is the penalty cost coefficient for wind and solar curtailment at time t, ΔP dis,t is the wind and solar power curtailment at time t;
[0027] Network loss cost C loss
[0028]
[0029] Where Ω is the branch set, P loss is the branch loss power, F i is the i-th branch.
[0030] Furthermore, the operating constraints of the active distribution network day-ahead dispatch model are as follows:
[0031] ADN power flow constraints:
[0032]
[0033] Where, Ω e is the set of nodes i, Ωs is the set of nodes k; branch ij indicates that the positive direction of the flow is from bus i to bus j; P ij,t , Q ij,t are the active and reactive power flows from node i to node j during period t; P jk,t , Q jk,t are the active and reactive power flows from node j to node k during period t; I ij,t 、V j,t are the current of branch ij and the voltage of bus j in time period t respectively; R ij 、X ij are the resistance and reactance of line ij respectively; Pin,j,t , Q in,j,t are the injected active and reactive power of node j in period t; Q svc,j,t is the reactive power output of the SVC at node j at time t, V i,t is the voltage of node i during period t;
[0034] Node voltage constraints: the voltage amplitude of each node in the distribution network must be maintained within the allowable range:
[0035]
[0036] Where, The minimum and maximum voltages of node i, plus the maximum and minimum current constraints
[0037] Branch current constraint
[0038]
[0039] Where, I ij is the branch current connecting nodes i and j, are the minimum and maximum current constraints of branch ij respectively
[0040] Reactive power compensation constraints, its adjustment range should meet the following requirements:
[0041]
[0042] Where, are the minimum and maximum reactive powers injected by the static VAR compensator SVC at node j at time t;
[0043] Diesel engine operating constraints:
[0044]
[0045] Where, P d (t-1), P d (t) are the output power of the diesel engine at time t-1 and time t, respectively, D d is the lower limit of diesel engine climbing grade, U d It is the upper limit of diesel engine climbing grade; are the minimum and maximum output power of the diesel engine respectively;
[0046] Microgrid power balance constraint, the operation of the microgrid should meet the power balance conditions, namely:
[0047]
[0048] Where, P L,j,t , Q L,j,t are respectively the active and reactive load demands of microgrid j in period t; P G,j,tis the active power output of the diesel engine of microgrid j during period t; P pv,j,t is the output active power of PV of microgrid j during period t; P wt,j,t is the output active power of WT of microgrid j during period t; P ess,j,t is the output active power of ESS of microgrid j in period t; P in,j,t , Q in,j,t are the system load active and reactive power of microgrid j in time period t, respectively;
[0049] Energy storage system constraints, energy storage charging and discharging power and state of charge constraints are:
[0050]
[0051] Where, They are the minimum and maximum values of the state of charge, respectively, where the SOC states are as follows:
[0052]
[0053] Where, SOC t , SOC t+1 are the state of charge at time t and t+1 respectively; η is the efficiency of ESS; P ess (t) is the output active power of the energy storage system ESS at time t; E cap is the rated capacity of the energy storage system.
[0054] Furthermore, the objective function and operating constraints of the intraday dispatch model of the microgrid group are as follows:
[0055] The target L of the intraday dispatch model of microgrid group total as follows:
[0056] L total =ζ1C mg +ζ2L mse Formula (16)
[0057] Where ζ1 and ζ2 are L mse and C mg The weight coefficient, L mse is the mean square error term of the microgrid controllable unit variables; the output of the controllable unit in the intraday scheduling stage is normalized, that is:
[0058]
[0059] Where N is the total number of controllable devices, ν i (t) is the actual control quantity of intraday scheduling, is the reference control quantity for day-ahead scheduling;
[0060] The microgrid operation constraints involved in the intraday dispatch model of the microgrid cluster are consistent with the day-ahead dispatch model of the active distribution network, including microgrid power balance constraints, equipment output range limitations, and energy storage system constraints. The constraints are shown in Formulas (11) to (14).
[0061] Furthermore, the specific definition of the day-ahead scheduling strategy in step S2 is as follows:
[0062] State space: the state space s of the cloud agent e Including historical load data of each microgrid, output of each microgrid and external electricity price, reactive output of active distribution network, namely:
[0063]
[0064] P cc is the power purchased from the upper level, P loss is the branch loss power, P pv,w is the output power of the w-th photovoltaic PV, P wt,v is the output power of the vth wind turbine WT, P ADN,MG,i is the active power transmitted from the active distribution network to the microgrid group i, P MG,ADN,i is the active power transmitted from microgrid group i to the active distribution network, is the time-of-use electricity price, U i is the voltage of node i, I ij is the branch current connecting nodes i and j, P L,j,t , Q L,j,t are the active and reactive load demands of microgrid j in period t respectively;
[0065] Action space: The action space of the cloud agent includes the reactive power change of the active distribution network and the output of the controllable units of each microgrid, namely:
[0066] a c ={ΔP pv,w ,ΔP wt,v ,ΔP ess,h ,ΔP mt,y ,ΔQ svc ,ΔP MG,ADN ,ΔP ADN,MG} Formula (24)
[0067] ΔP pv,w is the adjusted power of the w-th photovoltaic PV, ΔP wt,v is the adjusted power of the vth wind turbine WT, ΔP ess,h is the adjusted power of the hth energy storage system ESS, ΔP mt,y is the adjusted power of the yth micro fuel engine MT, ΔQ svc is the adjustment power of the static VAR compensator SVC, ΔPMG,ADN is the adjusted power transmitted from the microgrid group to the active distribution network, ΔP ADN,MG It is the adjusted power transmitted from the active distribution network to the microgrid group.
[0068] Reward function: While optimizing the active distribution network and microgrid target costs, P ADN,MG and P MG,ADN Therefore, a penalty function is added to the objective function of ADN to describe P ADN,MG,i and P MG,ADN,i At the same time, a voltage penalty term is introduced to maintain the stability of the node voltage. The reward function is designed as follows:
[0069]
[0070] Where k is the number of microgrids, is the power coordination penalty factor, is the voltage over-limit penalty factor, is the SOC state limit penalty factor, SOC t is the state of charge at time t, are the minimum and maximum values of the state of charge, respectively.
[0071] Furthermore, the overall interactive process of the intraday scheduling strategy in step S3 is: S→LSTM→π(S)→a t ; The specific definition of SAC strategy is as follows:
[0072] State space: E-agent state space s e Including the real-time operation status of the microgrid, day-ahead dispatch data, and distributed power sources, loads, and generator power in the area, as shown below:
[0073] s e ={P pv,w ,P wt,v ,P ADN,MG ,P MG,ADN ,P L,j,t} Formula (26)
[0074] Where, P pv,w is the active power output of the w-th photovoltaic PV, P wt,v is the output active power of the vth wind turbine WT, ΔP MG,ADN is the active power transmitted from the microgrid to the active distribution network, ΔP ADN,MG is the active power transmitted from the active distribution network to the microgrid group, P L,j,t is the active load demand of microgrid j at time t.
[0075] As the input of LSTM, it is used to model the state time series of the intraday scheduling process, extract key time series features, and encode the historical state results. With the current state Splicing, constructing enhanced state space s e , as shown below:
[0076]
[0077] Action space: the action space a of the edge agent e This includes power adjustment instructions for locally controllable devices, as shown in the following formula.
[0078] a e ={ΔP pv,w ,ΔP wt,v ,ΔP ess,h ,ΔP mt,y ,ΔP G} Formula (28)
[0079] Where ΔP pv,w is the adjusted power of the w-th photovoltaic PV, ΔP wt,v is the adjusted power of the vth wind turbine WT, ΔP ess,h is the adjusted power of the hth energy storage system ESS, ΔP mt,y w is the adjustment power of the yth micro fuel engine MT, ΔP G The adjusted power of the diesel engine;
[0080] Reward function, the reward function includes the microgrid operation cost and the mean square error of the controllable unit variables, that is:
[0081]
[0082] Where ζ1 and ζ2 are L mse and C mg The weight coefficient, L mse is the mean square error term of the microgrid controllable unit variables; is the SOC state limit penalty factor, SOC t is the state of charge at time t, are the minimum and maximum values of the state of charge, respectively.
[0083] Compared with the existing technology, the present invention has the following advantages: it establishes a two-stage edge-cloud collaborative scheduling framework for active distribution networks containing microgrid clusters based on deep reinforcement learning. It also establishes a two-stage deep reinforcement learning model for active distribution networks containing microgrid clusters based on the soft actor-critic (SAC) algorithm, and designs the state space, action space, and reward function of the deep reinforcement learning model. The SAC method of the present invention can more effectively coordinate the microgrids and controllable resources such as SVCs in the ADN, thereby reducing dependence on the external power grid.
[0084] In the intraday dispatch stage, in order to reduce the impact of microgrid intraday operation regulation deviation on ADN operation safety, the goal is to minimize the mean square error of the current microgrid operation cost and the day-ahead action deviation. A long short-term memory network (LSTM) is embedded in the actor strategy network front end of the SAC framework to extract the local state sequence of the microgrid and generate dispatch actions. This allows intraday dispatch to dynamically adjust the output of local microgrid controllable units while ensuring economic efficiency, maintain the continuity and stability of the dispatch process in both the day-ahead and intraday stages, and achieve a balance between cost and voltage safety. BRIEF DESCRIPTION OF THE DRAWINGS
[0085] Figure 1 Schematic diagram of the ADN structure including microgrid groups.
[0086] Figure 2 This is a flow chart of the two-stage scheduling strategy training proposed in the present invention.
[0087] Figure 3 It is an improved IEEE33 node system.
[0088] Figure 4 New energy,load diagrams for two scenarios.
[0089] Figure 5 The SAC scheduling results under the sunny scene. (a) is the active power output diagram, (b) is the reactive power output diagram, and (c) is the voltage fluctuation of each node.
[0090] Figure 6 The results of PSO scheduling under sunny weather scenarios are shown in Figure 1. (a) is the active power output diagram, (b) is the reactive power output diagram, and (c) is the voltage fluctuation diagram of each node.
[0091] Figure 7 This is the voltage diagram of each node at 18:00 on a sunny day.
[0092] Figure 8 This is the voltage node diagram at 10:30 on a cloudy day. DETAILED DESCRIPTION
[0093] The following will describe the embodiments of the present invention in detail with reference to the accompanying drawings and examples, so as to fully understand and implement the process of how the present invention applies technical means to solve technical problems and achieve technical effects.
[0094] like Figure 1 As shown in the figure, the ADN system structure with microgrid clusters consists of three microgrids MG1, MG2 and MG3. The microgrid includes wind turbines (WT), photovoltaics (PV), diesel engines (DE), microturbines (MT) and energy storage systems (ESS). In order to adjust the reactive balance of the ADN, a static var compensator (SVC) is added to the ADN. Based on the above model, a two-stage edge-cloud collaborative scheduling method for a campus distribution network with microgrid clusters is introduced in detail below. The specific steps are as follows:
[0095] Step S1: establishing an active distribution network dispatching model with the goal of minimizing the total operating cost of the active distribution network. The active distribution network dispatching model includes a day-ahead dispatching model of the active distribution network including a microgrid cluster and an intraday dispatching model of the microgrid cluster.
[0096] The day-ahead dispatch goal for an ADN with microgrid clusters is to achieve optimal operation of the ADN with microgrid clusters by optimizing the output of each microgrid and the controllable units in the distribution network through a cloud-based intelligent agent. This means minimizing the overall operating cost of the distribution network while satisfying operational constraints such as ADN voltage.
[0097]
[0098] Among them, C ADN is the total operating cost of the active distribution network, C pcc is the electricity purchase cost, C mg is the total operating cost of the microgrid, C loss is the network loss cost;
[0099] Power purchase cost C pcc : Cost of purchasing electricity from the active distribution network to the upper level
[0100]
[0101] Where T is the total scheduling time, is the time-of-use electricity price at time t, P cc (t) is the power purchased from the upper grid at time t, and Δt is the time interval;
[0102] The total operating cost of the microgrid C mg :
[0103]
[0104] Where C m is the comprehensive operation and maintenance cost of the microgrid, C curt is the cost of curtailing wind and solar power, C G is the power generation cost of the diesel engine, m mg is the number of microgrids;
[0105]
[0106] Where a i 、b i 、c i is the fuel cost coefficient of diesel engine power generation, λ i , ρ i is the additional fuel cost coefficient due to valve point effect, P Gi is the power of the i-th microgrid generator;
[0107] Comprehensive operation and maintenance cost C m for;
[0108]
[0109] Where, P pv,w 、P ess,h 、P wt,v 、P mt,y are the output power of PV, ESS, WT and MT respectively; c pv 、c ess 、c wt 、c mt are the operation and maintenance cost coefficients of PV, ESS, WT, and MT respectively;
[0110] Cost of curtailing wind and solar power C curt for:
[0111] C curt =c dis,t ΔP dis,t Δt Formula (6)
[0112] Where c dis,t is the penalty cost coefficient for wind and solar curtailment at time t, ΔP dis,t is the wind and solar power curtailment at time t;
[0113] Network loss cost C loss
[0114]
[0115] Where Ω is the branch set, P loss is the branch loss power, Fi is the i-th branch.
[0116] To ensure the safe and stable operation of the active distribution network, its operating characteristics and physical limitations must be considered (operational constraints include ADN constraints and microgrid constraints). The constraints are as follows.
[0117] ADN power flow constraints:
[0118]
[0119] Where, branch ij indicates that the positive direction of the flow is from busbar i to busbar j; P ij,t , Q ij,t are the active and reactive power flows from node i to node j during period t; I ij,t 、V j,t are the current of branch ij and the voltage of bus j in time period t respectively; R ij 、X ij are the resistance and reactance of line ij respectively; P in,j,t , Q in,j,t are the injected active and reactive power of node j in period t; Q svc,j,t is the reactive power output of the SVC at node j at time t, V i,t is the voltage of node i during period t;
[0120] Node voltage constraints: the voltage amplitude of each node in the distribution network must be maintained within the allowable range:
[0121]
[0122] Where, The minimum and maximum voltages of node i, plus the maximum and minimum current constraints
[0123] Branch current constraint
[0124]
[0125] Where, I ij is the branch current connecting nodes i and j, are the minimum and maximum current constraints of branch ij respectively
[0126] Reactive power compensation constraints, its adjustment range should meet the following requirements:
[0127]
[0128] Where, are the minimum and maximum reactive powers injected by the static VAR compensator SVC at node j at time t;
[0129] Diesel engine operating constraints:
[0130]
[0131] Where, P d is the output power of the diesel engine, D d is the lower limit of diesel engine climbing grade, U d It is the upper limit of diesel engine climbing grade; are the minimum and maximum output power of the diesel engine respectively;
[0132] Microgrid power balance constraint, the operation of the microgrid should meet the power balance conditions, namely:
[0133]
[0134] Where, P L,j,t , Q L,j,t are respectively the active and reactive load demands of microgrid j in period t; P G,j,t is the active power output of the diesel engine; P pv,j,t is the output active power of PV; P wt,j,t is the output active power of WT; P ess is the output active power of ESS; P in,j,t , Q in,j,t are the active and reactive power of the microgrid system load respectively;
[0135] Energy storage system constraints, energy storage charging and discharging power and state of charge constraints are:
[0136]
[0137] Where, They are the minimum and maximum values of the state of charge, respectively, where the SOC states are as follows:
[0138]
[0139] Where, SOC t , SOC t+1 are the state of charge at time t and t+1 respectively; η is the efficiency of ESS; E cap is the rated capacity of the energy storage system.
[0140] Intraday dispatch model of microgrid cluster
[0141] In order to reduce the voltage over-limit risk of distribution network operation during the day, an action consistency mechanism is introduced in the intraday dispatch of microgrids. Taking the day-ahead dispatch results as a priori, the intraday dispatch of microgrids is economically optimized. By adding the mean square error term between the intraday controllable resource variables and the day-ahead plan into the objective function, the intraday dispatch deviation of microgrids is suppressed, and the impact of microgrid nodes on the voltage fluctuation and operating cost of the ADN system is reduced. Therefore, the objective function established is as follows:
[0142] L total =ζ1C mg +ζ2L mse Formula (16)
[0143] Where ζ1 and ζ2 are L mse and C mg The weight coefficient, L mse is the mean square error term of the microgrid controllable unit variables; the output of the controllable unit in the intraday scheduling stage is normalized, that is:
[0144]
[0145] Where N is the total number of controllable devices, ν i (t) is the actual control quantity of intraday scheduling, It is the reference control quantity for day-ahead scheduling.
[0146] The microgrid operation constraints involved in the intraday dispatch model of the microgrid cluster are consistent with the day-ahead dispatch model of the active distribution network, including microgrid power balance constraints, equipment output range limitations, energy storage system constraints, etc. The constraints are shown in Formulas (11) to (14).
[0147] Step S2: A day-ahead dispatching strategy based on the soft policy gradient (SAC) algorithm is constructed. The cloud-based intelligent agent takes the minimum operating cost of the active distribution network as the goal, and uses the policy network to train the model of the input active distribution network status information. The cloud-based intelligent agent uses the execution result feedback to update the experience pool and optimize the SAC model.
[0148] Aiming at the problems of high-dimensional state space, continuous control variables and multi-source coupling in the two-stage edge-cloud scheduling of active distribution networks, the present invention designs a scheduling agent based on the actor-critic strategy structure.
[0149] The actor network controls the output of controllable units, and the critic network evaluates the performance of the actor strategy and guides optimization, thereby solving continuous decision-making problems in ADN multi-resource coordination scenarios. On this basis, the SAC algorithm introduces an entropy regularization mechanism to further optimize the actor strategy. The intelligent agent observes the current operating state s of the microgrid and distribution network, generates a dispatch action a to control various controllable units, and the distribution network feedbacks the immediate reward r and the next state s'. The four constitute the experience quadruple (s, a, r, s'), which is stored in the experience replay pool D for subsequent policy updates. When the number of caches in D is greater than the threshold, the network begins learning. The expected cumulative reward of the intelligent agent is expressed as:
[0150]
[0151] Where J is the objective function; ρπ is the state distribution under strategy π; E is the expected function of calculating the objective function J under strategy π; r(s t , a t ) is the reward function; s t is the state at time t; a t is the decision-making behavior at time t; α is the entropy regularization coefficient, H(π(·|s t )) is used to adjust the weight of the policy entropy H; π(·|s t ) represents the policy function in s t The probability distribution of the next action, the strategy update of the SAC algorithm is as follows:
[0152]
[0153] Where θ is the policy network parameter, Δ θ J(θ) is the gradient of the policy objective function; D is the experience replay pool, whose elements are (s t , a t , r t , s t+1 , dt), used for subsequent updates; E s~D is the expected return; Q1 and Q2 are dual Q networks, taking the minimum value to avoid overestimation; a' is the action generated by policy π. The Q value update strategy of the SAC algorithm is as follows:
[0154] L critic =E (s,a)~D [(Q(s,a)-y) 2 ] Formula (20)
[0155] Where, L critic is the critic loss function; y is the target Q value. The update strategy of the double Q network adopts soft update:
[0156] θ′ target =(1-τ)θ target +τθ Formula (21)
[0157] Where τ is the soft update coefficient. The algorithm convergence criterion is as follows:
[0158] |ΔJ(π)|<ε Formula (22)
[0159] Where, the strategy is considered to converge when ΔJ(π) is less than the convergence threshold ε.
[0160] In the day-ahead phase, the cloud-based intelligent agent adjusts the output of the microgrid cluster to achieve coordinated optimization at the ADN system level. The specific definitions are as follows:
[0161] (1) State space. The state space s of the cloud agent eIncluding historical load data of each microgrid, output of each microgrid and external electricity price, reactive output of active distribution network, etc., namely:
[0162]
[0163] Where, P ADN,MG,i is the active power transmitted from ADN to microgrid group i, P MG,ADN,i is the active power transmitted from microgrid cluster i to ADN;
[0164] (2) Action space. The action space of the cloud agent includes the reactive power change of the ADN and the output of the controllable units of each microgrid, that is:
[0165] a c ={ΔP pv,w ,ΔP wt,v ,ΔP ess,h ,ΔP mt,y ,ΔQ svc ,ΔP MG,ADN ,ΔP ADN,MG} Formula (24)
[0166] (3) Reward function. While optimizing the ADN and microgrid target costs, P ADN,MG and P MG,ADN Therefore, a penalty function is added to the objective function of ADN to describe P ADN,MG,i and P MG,ADN,i At the same time, a voltage penalty term is introduced to maintain the stability of the node voltage. The reward function is designed as follows:
[0167]
[0168] Where k is the number of microgrids, is the power coordination penalty factor, is the voltage over-limit penalty factor, It is the penalty factor for SOC state exceeding the limit.
[0169] Step S3: Construct an intraday dispatch strategy that combines the long short-term memory network and the soft policy gradient algorithm, with the goal of minimizing the mean square error of the current microgrid operating cost and the day-ahead action deviation. The long short-term memory network is embedded in the actor strategy network front end of the soft policy gradient algorithm framework to extract the local state sequence of the microgrid and generate dispatch actions.
[0170] The introduction of a mean square error term into the intraday scheduling objective function improves the intraday scheduling strategy's ability to track the day-ahead trend. Because the intraday scheduling cycle of this invention is 0.5 hours, the ADN exhibits stronger temporal correlation. Traditional SAC algorithms rely solely on the current real-time state of the microgrid and fail to consider the day-ahead scheduling strategy. This can lead to frequent fluctuations in the output of controllable units and create the risk of ADN voltage over-limits.
[0171] Therefore, we integrate reinforcement learning with the LSTM network with time-series perception to construct a SAC+LSTM edge agent to enhance its perception of microgrid disturbance trends. The latent feature vector output by LSTM processing is used as the input of the actor network, so that the actor network takes into account the current observed real-time state of the microgrid and the day-ahead dispatch state when generating scheduling actions, achieving consistency between intraday strategy output and day-ahead plan tracking. The overall interaction process is: S→LSTM→π(S)→a t The SAC strategy is designed as follows:
[0172] (1) State space. E-agent state space s e It includes the real-time operating status of the microgrid, day-ahead dispatch data, and the distributed power sources, loads, and generator power in the region, as shown in Equation (26).
[0173] s e ={P pv,w ,P wt,v ,P ADN,MG ,P MG,ADN ,P L,j,t} Formula (26)
[0174] As the input of LSTM, it can model the state time series of the intraday scheduling process and extract key time series features. With the current state Splicing, constructing enhanced state space s e , as shown in formula (27):
[0175]
[0176] (2) Action space. The action space of the edge agent a e Including power adjustment instructions for local controllable devices, as shown in formula (28).
[0177] a e ={ΔP pv,w ,ΔP wt,v ,ΔP ess,h ,ΔP mt,y ,ΔP G} Formula (28)
[0178] Wherein, each term is the power adjustment value relative to its current reference plan.
[0179] (3) Reward function. Considering the economic efficiency of microgrid operation and the consistency of day-ahead regulation, the reward function includes the microgrid operation cost and the mean square error of the controllable unit variables, namely:
[0180]
[0181] The two-stage dispatching training flow chart of the proposed active distribution network containing microgrids is shown in the figure below: Figure 2 It should be noted that the update steps of the policy network and the value network in this flowchart are all based on the expected return maximization function defined by formula (18).
[0182] During the day-ahead training phase, the cloud agent uses the policy network to train a model based on the input ADN state information, aiming to minimize ADN operating costs. The cloud agent uses execution results as feedback to update the experience pool and optimize the SAC model. During the intraday training phase, a time-enhanced DRL scheduling framework is constructed by combining SAC and LSTM. An LSTM module is embedded in the policy network input, encoding the temporal features of the microgrid's local state and actions. Training is also performed using historical microgrid operating data. During the execution phase, each edge agent controls the controllable resources of each microgrid cluster, thereby implementing a "cloud-based training, edge-based execution" control framework for ADNs involving microgrid clusters.
[0183] Take the improved IEEE33 node system as an example for verification. Figure 3 The system integrates three microgrids, and static VAR compensation devices are installed at the backbone nodes. The baseline capacity of each VAR is 400 kVAR. The operation and maintenance costs of each distributed power source are shown in Table 1, and the parameters of each unit are shown in Table 2.
[0184] In order to evaluate the performance of the proposed scheduling method in different environments, two typical scenarios, sunny and cloudy, are selected for analysis. The power distribution of renewable energy is as follows: Figure 4 shown.
[0185] Table 1 Parameters of each unit
[0186] Serial number <![CDATA[G1]]> <![CDATA[G2]]> <![CDATA[a / (¥ / kW 2 h)]]> 0.0004 0.0006 b / (¥ / kWh) 0.25 0.2 c / (¥ / h) 40 40 λ / (¥ / h) 300 200 ρ / (¥ / kWh) 0.035 0.042 <![CDATA[P min / kW]]> 200 100 <![CDATA[P max / kW]]> 2000 1500
[0187] Table 2 Parameters of each distributed power supply
[0188]
[0189] Analysis of the results of the previous stage
[0190] In order to verify the effectiveness of the proposed method, the traditional particle swarm optimization algorithm (PSO) and genetic algorithm (GA) are used for scheduling. The average voltage node fluctuation index is defined as follows:
[0191]
[0192] Where, is the average voltage of node i, γ i is the voltage node fluctuation rate, is the average voltage node fluctuation rate.
[0193] The results of different algorithms are shown in Table 3. The scheduling results of SAC and PSO under sunny scene are shown in Table 3. Figure 5 、 Figure 6 .
[0194] As shown in Table 3, under sunny conditions, the SAC method reduces total costs by 17.98% and 19.27% compared to the PSO and GA algorithms, respectively. Furthermore, the SAC method also achieves lower electricity purchase costs than the PSO and GA algorithms. Under cloudy conditions, the reduced output of renewable energy reduces the economic efficiency of ADN operations, resulting in a 17.17% increase in total costs compared to sunny conditions. However, these costs are still reduced by 13.80% and 13.30% compared to the PSO and GA methods, respectively. This demonstrates that the SAC method can more effectively coordinate controllable resources such as microgrids and SVCs within the ADN, thereby reducing dependence on the external grid.
[0195] Table 3. Results of running different algorithms
[0196]
[0197] In terms of voltage fluctuation control, under sunny weather scenarios, the average voltage fluctuation rate of ADN under SAC increased by 0.024% compared with PSO and decreased by 0.091% compared with GA. Both SAC and GA algorithms can maintain the overall node voltage within the limit by adjusting the reactive power and energy storage power of SVC. However, although PSO maintains a lower average voltage fluctuation rate, it is Figure 5 c and Figure 6 c shows that it performs insufficiently in handling the voltage limit problem of the ADN terminal nodes. For example, the voltage amplitudes of nodes 32 and 33 have reached 0.905 pu and 0.90 pu, causing the voltage to exceed the limit (the safety range is 0.95 pu to 1.05 pu).
[0198] Analysis of intraday stage scheduling results
[0199] Based on the day-ahead results, each microgrid is dispatched within the day. The results of the SAC algorithm and the SAC+LSTM algorithm are shown in Tables 4 and 5, respectively. The reward standard deviation in the table is used to measure the degree of fluctuation of the strategy output during the training process. The smaller the standard deviation, the more stable the strategy and the more consistent the dispatch behavior under different states. Two scenarios with large voltage fluctuations are analyzed respectively. The node voltage is as follows: Figure 7 and Figure 8 .
[0200] Table 4 SAC algorithm results
[0201]
[0202] Table 5 SAC+LSTM algorithm results
[0203]
[0204] It can be seen from Tables 4 and 5 that SAC has a greater advantage in terms of microgrid operation costs. For example, on sunny days in microgrid 2, the dispatching cost of SAC+LSTM increased by 15.37% compared with SAC. Analysis of the objective function values in Tables 4 and 5 shows that the SAC+LS TM method is lower in most scenarios. For example, the objective function value of microgrid 1 in scenario 1 decreased by 5.90%, and the reward standard deviation of SAC+LSTM during training was smaller. This shows that the SAC+LSTM strategy is more stable and the overall dispatching behavior is more consistent. This advantage is particularly prominent during periods when voltage disturbances are more obvious. Figure 7 At 18:00 on a sunny day, the system experienced significant voltage fluctuations. The voltage at the end node of SAC was approaching the safety critical value of 0.950pu, while SAC+LSTM was able to maintain the voltage in this area between 0.960 and 0.965pu, thus increasing the voltage margin and reducing the risk of voltage exceeding the limit. Figure 8 It can be seen that at 10:30 on a cloudy day, the voltage of the SAC at the middle and terminal nodes dropped to a minimum of 0.950pu, while the SAC+LSTM could stabilize above 0.957pu, preventing the system from entering the critical operation area.
[0205] Results show that while SAC achieves lower costs in some scenarios, its short-sighted optimization strategy may sacrifice voltage safety at local nodes. The proposed SAC+LSTM model leverages a temporal memory mechanism to enhance the modeling capability of microgrid state evolution trends. It uses the LSTM gating mechanism to filter historical microgrid state information, extract key feature vectors that reflect microgrid state change trends, and input them into the actor strategy network. This allows intraday scheduling to dynamically adjust the output of local microgrid controllable units while ensuring economic efficiency, maintaining the continuity and stability of the scheduling process both day-ahead and intraday, and achieving a balance between cost and voltage safety.
[0206] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in the present invention should be included in the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be based on the scope of protection of the claims.
Claims
1. A two-stage edge-cloud collaborative scheduling method for a campus distribution network containing microgrid clusters, characterized by: Follow the steps below; Step S1: establishing an active distribution network dispatching model with the goal of minimizing the total operating cost of the active distribution network. The active distribution network dispatching model includes a day-ahead dispatching model of the active distribution network including a microgrid cluster and an intraday dispatching model of the microgrid cluster. Step S2: Construct a day-ahead dispatch strategy based on the soft policy gradient (SAC) algorithm. The cloud agent takes minimizing the operating cost of the active distribution network as the goal. The cloud agent uses the policy network to train the model based on the input active distribution network status information. The cloud agent uses the execution results as feedback to update the experience pool and optimize the SAC model. Step S3: Construct an intraday dispatch strategy that combines the long short-term memory network and the soft policy gradient algorithm, with the goal of minimizing the mean square error of the current microgrid operating cost and the day-ahead action deviation. The long short-term memory network is embedded in the actor strategy network front end of the soft policy gradient algorithm framework to extract the local state sequence of the microgrid and generate dispatch actions.
2. The two-stage edge-cloud collaborative scheduling method for a campus distribution network containing a microgrid cluster according to claim 1 is characterized by: The objective function of the active distribution network day-ahead dispatch model in step S1 is as follows: Among them, C ADN is the total operating cost of the active distribution network, C pcc is the electricity purchase cost, C mg is the total operating cost of the microgrid, C loss is the network loss cost, T is the total dispatching time; Power purchase cost C pcc : Cost of purchasing electricity from the active distribution network to the upper level Where T is the total scheduling time, is the time-of-use electricity price at time t, P cc (t) is the power purchased from the upper grid at time t, and Δt is the time interval; The total operating cost of the microgrid C mg : Where C m is the comprehensive operation and maintenance cost of the microgrid, C curt is the cost of curtailing wind and solar power, C G is the power generation cost of the diesel engine, m mg is the number of microgrids; Where a i 、b i 、c i is the fuel cost coefficient of diesel engine power generation, λ i , ρ i is the additional fuel cost coefficient due to valve point effect, P Gi is the power of the i-th microgrid generator, is the minimum power of the i-th microgrid generator, and N is the number of generators; Comprehensive operation and maintenance cost C m for; Where, P pv,w is the output power of the w-th photovoltaic PV, P ess,h is the output power of the hth energy storage system ESS, P wt,v is the output power of the vth wind turbine WT, P mt,y is the output power of the yth micro fuel engine MT; c pv 、c ess 、c wt 、c mt are the operation and maintenance cost coefficients of PV, ESS, WT, and MT respectively, N pv is the number of photovoltaic cells, N ess is the number of energy storage systems, Nwt is the number of wind turbines, Nmt is the number of micro fuel engines; Cost of curtailing wind and solar power C curt for: C curt =c dis,t ΔP dis,t Δt Formula (6) Where c dis,t is the penalty cost coefficient for wind and solar curtailment at time t, ΔP dis,t is the wind and solar power curtailment at time t; Network loss cost C loss Where Ω is the branch set, P loss is the branch loss power, F i is the i-th branch.
3. The two-stage edge-cloud collaborative scheduling method for a campus distribution network containing a microgrid cluster according to claim 2 is characterized by: The operating constraints of the active distribution network day-ahead dispatch model are as follows: ADN power flow constraints: Where, Ω e is the set of nodes i, Ωs is the set of nodes k; branch ij indicates that the positive direction of the flow is from bus i to bus j; P ij,t , Q ij,t are the active and reactive power flows from node i to node j during period t; P jk,t , Q jk,t are the active and reactive power flows from node j to node k during period t; I ij,t 、V j,t are the current of branch ij and the voltage of bus j in time period t respectively; R ij 、X ij are the resistance and reactance of line ij respectively; P in,j,t , Q in,j,t are the injected active and reactive power of node j in period t; Q svc,j,t is the reactive power output of the SVC at node j at time t, V i,t is the voltage of node i during period t; Node voltage constraints: the voltage amplitude of each node in the distribution network must be maintained within the allowable range: Where, They are the minimum and maximum voltages of node i, plus the maximum and minimum current constraints and branch current constraints Where, I ij is the branch current connecting nodes i and j, They are the minimum and maximum current constraints of branch ij and the reactive compensation constraints, and their adjustment range should satisfy: Where, are the minimum and maximum reactive powers injected by the static VAR compensator SVC at node j at time t; Diesel engine operating constraints: Where, P d (t-1), P d (t) are the output power of the diesel engine at time t-1 and time t, respectively, D d is the lower limit of diesel engine climbing grade, U d It is the upper limit of diesel engine climbing grade; are the minimum and maximum output power of the diesel engine respectively; Microgrid power balance constraint, the operation of the microgrid should meet the power balance conditions, namely: Where, P L,j,t , Q L,j,t are respectively the active and reactive load demands of microgrid j in period t; P G,j,t is the active power output of the diesel engine of microgrid j during period t; P pv,j,t is the output active power of PV of microgrid j during period t; P wt,j,t is the output active power of WT of microgrid j during period t; P ess,j,t is the output active power of ESS of microgrid j in period t; P in,j,t , Q in,j,t are the system load active and reactive power of microgrid j in time period t, respectively; Energy storage system constraints, energy storage charging and discharging power and state of charge constraints are: Where, They are the minimum and maximum values of the state of charge, respectively, where the SOC states are as follows: Where, SOC t , SOC t+1 are the state of charge at time t and t+1 respectively; η is the efficiency of ESS; P ess (t) is the output active power of the energy storage system ESS at time t; E cap is the rated capacity of the energy storage system.
4. The two-stage edge-cloud collaborative scheduling method for a campus distribution network containing a microgrid cluster according to claim 3 is characterized by: The objective function and operating constraints of the intraday dispatch model of the microgrid group are as follows: The target L of the intraday dispatch model of microgrid group total as follows: L total =ζ1C mg +ζ2L mse Formula (16) Where ζ1 and ζ2 are L mse and C mg The weight coefficient, L mse is the mean square error term of the microgrid controllable unit variables; the output of the controllable unit in the intraday scheduling stage is normalized, that is: Where N is the total number of controllable devices, ν i (t) is the actual control quantity of intraday scheduling, is the reference control quantity for day-ahead scheduling; The microgrid operation constraints involved in the intraday dispatch model of the microgrid cluster are consistent with the day-ahead dispatch model of the active distribution network, including microgrid power balance constraints, equipment output range limitations, and energy storage system constraints. The constraints are shown in Formulas (11) to (14).
5. The two-stage edge-cloud collaborative scheduling method for a campus distribution network containing a microgrid cluster according to claim 4 is characterized in that: The specific definition of the day-ahead scheduling strategy in step S2 is as follows: State space: the state space s of the cloud agent e Including historical load data of each microgrid, output of each microgrid and external electricity price, reactive output of active distribution network, namely: P cc is the power purchased from the upper level, P loss is the branch loss power, P pv,w is the output power of the w-th photovoltaic PV, P wt, v is the output power of the vth wind turbine WT, P ADN,MG,i is the active power transmitted from the active distribution network to the microgrid group i, P MG,ADN,i is the active power transmitted from microgrid group i to the active distribution network, is the time-of-use electricity price, U i is the voltage of node i, I ij is the branch current connecting nodes i and j, P L,j,t , Q L,j,t are the active and reactive load demands of microgrid j in period t respectively; Action space: The action space of the cloud agent includes the reactive power change of the active distribution network and the output of the controllable units of each microgrid, namely: a c = {ΔP pv,w , ΔP wt,v , ΔP ess,h , ΔP mt,y , ΔQ svc , ΔP MG,ADN , ΔP ADN,MG}} Formula (24) ΔP pv,w is the adjusted power of the w-th photovoltaic PV, ΔP wt,v is the adjusted power of the vth wind turbine WT, ΔP ess, h is the adjustment power of the hth energy storage system ESS, ΔP mt,y is the adjusted power of the yth micro fuel engine MT, ΔQ svc is the adjustment power of the static VAR compensator SVC, ΔP MG,ADN is the adjusted power transmitted from the microgrid group to the active distribution network, ΔP ADN,MG It is the adjusted power transmitted from the active distribution network to the microgrid group. Reward function: While optimizing the active distribution network and microgrid target costs, P ADN,MG and P MG,ADN Therefore, a penalty function is added to the objective function of ADN to describe P ADN,MG,i and P MG,ADN,i At the same time, a voltage penalty term is introduced to maintain the stability of the node voltage. The reward function is designed as follows: Where k is the number of microgrids, is the power coordination penalty factor, is the voltage over-limit penalty factor, is the SOC state limit penalty factor, SOC t is the state of charge at time t, are the minimum and maximum values of the state of charge, respectively.
6. The two-stage edge-cloud collaborative scheduling method for a campus distribution network containing a microgrid cluster according to claim 5 is characterized by: The overall interactive process of the intraday scheduling strategy in step S3 is: S→LSTM→π(S)→a t ; The specific definition of SAC strategy is as follows: State space: E-agent state space s e Including the real-time operation status of the microgrid, day-ahead dispatch data, and distributed power sources, loads, and generator power in the area, as shown below: s e ={P pv,w ,P wt,v ,P ADN,MG ,P MG,ADN ,P L,j,t } Formula (26) Where, P pv,w is the active power output of the w-th photovoltaic PV, P wt,v is the output active power of the vth wind turbine WT, ΔP MG,ADN is the active power transmitted from the microgrid to the active distribution network, ΔP ADN,MG is the active power transmitted from the active distribution network to the microgrid group, P L,j,t is the active load demand of microgrid j at time t. As the input of LSTM, it is used to model the state time series of the intraday scheduling process, extract key time series features, and encode the historical state results. With the current state Splicing, constructing enhanced state space s e , as shown below: Action space: the action space a of the edge agent e This includes power adjustment instructions for locally controllable devices, as shown in the following formula. a e = {ΔP pv,w , ΔP wt,v , ΔP ess,h , ΔP mt,y , ΔP G}} Equation (28) Where ΔP pv,w is the adjusted power of the w-th photovoltaic PV, ΔP wt,v is the adjusted power of the vth wind turbine WT, ΔP ess,h is the adjusted power of the hth energy storage system ESS, ΔP mt,y w is the adjustment power of the yth micro fuel engine MT, ΔP G The adjusted power of the diesel engine; Reward function, the reward function includes the microgrid operation cost and the mean square error of the controllable unit variables, that is: Where ζ1 and ζ2 are L mse and C mg The weight coefficient, L mse is the mean square error term of the microgrid controllable unit variables; is the SOC state limit penalty factor, SOC t is the state of charge at time t, are the minimum and maximum values of the state of charge, respectively.