Hybrid microgrid system planning capacity configuration method and system based on two-layer optimization

Through a two-layer optimization model combined with deep Q network and mathematical programming, the multi-period capacity planning problem of the hybrid microgrid system is solved, the coordination of long-term and short-term capacity configuration is achieved, and the stability and economy of the system are improved.

CN119647996BActive Publication Date: 2025-10-03SHENYANG UNIVERSITY OF TECHNOLOGY +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411705350.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-26
Publication Date
2025-10-03
Estimated Expiration
2044-11-26

Smart Images

  • Figure CN119647996B_ABST
    Figure CN119647996B_ABST
Patent Text Reader

Abstract

A capacity configuration method and system for microgrid system planning based on double-layer optimization adopts a deep Q-network reinforcement learning method to determine the capacity configuration scheme for each decision cycle according to the state space and action space; the capacity configuration scheme for the long-term planning cycle is composed of the capacity configuration schemes of all decision cycles in the long-term planning cycle; a mathematical programming method is adopted to determine the capacity configuration scheme for each short-term planning cycle in the current decision cycle according to the constraint conditions of the target model and the upper limit of the configuration capacity of each short-term planning cycle in the current decision cycle; the economic index of the capacity configuration scheme for the long-term planning cycle is updated using the updated decision cycle reward function; when the updated economic index is less than the economic index before the update, the updated reward function of the implemented decision cycle is fed back to update the capacity configuration scheme for the long-term planning cycle, thereby solving the capacity configuration problem of microgrid system planning under multi-time scale uncertainty.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of microgrid system planning, and specifically relates to a microgrid system planning capacity configuration method and system based on a two-layer optimization model, which is used for two-layer optimization of multi-period capacity investment planning of a hybrid microgrid (HM). Background Art

[0002] Given that the share of renewable energy, primarily wind power, in the global energy mix is ​​expected to rise dramatically in the coming decades, a systematic approach is needed to plan long-term investment strategies.

[0003] Existing technologies cannot plan microgrid systems over long timescales, and the power trading market is subject to dynamic uncertainty. Rapid temporal fluctuations in renewable energy further complicate the problem, potentially leading to grid instability and unmet demand. Existing technologies combine energy storage devices and water electrolysis to produce hydrogen with energy management systems, but this further increases the complexity of the planning problem. Long-term capacity planning for hybrid microgrid systems can be formulated as a multi-period stochastic decision-making problem that accounts for uncertainties occurring on multiple timescales. Long-term capacity decisions are inherently linked to short-term energy dispatch and storage decisions, requiring simultaneous fulfillment. However, as the scale of hybrid microgrids increases, computational complexity increases. Existing technologies employ a two-stage stochastic programming (2SSP) model to address short-term uncertainties arising from weather and demand fluctuations. However, these models are short-term because they fail to account for the multi-period nature of capacity investment decisions and the impact of long-term uncertainty. Existing technologies also employ dynamic movement primitives (DMPs). While these DMPs account for the multi-period nature of capacity investment and evolving economic factors, they fail to account for the uncertainty inherent in these factors. Therefore, when the assumed trajectory of economic factors is far from the actual trajectory, the results of the DMP method may deviate significantly from the actual situation. Summary of the Invention

[0004] In order to address the deficiencies in the prior art, the present invention provides a hybrid microgrid system planning and capacity configuration method and system based on double-layer optimization, which solves the problem of microgrid system planning and capacity configuration under multi-time scale uncertainty, while meeting the requirements of stable energy management and maximizing economic efficiency.

[0005] The present invention adopts the following technical solutions.

[0006] The present invention proposes a capacity configuration method for microgrid system planning based on double-layer optimization. The microgrid system planning includes a long-term planning cycle and a short-term planning cycle. The facilities in the microgrid system include: wind turbines, energy storage equipment, and water electrolyzers; including:

[0007] Set the length of the decision cycle; a long-term planning cycle includes multiple decision cycles, and a decision cycle includes multiple short-term planning cycles;

[0008] The state space of a decision cycle is constructed using the allowable operating capacity of each facility within the decision cycle, and the action space of the decision cycle is constructed using the allowable expansion of the allowable operating capacity of each facility within the decision cycle. A deep Q-network reinforcement learning method is used to determine the capacity allocation plan for each decision cycle based on the state space and action space. The capacity allocation plan for the long-term planning cycle is constructed by combining the capacity allocation plans of all decision cycles within the long-term planning cycle.

[0009] The total allowable operating capacity of each facility in the current decision cycle is constructed by summing the state space of the previous decision cycle and the action space of the current decision cycle. The total allowable operating capacity of each facility in the current decision cycle is used as the upper limit of the capacity of each short-term planning cycle in the current decision cycle, and the maximum annual operating profit in the current decision cycle is used as the target model. A mathematical programming method is used to determine the capacity configuration plan for each short-term planning cycle in the current decision cycle according to the target model under the constraints of the target model and the upper limit of the configuration capacity of each short-term planning cycle in the current decision cycle.

[0010] After the capacity allocation plan of each short-term planning cycle is implemented, the actual power production value of each short-term planning cycle within a decision cycle is used to form the actual power production value of the decision cycle; the reward function of the decision cycle is updated according to the actual power production value of the decision cycle; the economic index of the capacity allocation plan of the long-term planning cycle is updated using the updated reward function of the decision cycle; when the updated economic index is less than the economic index before the update, the updated reward function of the decision cycle in which the capacity allocation plan has been implemented is fed back to update the capacity allocation plan of the long-term planning cycle.

[0011] Preferably, the length of the decision cycle is set to N years, and the long-term planning cycle includes W decision cycles, then the length of the long-term planning cycle is N×W years; a decision cycle includes m short-term planning cycles, and the length of the short-term planning cycle is W / m years.

[0012] Preferably, the state space of the decision cycle is constructed based on the allowed operating capacity of each facility within the decision cycle, satisfying the following relationship:

[0013] S P-1 =[s T,P-1 s B,P-1 sW,P-1 ] T (1)

[0014] Where S P-1 is the state space of the P-1th decision cycle, s T,P-1 、s B,P-1 、s W,P-1 They are the allowable operating capacity of the wind turbine, the allowable operating capacity of the energy storage device, and the allowable operating capacity of the water electrolyzer in the P-1th decision cycle.

[0015] Preferably, the action space of the decision cycle is constructed using the allowable expansion amount of the allowable operating capacity of each facility within the decision cycle, satisfying the following relationship:

[0016] A P =[ΔX T,P ΔX B,P ΔX W,P ] T (2)

[0017] Where A P is the action space of the Pth decision cycle, ΔX T,P , ΔX B,P , ΔX W,P They are respectively the allowable expansion amount of the allowable operating capacity of the wind turbine in the Pth decision cycle, the allowable expansion amount of the allowable operating capacity of the energy storage device, and the allowable expansion amount of the allowable operating capacity of the water electrolyzer.

[0018] Preferably, a deep Q-network reinforcement learning method is used to determine the capacity configuration scheme for each decision cycle based on the state space and action space, including:

[0019] 1) According to the state space S of the P-1th decision cycle P-1 and action space A P-1 Set the initial Q value Q of the facility in state s and action a (s,a) , where s∈S P-1 , a∈A P-1 ;

[0020] 2) The state space S in the P-1th decision cycle P-1 Next, select and execute the action space A of the Pth decision cycle P , to obtain the state space S of the Pth decision cycle P and reward function;

[0021] The reward function of the Pth decision cycle satisfies the following relationship:

[0022]

[0023] Where, is the reward function of the facility in state s and action a in the Pth decision cycle, C total,P is the total installation cost of each facility in the Pth decision cycle, R op,P,y is the total operating income of each facility in the yth year during the Pth decision cycle, r is the discount rate, and N is the number of years in a decision cycle;

[0024] 3) Update the Q value according to the deep Q network reinforcement learning rules to satisfy the following relationship:

[0025]

[0026] Where Q (s,a),P is the Q value in the Pth decision cycle, γ is the discount factor, satisfying The discount factor is used to discount the value of future rewards to the current value; P (s′|s,a) To perform action a∈A P Then from the current state s∈S P-1 Transfer to the updated state s′∈S P The probability of max a′ Q (s′,a′) To make Q in the updated state s′ (s′,a′) The action a′∈A that reaches the maximum value P ;

[0027] 4) In the updated state s′, repeat steps 2) and 3) to iteratively update the Q value until the predetermined number of iterations is reached;

[0028] 5) In the Pth decision cycle, the action space corresponding to the maximum Q value in the state space is used as the capacity configured in the decision cycle, satisfying the following relationship:

[0029] π(s)=max a Q (s,a) (6)

[0030] Where π(s) is the capacity configured in the decision cycle.

[0031] Preferably, the total allowed operating capacity of each facility in the current decision cycle is constructed by the sum of the state space of the previous decision cycle and the action space of the current decision cycle, satisfying the following relationship:

[0032]

[0033] Where, TX T,P TX B,P TX W,Pare the total allowable operating capacity of wind turbines, the total allowable operating capacity of energy storage devices, and the total allowable operating capacity of water electrolyzers in the Pth decision cycle respectively.

[0034] Preferably, the total allowed operating capacity of each facility in the current decision cycle is used as the upper limit of the capacity of each short-term planning cycle in the current decision cycle, and the maximum annual operating profit in the current decision cycle is used as the target model.

[0035] Satisfies the following relationship:

[0036] max∏ P,y =Rev P,y -C op,P,y -C main,P,y (8)

[0037] Where, π P,y is the annual operating profit based on electricity price and operating cost in the yth year of the Pth decision cycle, Rev P,y is the total annual income in the yth year of the Pth decision cycle, C op,P,y is the total annual operating cost in the yth year of the Pth decision cycle, C main,P,y is the total annual maintenance cost in year y during the P-th decision cycle;

[0038] The total annual revenue includes electricity sales and hydrogen sales, and satisfies the following relationship:

[0039]

[0040] Where EPrice is the market price of electricity, ESupply P,y,t is the amount of electricity supplied to the market in period t in year y during the Pth decision cycle, H Price is the market price of hydrogen, and H Pr od P,y,t is the hydrogen supplied to the market in time period t in year y during the Pth decision cycle, where T is the total number of time periods;

[0041] The total annual operating cost satisfies the following relationship:

[0042]

[0043] Where C penalty is the penalty cost for not meeting electricity demand, D P,y,t is the electricity demand in period t of year y in the Pth decision cycle, E sup ply,P,y,t is the amount of electricity supplied in time period t in year y during the Pth decision cycle, C carbon is the carbon tax cost per unit of electricity, E purchase,P,y,t is the amount of electricity purchased from the main grid in time period t in year y during the P-th decision cycle, where T is the total number of time periods;

[0044] The total annual maintenance cost satisfies the following relationship:

[0045] C main,P,y =(TX j,P ×IC j,P ×Ma int C ofef j ) (11)

[0046] Where, TX j,P is the total allowed operating capacity of facilities of type j in the Pth decision cycle, IC j,P is the unit investment cost of facility type j in the Pth decision cycle, MaintCoef j is the maintenance cost coefficient of facility type j.

[0047] Preferably, the constraints of the target model include:

[0048] 1) Constraints on the total operating capacity of each facility

[0049] The total permissible operating capacity of the wind turbine generator satisfies the following relationship:

[0050] TX T,P =ΔX T,P +TX T,P-1 (12)

[0051] Where, TX T,P is the total allowed operating capacity of wind turbines in the Pth decision cycle, ΔX T,P is the allowable expansion of the wind turbine operating capacity in the Pth decision cycle, TX T,P-1 is the total allowed operating capacity of wind turbines in the P-1th decision cycle;

[0052] The total allowable operating capacity of the energy storage device and the water electrolyzer satisfies the following relationship:

[0053] TX B,P +TX W,P =TX B,P-1 +TX W,P-1 +ΔX B,P +ΔX W,P (13)

[0054] Where ΔX B,P is the allowable expansion of the operating capacity of the energy storage device in the Pth decision cycle, ΔX W,P is the allowable expansion of the operating capacity of the water electrolyzer in the Pth decision cycle, TX B,P-1 TX W,P-1 are the total allowed operating capacities of the energy storage device and the water electrolyzer in the P-1th decision cycle, respectively;

[0055] 2) The energy balance constraint condition satisfies the following relationship:

[0056]

[0057] Where, GenPower P,t is the amount of electricity generated by the wind turbine in time period t during the Pth decision cycle, Eff grid ImportPower is the AC to AC conversion efficiency of the power grid. P,t is the amount of electricity purchased from the main grid during time period t in the Pth decision cycle, Discharge P,t Demand is the discharge amount of the energy storage device in time period t during the Pth decision cycle. P,t is the power demand in time period t during the Pth decision cycle, H2Prod new,P,t H2Prod is the amount of hydrogen produced by the newly installed water electrolyzer in the Pth decision cycle during time period t, existing,P,t is the amount of hydrogen produced by the previously installed water electrolyzer in time period t during the Pth decision cycle, Curtail P,t is the amount of power curtailed in the Pth decision cycle during time period t due to grid instability or unmet demand;

[0058] 3) The energy storage system constraints must satisfy the following relationship:

[0059]

[0060] InitialSorage P -FinalStorage P =0 (16)

[0061] Where, ESS P,y,t+1 is the energy storage system status at the end of time period t+1 in the yth year of the Pth decision cycle, ESS P,y,t is the energy storage system state at the beginning of time period t in the yth year of the Pth decision cycle, σ B is the self-discharge rate of the energy storage device, η AC / DC is the efficiency of the rectifier, η B is the discharge efficiency of the energy storage device, Charge P,y,t is the charging power in time period t in year y during the Pth decision cycle, Discharging P,y,t is the discharge power in time period t in year y during the Pth decision cycle, InitialStorage P and FinalStorage P are the energy storage device power at the beginning and end of the Pth decision cycle, respectively, and Δt is the time increment;

[0062] 4) Supply and demand constraints satisfy the following relationship:

[0063] sup P,y,t ≤D P,y,t (17)

[0064] In the formula, sup P,y,t is the power supply in time period t in year y during the Pth decision cycle, D P,y,t is the electricity demand in time period t in year y during the Pth decision cycle.

[0065] Preferably, the capacity allocation plan for the decision cycle includes a target value for power production of each facility within the decision cycle;

[0066] Capacity allocation plans for the long-term planning period, including target power production values ​​for the microgrid system;

[0067] Capacity allocation plan for the short-term planning period, including power production target values ​​and dispatch instructions for each facility.

[0068] Preferably, the economic performance index of the capacity allocation solution in the long-term planning period is updated using the following relationship using the reward function of the updated decision period:

[0069]

[0070] Where maxN PV is the economic index of the capacity configuration plan in the long-term planning period, To find the expected value of the function, P finish is the decision cycle number of the capacity allocation plan that has been implemented, P plan The decision cycle number in which the capacity allocation plan has not yet been implemented. The updated reward function for the decision cycle where the capacity allocation scheme has been implemented, The reward function for decision cycles where the capacity allocation plan has not yet been implemented.

[0071] The present invention also proposes a microgrid system planning capacity configuration system based on double-layer optimization. The microgrid system planning includes a long-term planning cycle and a short-term planning cycle. The facilities in the microgrid system include: wind turbines, energy storage equipment and water electrolyzers; including:

[0072] The cycle setting module is used to set the length of the decision cycle; a long-term planning cycle includes multiple decision cycles, and a decision cycle includes multiple short-term planning cycles;

[0073] The long-term planning scheme configuration module is used to construct the state space of the decision cycle based on the allowable operating capacity of each facility within the decision cycle, and to construct the action space of the decision cycle based on the allowable expansion of the allowable operating capacity of each facility within the decision cycle. The module uses a deep Q-network reinforcement learning method to determine the capacity configuration scheme for each decision cycle based on the state space and action space. The module then constructs the capacity configuration scheme for the long-term planning cycle based on the capacity configuration schemes of all decision cycles within the long-term planning cycle.

[0074] The short-term planning scheme configuration module is used to construct the total allowable operating capacity of each facility in the current decision cycle based on the sum of the state space of the previous decision cycle and the action space of the current decision cycle; the total allowable operating capacity of each facility in the current decision cycle is used as the upper limit of the capacity of each short-term planning cycle in the current decision cycle, and the maximum annual operating profit in the current decision cycle is used as the target model; mathematical programming methods are used to determine the capacity configuration scheme of each short-term planning cycle in the current decision cycle according to the target model under the constraints of the target model and the upper limit of the configuration capacity of each short-term planning cycle in the current decision cycle;

[0075] The long-term planning scheme update module is used to construct the actual value of power output of the decision cycle with the actual power output value of each short-term planning cycle within a decision cycle after the capacity allocation scheme of each short-term planning cycle is implemented; the reward function of the decision cycle is updated according to the actual power output value of the decision cycle; the economic index of the capacity allocation scheme of the long-term planning cycle is updated using the updated reward function of the decision cycle; when the updated economic index is less than the economic index before the update, the updated reward function of the decision cycle in which the capacity allocation scheme has been implemented is fed back to update the capacity allocation scheme of the long-term planning cycle.

[0076] Assume that the length of the decision cycle is N years, and the long-term planning cycle includes W decision cycles, then the length of the long-term planning cycle is N×W years; a decision cycle includes m short-term planning cycles, and the length of the short-term planning cycle is W / m years.

[0077] The state space of the decision cycle is constructed based on the allowed operating capacity of each facility within the decision cycle, satisfying the following relationship:

[0078] S P-1 =[s T,P-1 s B,P-1 s W,P-1 ] T (1)

[0079] Where S P-1 is the state space of the P-1th decision cycle, s T,P-1 、s B,P-1 、s W,P-1They are the allowable operating capacity of the wind turbine, the allowable operating capacity of the energy storage device, and the allowable operating capacity of the water electrolyzer in the P-1th decision cycle.

[0080] The action space of the decision cycle is constructed based on the allowable expansion of the allowable operating capacity of each facility within the decision cycle, satisfying the following relationship:

[0081] A P =[ΔX T,P ΔX B,P Δ XW,P ] T (2)

[0082] Where A P is the action space of the Pth decision cycle, ΔX T,P , ΔX B,P , ΔX W,P They are respectively the allowable expansion amount of the allowable operating capacity of the wind turbine in the Pth decision cycle, the allowable expansion amount of the allowable operating capacity of the energy storage device, and the allowable expansion amount of the allowable operating capacity of the water electrolyzer.

[0083] Using a deep Q-network reinforcement learning method, the capacity allocation scheme for each decision cycle is determined based on the state space and action space, including:

[0084] 1) According to the state space S of the P-1th decision cycle P-1 and action space A P-1 Set the initial Q value Q of the facility in state s and action a (s,a) , where s∈S P-1 , a∈A P-1 ;

[0085] 2) The state space S in the P-1th decision cycle P-1 Next, select and execute the action space A of the Pth decision cycle P , to obtain the state space S of the Pth decision cycle P and reward function;

[0086] The reward function of the Pth decision cycle satisfies the following relationship:

[0087]

[0088] Where, is the reward function of the facility in state s and action a in the Pth decision cycle, C total,P is the total installation cost of each facility in the Pth decision cycle, R op,P,y is the total operating income of each facility in the yth year during the Pth decision cycle, r is the discount rate, and N is the number of years in a decision cycle;

[0089] 3) Update the Q value according to the deep Q network reinforcement learning rules to satisfy the following relationship:

[0090]

[0091] Where Q (s,a),P is the Q value in the Pth decision cycle, γ is the discount factor, satisfying The discount factor is used to discount the value of future rewards to the current value; P (s′|s,a) To perform action a∈A P Then from the current state s∈S P-1 Transfer to the updated state s′∈S P The probability of max a′ Q (s′,a′) To make Q in the updated state s′ (s′,a′) The action a′∈A that reaches the maximum value P ;

[0092] 4) In the updated state s′, repeat steps 2) and 3) to iteratively update the Q value until the predetermined number of iterations is reached;

[0093] 5) In the Pth decision cycle, the action space corresponding to the maximum Q value in the state space is used as the capacity configured in the decision cycle, satisfying the following relationship:

[0094] π(s)=max a Q (s,a) (6)

[0095] Where π(s) is the capacity configured in the decision cycle.

[0096] The total allowed operating capacity of each facility in the current decision cycle is constructed by the sum of the state space of the previous decision cycle and the action space of the current decision cycle, satisfying the following relationship:

[0097]

[0098] Where, TX T,P TX B,P TX W,P are the total allowable operating capacity of wind turbines, the total allowable operating capacity of energy storage devices, and the total allowable operating capacity of water electrolyzers in the Pth decision cycle respectively.

[0099] The total allowed operating capacity of each facility in the current decision cycle is used as the upper limit of the capacity of each short-term planning cycle in the current decision cycle, and the maximum annual operating profit in the current decision cycle is used as the target model.

[0100] Satisfies the following relationship:

[0101] max∏ P,y =Re v P,y -C op,P,y -C main,P,y (8)

[0102] Where, π P,y is the annual operating profit based on electricity price and operating cost in the yth year of the Pth decision cycle, Rev P,y is the total annual income in the yth year of the Pth decision cycle, C op,P,y is the total annual operating cost in the yth year of the Pth decision cycle, C main,P,y is the total annual maintenance cost in year y during the P-th decision cycle;

[0103] The total annual revenue includes electricity sales and hydrogen sales, and satisfies the following relationship:

[0104]

[0105] Where, EPrice is the market price of electricity, ESupply P,y,t is the amount of electricity supplied to the market in period t in year y during the Pth decision cycle, H Price is the market price of hydrogen, and HProd P,y,t is the hydrogen supplied to the market in time period t in year y during the Pth decision cycle, where T is the total number of time periods;

[0106] The total annual operating cost satisfies the following relationship:

[0107]

[0108] Where C penalty is the penalty cost for not meeting electricity demand, D P,y,t is the electricity demand in period t in year t during the Pth decision cycle, E sup ply,P,y,t is the amount of electricity supplied in time period t in year y during the Pth decision cycle, C carbon is the carbon tax cost per unit of electricity, F purchase,P,y,t is the amount of electricity purchased from the main grid in time period t in year y during the P-th decision cycle, where T is the total number of time periods;

[0109] The total annual maintenance cost satisfies the following relationship:

[0110] C main,P,y =(TX j,P ×IC j,P ×Maint Coef j ) (11)

[0111] Where, TX j,Pis the total allowed operating capacity of facilities of type j in the Pth decision cycle, IC j,P is the unit investment cost of facility type j in the Pth decision cycle, MaintCoef j is the maintenance cost coefficient of facility type j.

[0112] The constraints of the target model include:

[0113] 1) Constraints on the total operating capacity of each facility

[0114] The total permissible operating capacity of the wind turbine generator satisfies the following relationship:

[0115] TX T,P =ΔX T,P +TX T,P-1 (12)

[0116] Where, TX T,P is the total allowed operating capacity of wind turbines in the Pth decision cycle, ΔX T,P is the allowable expansion of the wind turbine operating capacity in the Pth decision cycle, TX T,P-1 is the total allowed operating capacity of wind turbines in the P-1th decision cycle;

[0117] The total allowable operating capacity of the energy storage device and the water electrolyzer satisfies the following relationship:

[0118] TX B,P +TX W,P =TX B,P-1 +TX W,P-1 +ΔX B,P +ΔX W,P (13)

[0119] Where ΔX B,P is the allowable expansion of the operating capacity of the energy storage device in the Pth decision cycle, ΔX W,P is the allowable expansion of the operating capacity of the water electrolyzer in the Pth decision cycle, TX B,P-1 TX W,P-1 are the total allowed operating capacities of the energy storage device and the water electrolyzer in the P-1th decision cycle, respectively;

[0120] 2) The energy balance constraint condition satisfies the following relationship:

[0121]

[0122] Where, GenPower P,t is the amount of electricity generated by the wind turbine in time period t during the Pth decision cycle, Eff gridImportPower is the AC to AC conversion efficiency of the power grid. P,t is the amount of electricity purchased from the main grid during period t in the Pth decision cycle, Discharging P,t Demand is the discharge amount of the energy storage device in time period t during the Pth decision cycle. P,t is the power demand in time period t during the Pth decision cycle, H2Prod new,P,t H2Prod is the amount of hydrogen produced by the newly installed water electrolyzer in the Pth decision cycle during time period t, existing,P,t is the amount of hydrogen produced by the previously installed water electrolyzer in time period t during the Pth decision cycle, Curtail P,t is the amount of power curtailed in the Pth decision cycle during time period t due to grid instability or unmet demand;

[0123] 3) The energy storage system constraints must satisfy the following relationship:

[0124]

[0125] IiitilStorage P -FinalStorage P =0 (16)

[0126] Where, ESS P,y,t+1 is the energy storage system status at the end of time period t+1 in the yth year of the Pth decision cycle, ESS P,y,t is the energy storage system state at the beginning of time period t in the yth year of the Pth decision cycle, σ B is the self-discharge rate of the energy storage device, η AC / DC is the efficiency of the rectifier, η B is the discharge efficiency of the energy storage device, Charge P,y,t is the charging power in time period t in year y during the Pth decision cycle, Discharge P,y,t is the discharge power in time period t in year y during the Pth decision cycle, InitialStorage P and FinalStorage P are the energy storage device power at the beginning and end of the Pth decision cycle, respectively, and Δt is the time increment;

[0127] 4) Supply and demand constraints satisfy the following relationship:

[0128] sup P,y,t ≤D P,y,t (17)

[0129] In the formula, sup P,y,tis the power supply in time period t in year y during the Pth decision cycle, D P,y,t is the electricity demand in time period t in year y during the Pth decision cycle.

[0130] The capacity allocation plan for the decision cycle, including the target power production value for each facility during the decision cycle;

[0131] Capacity allocation plans for the long-term planning period, including target power production values ​​for the microgrid system;

[0132] Capacity allocation plan for the short-term planning period, including power production target values ​​and dispatch instructions for each facility.

[0133] Using the updated reward function of the decision cycle, the economic indicators of the capacity allocation plan in the long-term planning cycle are updated according to the following relationship:

[0134]

[0135] Where max N PV is the economic index of the capacity allocation plan in the long-term planning period, To find the expected value of the function, P finish is the decision cycle number of the capacity allocation plan that has been implemented, P plan The decision cycle number in which the capacity allocation plan has not yet been implemented. The updated reward function for the decision cycle where the capacity allocation scheme has been implemented, The reward function for decision cycles where the capacity allocation plan has not yet been implemented.

[0136] A terminal includes a processor and a storage medium; the storage medium is used to store instructions; the processor is used to operate according to the instructions to execute the steps of the method.

[0137] A computer-readable storage medium stores a computer program thereon, which implements the steps of the method when executed by a processor.

[0138] The present invention has the following advantages over existing technologies: The proposed coordination relationship between long-term and short-term capacity allocation is achieved through a two-layer optimization framework. This framework combines reinforcement learning (RL) methods to make upper-level decisions for long-term capacity allocation with mathematical programming (MP) methods to make lower-level decisions for short-term capacity allocation. The upper-level decision sets capacity constraints for the lower-level decision, and the lower-level decision is optimized based on the realized short-term uncertainty in each period. This coupling relationship ensures coordination between long-term and short-term capacity allocation.

[0139] The method proposed in the present invention realizes multi-period coupled optimization of hybrid microgrid system capacity planning, and can also be applied to other similar multi-time-scale optimization problems with random uncertainties. BRIEF DESCRIPTION OF THE DRAWINGS

[0140] Figure 1 This is a flow chart of a microgrid system planning capacity configuration method based on double-layer optimization proposed by the present invention. DETAILED DESCRIPTION

[0141] To make the objectives, technical solutions, and advantages of the present invention more clear, the technical solutions of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. The embodiments described in this application are only part of the embodiments of the present invention, not all of them. Based on the spirit of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0142] The present invention proposes a capacity configuration method for microgrid system planning based on double-layer optimization. The microgrid system planning includes a long-term planning cycle and a short-term planning cycle. The facilities in the microgrid system include: wind turbines, energy storage equipment and water electrolyzers; Figure 1 As shown, the method includes:

[0143] Step 1: Set the length of the decision cycle; a long-term planning cycle includes multiple decision cycles, and a decision cycle includes multiple short-term planning cycles.

[0144] The length of the decision cycle is set to N years. In the embodiment, the length of a decision cycle is 5 years. The length of the decision cycle is adjusted according to the system planning scale requirements. The long-term planning cycle includes W decision cycles. In the embodiment, W = 3, then the length of the long-term planning cycle is N × W = 15 years. A decision cycle includes m short-term planning cycles. In the embodiment, m = 5, then the length of the short-term planning cycle is W / m = 1 year.

[0145] Step 2: Construct the state space of the decision cycle with the allowable operating capacity of each facility within the decision cycle, and construct the action space of the decision cycle with the allowable expansion of the allowable operating capacity of each facility within the decision cycle; adopt the deep Q network reinforcement learning method to determine the capacity configuration plan of each decision cycle based on the state space and action space; and construct the capacity configuration plan of the long-term planning cycle with the capacity configuration plans of all decision cycles within the long-term planning cycle.

[0146] Capacity allocation plan for the decision cycle, including but not limited to the power production target value of each facility.

[0147] Capacity allocation plans for the long-term planning period, including but not limited to the power production target value of the microgrid system.

[0148] Specifically, step 2 includes:

[0149] Step 2.1, construct the state space of the decision cycle based on the allowed operating capacity of each facility in the decision cycle;

[0150] The state space represents the physical state and information state of the hybrid microgrid system. In any decision cycle, the following relationship is satisfied:

[0151] S P-1 =[s T,P-1 s B,P-1 s W,P-1 ] T (1)

[0152] Where S P-1 is the state space of the P-1th decision cycle, s T,P-1 、s B,P-1 、s W,P-1 They are the allowable operating capacity of the wind turbine, the allowable operating capacity of the energy storage device, and the allowable operating capacity of the water electrolyzer in the P-1th decision cycle.

[0153] The permitted operating capacity (AOC) refers to the total capacity available for generating or storing energy during a decision cycle, taking into account the remaining life of existing facilities and the capacity of newly installed facilities. The permitted operating capacity (AOC) determines the range of resources within which scheduling decisions can be made during a decision cycle.

[0154] Step 2.2, construct the action space of the decision cycle based on the allowed expansion of the allowed operating capacity of each facility in the decision cycle;

[0155] In any decision cycle, the action space of the decision cycle is constructed with the allowable expansion of the operating capacity of each facility, satisfying the following relationship:

[0156] A P =[ΔX T,P ΔX B,P ΔX W,P ] T (2)

[0157] Where A P is the action space of the Pth decision cycle, ΔX T,P , ΔX B,P , ΔX W,P They are respectively the allowable expansion amount of the allowable operating capacity of the wind turbine in the Pth decision cycle, the allowable expansion amount of the allowable operating capacity of the energy storage device, and the allowable expansion amount of the allowable operating capacity of the water electrolyzer.

[0158] The allowable expansion of operational capacity refers to the ability to expand existing energy facilities within a decision cycle based on current energy demand, technological development trends, economic factors, and environmental uncertainties. This allowable expansion is determined within a long-term planning framework by considering multi-timescale uncertainties and decision-making processes.

[0159] In step 2.3, a deep Q-network reinforcement learning method is used to determine the capacity allocation scheme for each decision cycle based on the state space and action space, including:

[0160] 1) According to the state space S of the P-1th decision cycle P-1 and action space A P-1 Set the initial Q value Q of the facility in state s and action a (s,a) , where s∈S P-1 , a∈A P-1 ;

[0161] 2) The state space S in the P-1th decision cycle P-1 Next, select and execute the action space A of the Pth decision cycle P , to obtain the state space S of the Pth decision cycle P and reward function;

[0162] The reward function is used to evaluate the long-term benefits of taking a specific action in a given state. It can simultaneously consider multiple objectives such as cost, benefit, and risk, achieving multi-objective optimization coupling. The reward function is introduced as a feedback value to dynamically adjust the Q value to find the optimal strategy.

[0163] The reward function of the Pth decision cycle satisfies the following relationship:

[0164]

[0165] Where, is the reward function of the facility in state s and action a in the Pth decision cycle, C total,P is the total installation cost of each facility in the Pth decision cycle, R op,P,y is the total operating income of each facility in the yth year during the Pth decision cycle, r is the discount rate, and N is the number of years in a decision cycle.

[0166] Among them, the total installation cost C of the facility in the Pth decision cycle is total,P Satisfies the following relationship:

[0167] C total,P =∑ j [C unit,j,P ×Q j,P ] (4)

[0168] Where C unit,j,P is the installation cost of facility type j in the Pth decision cycle, Q j,P is the number of facilities of type j installed in the Pth decision cycle, where j = T represents a wind turbine, j = B represents an energy storage device, and j = W represents a water electrolyzer.

[0169] 3) Update the Q value according to the deep Q network reinforcement learning rules to satisfy the following relationship:

[0170]

[0171] Where Q (s,a),P is the Q value in the Pth decision cycle, γ is the discount factor, satisfying The discount factor is used to discount the value of future rewards to the current value; P (s′|s,a) To perform action a∈A P Then from the current state s∈S P-1 Transfer to the updated state s′∈S P The probability of max a′ Q (s′,a′) To make Q in the updated state s′ (s′,a′) The action a′∈A that reaches the maximum value P .

[0172] The reward function is introduced as a feedback value in the Q-value update process, and the optimal Q-value is determined by coupling the configuration capacity with economic indicators.

[0173] 4) In the updated state s′, steps 2) and 3) are repeated to iteratively update the Q value until a predetermined number of iterations is reached.

[0174] 5) In the Pth decision cycle, the action space corresponding to the maximum Q value in the state space is used as the capacity configured in the decision cycle, satisfying the following relationship:

[0175] π(s)=max a Q (s,a) (6)

[0176] Where π(s) is the capacity configured in the decision cycle.

[0177] In step 2.4, the capacity allocation plan for the long-term planning period is formed by combining the capacity allocation plans for all decision cycles within the long-term planning period.

[0178] Step 3: Construct the total allowable operating capacity of each facility in the current decision cycle by taking the sum of the state space of the previous decision cycle and the action space of the current decision cycle; take the total allowable operating capacity of each facility in the current decision cycle as the upper limit of the capacity of each short-term planning cycle in the current decision cycle, and take the maximum annual operating profit in the current decision cycle as the target model; adopt a mathematical programming method, under the constraints of the target model and the upper limit of the configuration capacity of each short-term planning cycle in the current decision cycle, determine the capacity configuration plan of each short-term planning cycle in the current decision cycle according to the target model.

[0179] Capacity allocation plans for the short-term planning period, including but not limited to power production targets and dispatch instructions for each facility.

[0180] The present invention realizes an upper-layer optimization of the capacity of the long-term planning period configuration in step 2, and the capacity of the long-term planning period configuration obtained after the upper-layer optimization becomes the feasible space for the lower-layer optimization.

[0181] Specifically, step 3 includes:

[0182] In step 3.1, the sum of the state space of the P-1th decision cycle and the action space of the Pth decision cycle is used as the total allowable operating capacity of each facility in the Pth decision cycle, satisfying the following relationship:

[0183]

[0184] Where, TX T,P TX B,P TX W,P are the total allowable operating capacity of wind turbines, the total allowable operating capacity of energy storage devices, and the total allowable operating capacity of water electrolyzers in the Pth decision cycle respectively.

[0185] In step 3.2, the total allowed operating capacity of each facility in the Pth decision cycle is used as the upper limit of the capacity configuration in each short-term planning period in the Pth decision cycle, and the annual operating profit in the Pth decision cycle is maximized as the target model, satisfying the following relationship:

[0186] max∏ P,y =Rev P,y -C op,P,y -C main,P,y (8)

[0187] Where, π P,y is the annual operating profit based on electricity price and operating cost in the yth year of the Pth decision cycle, Rev P,y is the total annual income in the yth year of the Pth decision cycle, C op,P,y is the total annual operating cost in the yth year of the Pth decision cycle, C main,P,yis the total annual maintenance cost in year y during the P-th decision cycle.

[0188] The total annual revenue includes electricity sales and hydrogen sales, and satisfies the following relationship:

[0189]

[0190] Where, EPrice is the market price of electricity, ESupply P,y,t is the amount of electricity supplied to the market in period t in year y during the Pth decision cycle, H Price is the market price of hydrogen, and H Pr od P,y,t is the hydrogen supplied to the market in time period t in year y during the Pth decision cycle, where T is the total number of time periods;

[0191] The total annual operating cost satisfies the following relationship:

[0192]

[0193] Where C penalty is the penalty cost for not meeting electricity demand, D P,y,t is the electricity demand in period t of year y in the Pth decision cycle, E sup ply,P,y,t is the amount of electricity supplied in time period t in year y during the Pth decision cycle, C carbon is the carbon tax cost per unit of electricity, E purchase,P,y,t is the amount of electricity purchased from the main grid in time period t in year y during the P-th decision cycle, where T is the total number of time periods;

[0194] The total annual maintenance cost satisfies the following relationship:

[0195] C main,P,y =(TX j,P ×IC j,P ×MaintCoef j ) (11)

[0196] Where, TX j,P is the total allowed operating capacity of facilities of type j in the Pth decision cycle, IC j,P is the unit investment cost of facility type j in the Pth decision cycle, MaintCoef j is the maintenance cost coefficient of facility type j.

[0197] Step 3.3, establish the constraints of the target model;

[0198] Based on real-time data and forecast information, in order to ensure the balance between power supply and power demand, a set of constraints are required to complete the optimization problem of the target model. The constraints of the target model include:

[0199] 1) Constraints on the total operating capacity of each facility

[0200] The total permissible operating capacity of the wind turbine generator satisfies the following relationship:

[0201] TX T,P =ΔX T,P +TX T,P-1 (12)

[0202] Where, TX T,P is the total allowed operating capacity of wind turbines in the Pth decision cycle, ΔX T,P is the allowable expansion of the wind turbine operating capacity in the Pth decision cycle, TX T,P-1 is the total allowed operating capacity of wind turbines in the P-1th decision cycle.

[0203] The total allowable operating capacity of the energy storage device and the water electrolyzer satisfies the following relationship:

[0204] TX B,P +TX W,P =TX B,P-1 +TX W,P-1 +ΔX B,P +ΔX W,P (13)

[0205] Where ΔX B,P is the allowable expansion of the operating capacity of the energy storage device in the Pth decision cycle, ΔX W,P is the allowable expansion of the operating capacity of the water electrolyzer in the Pth decision cycle, TX B,P-1 TX W,P-1 are the total allowed operating capacities of the energy storage equipment and water electrolyzer in the P-1th decision cycle, respectively.

[0206] 2) The energy balance constraint condition satisfies the following relationship:

[0207]

[0208] Where, GenPower P,t is the amount of electricity generated by the wind turbine in time period t during the Pth decision cycle, Eff grid ImportPower is the AC to AC conversion efficiency of the power grid. P,t is the amount of electricity purchased from the main grid during period t in the Pth decision cycle, Discharging P,t Demand is the discharge amount of the energy storage device in time period t during the Pth decision cycle. P,t is the power demand in time period t during the Pth decision cycle, H2Prodnew,P,t H2Prod is the amount of hydrogen produced by the newly installed water electrolyzer in the Pth decision cycle during time period t, existing,P,t is the amount of hydrogen produced by the previously installed water electrolyzer in time period t during the Pth decision cycle, Curtail P,t is the amount of power cut in the Pth decision cycle during time period t due to grid instability or unmet demand.

[0209] Energy balance constraints ensure the balance between power supply and demand and are a key component in microgrid energy management systems (EMS).

[0210] 3) The energy storage system constraints must satisfy the following relationship:

[0211]

[0212] InitialStorage P -FinalStorage P =0 (16)

[0213] Where, ESS P,y,t+1 is the energy storage system status at the end of time period t+1 in the yth year of the Pth decision cycle, ESS P,y,t is the energy storage system state at the beginning of time period t in the yth year of the Pth decision cycle, σ B is the self-discharge rate of the energy storage device, η AC / DC is the efficiency of the rectifier, η B is the discharge efficiency of the energy storage device, Charge P,y,t is the charging power in time period t in year y during the Pth decision cycle, Disch arg ep ,y,t is the discharge power in time period t in year y during the Pth decision cycle, InitialStorage P and FinalStorage P are the energy storage device power at the beginning and end of the Pth decision cycle, respectively. Δt is the time increment, which represents the time step in the battery charge and discharge calculation, usually in hours.

[0214] 4) To prevent oversupply of a given demand, set supply and demand constraints to satisfy the following relationship:

[0215] sup P,y,t ≤D P,y,t (17)

[0216] In the formula, sup P,y,t is the power supply in time period t in year y during the Pth decision cycle, D P,y,tis the electricity demand in time period t in year y during the Pth decision cycle.

[0217] In step 3.4, a mathematical programming method is used to determine the configuration capacity of each short-term planning period in the current decision cycle according to the target model under the constraints of the target model and the upper limit of the configuration capacity of each short-term planning period in the current decision cycle.

[0218] Long-term capacity decisions are integrated with short-term energy scheduling and storage decisions to solve these problems at the same time, rather than separately. This integrated approach can more effectively optimize the performance of the overall system. Specifically, after the upper-level decision determines the long-term capacity configuration plan, the lower-level decision uses a mathematical programming (MP) method to formulate a short-term capacity configuration plan. In the embodiment, a short-term goal of maximizing annual operating profit is first formulated, followed by the calculation of the total annual revenue, total operating costs and total maintenance costs, and finally the annual operating profit is calculated based on the above revenue and costs, and the mathematical programming problem is solved to determine the optimal short-term configuration plan. The short-term planning cycle configures capacity, including but not limited to hourly scheduling decisions, such as wind turbine power generation and wind curtailment, charging and discharging of energy storage equipment, and whether the water electrolyzer starts electrolyzing water to produce hydrogen.

[0219] The current total allowed operating capacity is determined based on a long-term capacity allocation plan, which is derived through reinforcement learning and takes into account environmental uncertainties such as investment costs and the scale of electricity demand. The long-term capacity allocation plan involves the capacity allocation of each facility (such as wind turbines, energy storage equipment, and water electrolyzers), which is based on a comprehensive consideration of factors such as forecasts of future electricity demand growth and the remaining life of existing facilities. At the beginning of each decision cycle, the current total allowed operating capacity determined according to the long-term capacity allocation plan is the total capacity that can be used for power generation or storage during that cycle. The current total allowed operating capacity refers to the total capacity that can be used for power generation or storage during the short-term operating cycle, taking into account the remaining life of installed facilities and the capacity of newly installed facilities.

[0220] Initializing the short-term operation plan includes determining the current total allowed operating capacity, setting the initial state of the short-term operation cycle, formulating and implementing energy management strategies based on real-time data and forecast information during the short-term operation cycle, and then starting to implement the short-term operation plan. Performance is continuously monitored during the cycle to ensure the effective implementation of the strategy and make adjustments as needed. The execution results of short-term operations will be fed back into long-term planning to guide future long-term capacity configuration decisions, forming a dynamic optimization and adjustment process.

[0221] The present invention has a different logical order from the existing power grid planning ideas. Traditional power grid planning usually adopts a top-down approach, first directly determining the long-term planning capacity, and then decomposing it into short-term planning capacity, so that the short-term planning capacity is constrained by the long-term planning capacity. The present invention adopts a two-layer optimization framework, in which the capacity of each decision cycle configured in the upper decision stage (i.e., the long-term planning cycle) is determined by the reinforcement learning method based on the deep Q network and is expressed as the optimal state space and optimal action space within each decision cycle. The sum of the capacity configured in each decision cycle constitutes the capacity configured in the long-term planning cycle, and then the capacity configured in the long-term planning cycle is used as the capacity constraint of the lower decision stage (i.e., the short-term planning cycle). Therefore, the optimal state space and optimal action space in each decision cycle jointly determine the feasible space of the short-term planning cycle. The feasible space provides the boundaries of the physical state and information state for the short-term planning cycle. Within the feasible space, a mathematical programming method is used to optimize the capacity for the uncertainty factors in each short-term planning cycle, thereby determining the capacity configured in the short-term planning cycle.

[0222] In step 4, after the capacity allocation plan of each short-term planning cycle is implemented, the actual power output value of each short-term planning cycle within a decision cycle is used to form the actual power output value of the decision cycle; the reward function of the decision cycle is updated according to the actual power output value of the decision cycle; the economic index of the capacity allocation plan of the long-term planning cycle is updated using the updated reward function of the decision cycle; when the updated economic index is less than the economic index before the update, the updated reward function of the decision cycle in which the capacity allocation plan has been implemented is fed back to step 2 to update the capacity allocation plan of the long-term planning cycle.

[0223] Specifically, step 4 includes:

[0224] Step 4.1: After the capacity allocation plan of each short-term planning period is implemented, the actual power production value of each short-term planning period within a decision period is used to form the actual power production value of the decision period.

[0225] In step 4.2, the reward function of the decision cycle is updated according to the actual value of the electricity production in the decision cycle.

[0226] Specifically, the total operating income R of each facility in the decision cycle is calculated based on the actual value of power production in the Pth decision cycle. op,P,y , thereby updating the reward function of the Pth decision cycle to satisfy the following relationship:

[0227]

[0228] Where, is the reward function of the facility in state s and action a in the Pth decision cycle after the update, C total,Pis the total installation cost of each facility in the Pth decision cycle, is the updated total operating income of each facility in the Pth decision cycle in year y, r is the discount rate, and N is the number of years in a decision cycle.

[0229] In step 4.3, the economic performance indicator of the capacity allocation scheme in the long-term planning period is updated using the following relationship using the updated reward function of the decision cycle:

[0230]

[0231] Where max N PV is the economic index of the capacity allocation plan in the long-term planning period, To find the expected value of the function, P finish is the decision cycle number of the capacity allocation plan that has been implemented, P plan The decision cycle number in which the capacity allocation plan has not yet been implemented. The updated reward function for the decision cycle where the capacity allocation scheme has been implemented, The reward function for decision cycles where the capacity allocation plan has not yet been implemented.

[0232] The economic indicator represents the maximization of the net present value of the capacity allocation plan in the long-term planning cycle and is the average of the weighted sum of the reward functions of each decision cycle. The reward function of the decision cycle, updated according to the actual power output value of the decision cycle, not only represents the impact of the uncertainty and randomness of renewable energy output in the microgrid system, but also reflects the technical and economic impacts brought about by the scheduling decision. The reward function determined by the power output target value of the decision cycle is an important factor in Q-value optimization and is the key indicator of the capacity allocation plan in the decision cycle. The weighted sum of the two is a combination of the actual execution effect and the expected planning effect. Therefore, the economic indicator is updated promptly after the implementation of each decision cycle. The updated reward function of the decision cycle in which the capacity allocation plan has been implemented is used as the feedback value. The deep Q-network reinforcement learning method is again used to update the capacity allocation plan of each decision cycle based on the state space and action space, thereby updating the capacity allocation plan of the long-term planning cycle and achieving coordination and consistency between the capacity allocation plan of the long-term planning cycle and the capacity allocation plan of the short-term planning cycle.

[0233] In addition, the economic indicators are constructed by taking the average value of the weighted sum of the reward functions, making full use of the characteristic that the reward function is a comprehensive function of multiple objectives such as cost, benefit and risk. The realization of economic indicators can comprehensively consider different operating scenarios, which is conducive to improving the applicability of capacity configuration plans in long-term planning cycles.

[0234] In step 4.4, when the updated economic index is less than the economic index before the update, the updated reward function of the decision cycle in which the capacity allocation plan has been implemented is fed back to step 2 to update the capacity allocation plan for the long-term planning cycle.

[0235] In this invention, a reward function is used as a coupling variable between the long-term capacity allocation decisions implemented by upper-level optimization and the short-term operational decisions implemented by lower-level optimization. This allows feedback from short-term operations to guide and adjust long-term capacity allocation decisions. This coupling ensures coordination between long-term and short-term decisions. Secondly, the reward function is used in reinforcement learning to evaluate the long-term benefits of taking a specific action in a given state, which includes multiple objectives such as cost, benefit, and risk. This approach takes into account the evolving trends of technological, economic, and other uncertainties, providing a more dynamic and flexible decision-making process. Thirdly, the design of the reward function allows the model to make effective investment and operational decisions in the face of long-term and short-term uncertainties, such as technological development, price changes, and demand growth. This handling of uncertainty is relatively uncommon in the prior art. Finally, by updating the reward function, more consistent and predictable transitions are achieved across multiple time periods of different scales.

[0236] The present invention also proposes a microgrid system planning capacity configuration system based on double-layer optimization. The microgrid system planning includes a long-term planning cycle and a short-term planning cycle. The facilities in the microgrid system include: wind turbines, energy storage equipment and water electrolyzers; including:

[0237] The cycle setting module is used to set the length of the decision cycle; a long-term planning cycle includes multiple decision cycles, and a decision cycle includes multiple short-term planning cycles;

[0238] The long-term planning scheme configuration module is used to construct the state space of the decision cycle based on the allowable operating capacity of each facility within the decision cycle, and to construct the action space of the decision cycle based on the allowable expansion of the allowable operating capacity of each facility within the decision cycle. The module uses a deep Q-network reinforcement learning method to determine the capacity configuration scheme for each decision cycle based on the state space and action space. The module then constructs the capacity configuration scheme for the long-term planning cycle based on the capacity configuration schemes of all decision cycles within the long-term planning cycle.

[0239] The short-term planning scheme configuration module is used to construct the total allowable operating capacity of each facility in the current decision cycle based on the sum of the state space of the previous decision cycle and the action space of the current decision cycle; the total allowable operating capacity of each facility in the current decision cycle is used as the upper limit of the capacity of each short-term planning cycle in the current decision cycle, and the maximum annual operating profit in the current decision cycle is used as the target model; mathematical programming methods are used to determine the capacity configuration scheme of each short-term planning cycle in the current decision cycle according to the target model under the constraints of the target model and the upper limit of the configuration capacity of each short-term planning cycle in the current decision cycle;

[0240] The long-term planning scheme update module is used to construct the actual value of power output of the decision cycle with the actual power output value of each short-term planning cycle within a decision cycle after the capacity allocation scheme of each short-term planning cycle is implemented; the reward function of the decision cycle is updated according to the actual power output value of the decision cycle; the economic index of the capacity allocation scheme of the long-term planning cycle is updated using the updated reward function of the decision cycle; when the updated economic index is less than the economic index before the update, the updated reward function of the decision cycle in which the capacity allocation scheme has been implemented is fed back to update the capacity allocation scheme of the long-term planning cycle.

[0241] Assume that the length of the decision cycle is N years, and the long-term planning cycle includes W decision cycles, then the length of the long-term planning cycle is N×W years; a decision cycle includes m short-term planning cycles, and the length of the short-term planning cycle is W / m years.

[0242] The state space of the decision cycle is constructed based on the allowed operating capacity of each facility within the decision cycle, satisfying the following relationship:

[0243] S P-1 =[s T,P-1 s B,P-1 s W,P-1 ] T (1)

[0244] Where S P-1 is the state space of the P-1th decision cycle, s T,P-1 、s B,P-1 、s W,P-1 They are the allowable operating capacity of the wind turbine, the allowable operating capacity of the energy storage device, and the allowable operating capacity of the water electrolyzer in the P-1th decision cycle.

[0245] The action space of the decision cycle is constructed based on the allowable expansion of the allowable operating capacity of each facility within the decision cycle, satisfying the following relationship:

[0246] A P =[ΔX T,P ΔX B,P ΔXW,P ] T (2)

[0247] Where A P is the action space of the Pth decision cycle, ΔX T,P , ΔX B,P , ΔX W,P They are respectively the allowable expansion amount of the allowable operating capacity of the wind turbine in the Pth decision cycle, the allowable expansion amount of the allowable operating capacity of the energy storage device, and the allowable expansion amount of the allowable operating capacity of the water electrolyzer.

[0248] Using a deep Q-network reinforcement learning method, the capacity allocation scheme for each decision cycle is determined based on the state space and action space, including:

[0249] 1) According to the state space S of the P-1th decision cycle P-1 and action space A P-1 Set the initial Q value Q of the facility in state s and action a (s,a) , where s∈S P-1 , a∈A P-1 ;

[0250] 2) The state space S in the P-1th decision cycle P-1 Next, select and execute the action space A of the Pth decision cycle P , to obtain the state space S of the Pth decision cycle P and reward function;

[0251] The reward function of the Pth decision cycle satisfies the following relationship:

[0252]

[0253] Where, is the reward function of the facility in state s and action a in the Pth decision cycle, C total,P is the total installation cost of each facility in the Pth decision cycle, R op,P,y is the total operating income of each facility in the yth year during the Pth decision cycle, r is the discount rate, and N is the number of years in a decision cycle;

[0254] 3) Update the Q value according to the deep Q network reinforcement learning rules to satisfy the following relationship:

[0255]

[0256] Where Q (s,a),P is the Q value in the Pth decision cycle, γ is the discount factor, satisfying The discount factor is used to discount the value of future rewards to the current value; P (s′|s,a)To perform action a∈A P Then from the current state s∈S P-1 Transfer to the updated state s′∈S P The probability of max a′ Q (s′,a′) To make Q in the updated state s′ (s′,a′) The action a′∈A that reaches the maximum value P ;

[0257] 4) In the updated state s′, repeat steps 2) and 3) to iteratively update the Q value until the predetermined number of iterations is reached;

[0258] 5) In the Pth decision cycle, the action space corresponding to the maximum Q value in the state space is used as the capacity configured in the decision cycle, satisfying the following relationship:

[0259] π(s)=max a Q (s,a) (6)

[0260] Where π(s) is the capacity configured in the decision cycle.

[0261] The total allowed operating capacity of each facility in the current decision cycle is constructed by the sum of the state space of the previous decision cycle and the action space of the current decision cycle, satisfying the following relationship:

[0262]

[0263] Where, TX T,P TX B,P TX W,P are the total allowable operating capacity of wind turbines, the total allowable operating capacity of energy storage devices, and the total allowable operating capacity of water electrolyzers in the Pth decision cycle respectively.

[0264] The total allowed operating capacity of each facility in the current decision cycle is used as the upper limit of the capacity of each short-term planning cycle in the current decision cycle, and the maximum annual operating profit in the current decision cycle is used as the target model.

[0265] Satisfies the following relationship:

[0266] max∏ P,y =Re v P,y -C op,P,y -C main,P,y (8)

[0267] Where, π P,y is the annual operating profit based on electricity price and operating cost in the yth year of the Pth decision cycle, Rev P,y is the total annual income in the yth year of the Pth decision cycle, Cop,P,y is the total annual operating cost in the yth year of the Pth decision cycle, C main,P,y is the total annual maintenance cost in year y during the P-th decision cycle;

[0268] The total annual revenue includes electricity sales and hydrogen sales, and satisfies the following relationship:

[0269]

[0270] Where EPrice is the market price of electricity, ESupply P,y,t is the amount of electricity supplied to the market in period t in year y during the Pth decision cycle, H Price is the market price of hydrogen, and H Prod P,y,t is the hydrogen supplied to the market in time period t in year y during the Pth decision cycle, where T is the total number of time periods;

[0271] The total annual operating cost satisfies the following relationship:

[0272]

[0273] Where C penalty is the penalty cost for not meeting electricity demand, D P,y,t is the electricity demand in period t of year y in the Pth decision cycle, E sup ply,P,y,t is the amount of electricity supplied in time period t in year y during the Pth decision cycle, C carbon is the carbon tax cost per unit of electricity, E purchase,P,y,t is the amount of electricity purchased from the main grid in time period t in year y during the P-th decision cycle, where T is the total number of time periods;

[0274] The total annual maintenance cost satisfies the following relationship:

[0275] C main,P,y =(TX j,P ×IC j,P ×Maint Coef j ) (11)

[0276] Where, TX j,P is the total allowed operating capacity of facilities of type j in the Pth decision cycle, IC j,P is the unit investment cost of facility type j in the Pth decision cycle, MaintCoef j is the maintenance cost coefficient of facility type j.

[0277] The constraints of the target model include:

[0278] 1) Constraints on the total operating capacity of each facility

[0279] The total permissible operating capacity of the wind turbine generator satisfies the following relationship:

[0280] TX T,P =ΔX T,P +TX T,P-1 (12)

[0281] Where, TX T,P is the total allowed operating capacity of wind turbines in the Pth decision cycle, ΔX T,P is the allowable expansion of the wind turbine operating capacity in the Pth decision cycle, TX T,P-1 is the total allowed operating capacity of wind turbines in the P-1th decision cycle;

[0282] The total allowable operating capacity of the energy storage device and the water electrolyzer satisfies the following relationship:

[0283] TX B,P +TX W,P =TX B,P-1 +TX W,P-1 +ΔX B,P +ΔX W,P (13)

[0284] Where ΔX B,P is the allowable expansion of the operating capacity of the energy storage device in the Pth decision cycle, ΔX W,P is the allowable expansion of the operating capacity of the water electrolyzer in the Pth decision cycle, TX B,P-1 TX W,P-1 are the total allowed operating capacities of the energy storage device and the water electrolyzer in the P-1th decision cycle, respectively;

[0285] 2) The energy balance constraint condition satisfies the following relationship:

[0286]

[0287] Where, GenPower P,t is the amount of electricity generated by the wind turbine in time period t during the Pth decision cycle, Eff grid ImportPower is the AC to AC conversion efficiency of the power grid. P,t is the amount of electricity purchased from the main grid during period t in the Pth decision cycle, Discharging P,t Demand is the discharge amount of the energy storage device in time period t during the Pth decision cycle. P,t is the power demand in time period t during the Pth decision cycle, H2Prod new,P,t is the amount of hydrogen produced by the newly installed water electrolyzer in the Pth decision cycle during time period t, H2Pr od existiong,P,tis the amount of hydrogen produced by the previously installed water electrolyzer in time period t during the Pth decision cycle, Curtail P,t is the amount of power curtailed in the Pth decision cycle during time period t due to grid instability or unmet demand;

[0288] 3) The energy storage system constraints must satisfy the following relationship:

[0289]

[0290]

[0291] InitialStorage P -FinalStrorage P =0 (16)

[0292] Where, ESS P,y,t+1 is the energy storage system status at the end of time period t+1 in the yth year of the Pth decision cycle, ESS P,y,t is the energy storage system state at the beginning of time period t in the yth year of the Pth decision cycle, σ B is the self-discharge rate of the energy storage device, η AC / DC is the efficiency of the rectifier, η B is the discharge efficiency of the energy storage device, Charge P,y,t is the charging power in time period t in year y during the Pth decision cycle, Discharging P,y,t is the discharge power in time period t in year y during the Pth decision cycle, InitialStorage P and FinalStorage P are the energy storage device power at the beginning and end of the Pth decision cycle, respectively, and Δt is the time increment;

[0293] 4) Supply and demand constraints satisfy the following relationship:

[0294] sup P,y,t ≤D P,y,t (17)

[0295] In the formula, sup P,y,t is the power supply in time period t in year y during the Pth decision cycle, D P,y,t is the electricity demand in time period t in year y during the Pth decision cycle.

[0296] The capacity allocation plan for the decision cycle, including the target power production value for each facility during the decision cycle;

[0297] Capacity allocation plans for the long-term planning period, including target power production values ​​for the microgrid system;

[0298] Capacity allocation plan for the short-term planning period, including power production target values ​​and dispatch instructions for each facility.

[0299] Using the updated reward function of the decision cycle, the economic indicators of the capacity allocation plan in the long-term planning cycle are updated according to the following relationship:

[0300]

[0301] Where max N PV is the economic index of the capacity allocation plan in the long-term planning period, To find the expected value of the function, P finish is the decision cycle number of the capacity allocation plan that has been implemented, P plan The decision cycle number in which the capacity allocation plan has not yet been implemented. The updated reward function for the decision cycle where the capacity allocation scheme has been implemented, The reward function for decision cycles where the capacity allocation plan has not yet been implemented.

[0302] The technical feature of coupling lower-level decision-making to upper-level decision-making is that the lower-level decision-making formulates and executes energy management strategies based on real-time data and forecast information in each period to ensure a balance between power supply and demand. The feedback from these short-term operations is used to guide and adjust the upper-level long-term capacity configuration decisions, forming a dynamic optimization process. This coupling ensures coordination between long-term and short-term decisions, so that long-term investment plans not only meet the actual needs of short-term operations, but also can adapt to environmental changes and uncertainties, achieving more flexible and adaptable microgrid system planning. The present invention has a different logical order from the existing power grid planning ideas. Traditional power grid planning usually adopts a top-down approach, first determining long-term goals and then breaking them down into short-term actions. The method proposed in the present invention adopts a two-layer optimization, allowing long-term decisions to take into account feedback from short-term operations and adapt to environmental changes through reinforcement learning, achieving more flexible and adaptable planning.

[0303] The hybrid microgrid in this embodiment has installed fossil fuel-based energy sources. On shorter timescales, the energy management system (EMS) makes operational decisions to manage hourly fluctuations in wind power and electricity demand. The EMS provides electricity to meet demand, uses batteries for energy storage, and produces green hydrogen through water electrolyzers. Curtailment refers to the deliberate idling and excessive outages of wind turbines to avoid overload. In overload or underload situations, the EMS can implement curtailment or purchase electricity from the main grid at any time, subject to carbon emission penalties. While uncertainty and short-term fluctuations in demand and wind speed can interfere, the controller (EMS) attempts to maximize profits through green hydrogen sales after meeting electricity demand. For the hybrid microgrid system in this embodiment, a two-level optimization approach proposed in this invention addresses uncertainty at multiple timescales. At a higher level, strategic plans are formulated to maximize long-term profitability given the aggregated state, system description, and uncertainty beyond the current planning horizon. At a lower level, short-term operational decisions are made to ensure execution of the plan while considering various constraints. Capacity investment decisions are determined by reinforcement learning (RL) and are influenced by Markov processes that represent environmental uncertainties, such as investment costs and demand size. High-level decisions then set capacity constraints for more frequent low-level operational decisions in the corresponding period, which are formulated as mathematical programs (MPs) based on realized fast timescale uncertainties (such as wind speed and demand profiles). Net present value (NPV) is calculated through low-level simulations of various scenarios. Capacity investments in the hybrid microgrid system of the embodiment are made based on current realized information and future probability distributions of the stochastic and time-varying environment and the remaining life of previously installed facilities. Electricity demand growth and learning rate are considered as random parameters, while technological development is a time-varying parameter. The state consists of physical state variables that contain information about previously installed capacity and the remaining life of facilities. Other state variables contain current realized information of random environmental factors, such as electricity demand, size, and investment cost of each facility in the year of investment decision. Markov Decision Process Formulation for High-Level Decision Making This study proposes a long-term planning of a hybrid microgrid (HM) involving capacity investment decisions for each facility with P N-year-long decision intervals that can also serve as the basis for future scenarios due to the Markov nature of the uncertainties. Compared to the 2SSP method, the proposed method produces capacity deployment plans that enable more robust and cost-effective energy system planning decisions and achieve more consistent and predictable transitions. This improvement comes from the method's ability to consider evolving trends in techno-economic and other uncertainties in the decision-making process. The proposed method also outperforms the DMP method in enabling more robust and cost-effective energy system planning decisions. The strategic configuration prioritizes increased investments in energy storage equipment and investments in water electrolyzers to mitigate long-term uncertainties.The approach proposed in this invention also emphasizes the importance of reflecting the multi-periodic nature of capacity investment decisions, taking into account not only long-term changes but also the uncertainty of these changes in relevant economic factors, and using optimization models with sufficiently fine temporal resolution.

[0304] In addition, the method proposed in this invention adopts a two-layer framework, which provides constraints for short-term operational decisions through long-term capacity configuration, ensuring consistency between long-term goals and short-term actions. Through RL, the two-layer optimization framework can dynamically adapt to environmental changes and update decision-making strategies in real time, while 2SSP and DMP usually only make decisions once at each stage and lack dynamic adaptability. The introduction of RL enables the two-layer optimization framework to learn from experience and discover and utilize potential patterns and laws. When the external environment or policies change, the two-layer optimization framework can adjust its strategy more quickly and adapt to the new situation, while traditional 2SSP and DMP may require re-planning, which will lead to cumbersome computational problems.

[0305] In the embodiments, the system proposed in the present invention is applied to various occasions, including: hybrid microgrid (HM), which uses multiple energy sources and supervises energy infrastructure in a specific area through decentralized management; green hydrogen production, which uses renewable energy to produce hydrogen through a water electrolysis process; multi-time scale decision-making, which considers long-term capacity decisions and energy scheduling and storage decisions on short-term time scales; energy storage system (ESS) and water electrolysis hydrogen production technology to balance supply and demand and improve system reliability.

[0306] The present disclosure may be a system, method and / or computer program product. The computer program product may include a computer-readable storage medium carrying computer-readable program instructions for causing a processor to implement various aspects of the present disclosure.

[0307] A computer-readable storage medium can be a tangible device that can hold and store instructions for use by an instruction execution device. A computer-readable storage medium can be, for example, but not limited to, an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination thereof. More specific examples (a non-exhaustive list) of computer-readable storage media include: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disk (DVD), a memory stick, a floppy disk, a mechanical encoding device, such as a punch card or a raised structure in a groove on which instructions are stored, and any suitable combination thereof. As used herein, a computer-readable storage medium is not to be construed as a transient signal per se, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagating through a waveguide or other transmission medium (e.g., a light pulse through a fiber optic cable), or an electrical signal transmitted through an electrical wire.

[0308] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to each computing / processing device, or downloaded to an external computer or external storage device via a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network can include copper transmission cables, fiber optic transmission, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. The network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions to be stored in the computer-readable storage medium in each computing / processing device.

[0309] The computer program instructions for performing the operations of the present disclosure may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages ​​such as Smalltalk, C++, and conventional procedural programming languages ​​such as "C" language or similar programming languages. Computer-readable program instructions may be executed entirely on a user's computer, partially on a user's computer, as an independent software package, partially on a user's computer, partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., utilizing an Internet service provider to connect via the Internet). In some embodiments, an electronic circuit, such as a programmable logic circuit, a field programmable gate array (FPGA), or a programmable logic array (PLA), may be personalized by utilizing the state information of the computer-readable program instructions. The electronic circuit may execute the computer-readable program instructions, thereby realizing various aspects of the present disclosure.

[0310] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the above embodiments, ordinary technicians in the field should understand that the specific implementation methods of the present invention can still be modified or replaced by equivalents. Any modification or equivalent replacement that does not depart from the spirit and scope of the present invention should be covered by the scope of protection of the claims of the present invention.

Claims

1. A capacity configuration method for microgrid system planning based on two-layer optimization, wherein the microgrid system planning includes a long-term planning cycle and a short-term planning cycle, wherein: The facilities within the microgrid system include: wind turbines, energy storage equipment, and water electrolyzers; and are characterized by: Set the length of the decision cycle; a long-term planning cycle includes multiple decision cycles, and a decision cycle includes multiple short-term planning cycles; The state space of a decision cycle is constructed using the allowable operating capacity of each facility within the decision cycle, and the action space of the decision cycle is constructed using the allowable expansion of the allowable operating capacity of each facility within the decision cycle. A deep Q-network reinforcement learning method is used to determine the capacity allocation plan for each decision cycle based on the state space and action space. The capacity allocation plan for the long-term planning cycle is constructed by combining the capacity allocation plans of all decision cycles within the long-term planning cycle. The total allowable operating capacity of each facility in the current decision cycle is constructed by summing the state space of the previous decision cycle and the action space of the current decision cycle. The total allowable operating capacity of each facility in the current decision cycle is used as the upper limit of the capacity of each short-term planning cycle in the current decision cycle, and the maximum annual operating profit in the current decision cycle is used as the target model. A mathematical programming method is used to determine the capacity configuration plan for each short-term planning cycle in the current decision cycle according to the target model under the constraints of the target model and the upper limit of the configuration capacity of each short-term planning cycle in the current decision cycle. After the capacity allocation plan of each short-term planning cycle is implemented, the actual power production value of each short-term planning cycle within a decision cycle is used to form the actual power production value of the decision cycle; the reward function of the decision cycle is updated according to the actual power production value of the decision cycle; the economic index of the capacity allocation plan of the long-term planning cycle is updated using the updated reward function of the decision cycle; when the updated economic index is less than the economic index before the update, the updated reward function of the decision cycle in which the capacity allocation plan has been implemented is fed back to update the capacity allocation plan of the long-term planning cycle.

2. The microgrid system planning capacity configuration method based on double-layer optimization according to claim 1 is characterized in that: Assume that the length of the decision cycle is N years, and the long-term planning cycle includes W decision cycles, then the length of the long-term planning cycle is N×W years; a decision cycle includes m short-term planning cycles, and the length of the short-term planning cycle is W / m years.

3. The microgrid system planning capacity configuration method based on double-layer optimization according to claim 1 is characterized in that: The state space of the decision cycle is constructed based on the allowed operating capacity of each facility within the decision cycle, satisfying the following relationship: S P-1 =[s T,P-1 s B,P-1 s W,P-1 ] T (1) Where S P-1 is the state space of the P-1th decision cycle, s T,P-1 、s B,P-1 、s W,P-1 They are the allowable operating capacity of the wind turbine, the allowable operating capacity of the energy storage device, and the allowable operating capacity of the water electrolyzer in the P-1th decision cycle.

4. The microgrid system planning capacity configuration method based on double-layer optimization according to claim 3 is characterized in that: The action space of the decision cycle is constructed based on the allowable expansion of the allowable operating capacity of each facility within the decision cycle, satisfying the following relationship: A P =[ΔX T,P ΔX B,P ΔX W,P ] T (2) Where A P is the action space of the Pth decision cycle, ΔX T,P , ΔX B,P , ΔX W,P They are respectively the allowable expansion amount of the allowable operating capacity of the wind turbine in the Pth decision cycle, the allowable expansion amount of the allowable operating capacity of the energy storage device, and the allowable expansion amount of the allowable operating capacity of the water electrolyzer.

5. The microgrid system planning capacity configuration method based on double-layer optimization according to claim 4 is characterized in that: Using a deep Q-network reinforcement learning method, the capacity allocation scheme for each decision cycle is determined based on the state space and action space, including: 1) According to the state space S of the P-1th decision cycle P-1 and action space A P-1 Set the initial Q value Q of the facility in state s and action a (s,a) , where s∈S P-1 , a∈A P-1 ; 2) The state space S in the P-1th decision cycle P-1 Next, select and execute the action space A of the Pth decision cycle P , to obtain the state space S of the Pth decision cycle P and reward function; The reward function of the Pth decision cycle satisfies the following relationship: Where, is the reward function of the facility in state s and action a in the Pth decision cycle, C total,P is the total installation cost of each facility in the Pth decision cycle, R op,P,y is the total operating income of each facility in the yth year during the Pth decision cycle, r is the discount rate, and N is the number of years in a decision cycle; 3) Update the Q value according to the deep Q network reinforcement learning rules to satisfy the following relationship: Where Q (s,a),P is the Q value in the Pth decision cycle, γ is the discount factor, satisfying The discount factor is used to discount the value of future rewards to the current value; P (s′|s,a) To perform action a∈A P Then from the current state s∈S P-1 Transfer to the updated state s′∈S P The probability of max a′ Q (s′,a′) To make Q in the updated state s′ (s′,a′) The action a′∈A that reaches the maximum value P ; 4) In the updated state s′, repeat steps 2) and 3) to iteratively update the Q value until the predetermined number of iterations is reached; 5) In the Pth decision cycle, the action space corresponding to the maximum Q value in the state space is used as the capacity configured in the decision cycle, satisfying the following relationship: π(s)=max a Q (s,a) (6) Where π(s) is the capacity configured in the decision cycle.

6. The microgrid system planning capacity configuration method based on double-layer optimization according to claim 5 is characterized in that: The total allowed operating capacity of each facility in the current decision cycle is constructed by the sum of the state space of the previous decision cycle and the action space of the current decision cycle, satisfying the following relationship: Where, TX T,P TX B,P TX W,P are the total allowable operating capacity of wind turbines, the total allowable operating capacity of energy storage devices, and the total allowable operating capacity of water electrolyzers in the Pth decision cycle respectively.

7. The microgrid system planning capacity configuration method based on double-layer optimization according to claim 6 is characterized in that: The total allowed operating capacity of each facility in the current decision cycle is used as the upper limit of the capacity of each short-term planning cycle in the current decision cycle, and the maximum annual operating profit in the current decision cycle is used as the target model. Satisfies the following relationship: max∏ P,y =Re v P,y -C oP,P,y -C main,P,y (8) Where, π P,y is the annual operating profit based on electricity price and operating cost in the yth year of the Pth decision cycle, Rev P,y is the total annual income in the yth year of the Pth decision cycle, C oP,P,y is the total annual operating cost in the yth year of the Pth decision cycle, C main,P,y is the total annual maintenance cost in year y during the P-th decision cycle; The total annual revenue includes electricity sales and hydrogen sales, and satisfies the following relationship: Where EPrice is the market price of electricity, ESupply P,y,t is the amount of electricity supplied to the market in period t in year y during the Pth decision cycle, HPrice is the market price of hydrogen, and HProd P,y,t is the hydrogen supplied to the market in time period t in year y during the Pth decision cycle, where T is the total number of time periods; The total annual operating cost satisfies the following relationship: Where C penalty is the penalty cost for not meeting electricity demand, D P,y,t is the electricity demand in period t of year y in the Pth decision cycle, E supply,P,y,t is the amount of electricity supplied in time period t in year y during the Pth decision cycle, C carbon is the carbon tax cost per unit of electricity, E purchase,P,y,t is the amount of electricity purchased from the main grid in time period t in year y during the P-th decision cycle, where T is the total number of time periods; The total annual maintenance cost satisfies the following relationship: C main,P,y =(TX j,P ×IC j,P ×MaintCoef j ) (11) Where, TX j,P is the total allowed operating capacity of facilities of type j in the Pth decision cycle, IC j,P is the unit investment cost of facility type j in the Pth decision cycle, MaintCoef j is the maintenance cost coefficient of facility type j.

8. The microgrid system planning capacity configuration method based on double-layer optimization according to claim 7 is characterized in that: The constraints of the target model include: 1) Constraints on the total operating capacity of each facility The total permissible operating capacity of the wind turbine generator satisfies the following relationship: TX T,P =ΔX T,P +TX T,P-1 (12) Where, TX T,P is the total allowed operating capacity of wind turbines in the Pth decision cycle, ΔX T,P is the allowable expansion of the wind turbine operating capacity in the Pth decision cycle, TX T,P-1 is the total allowed operating capacity of wind turbines in the P-1th decision cycle; The total allowable operating capacity of the energy storage device and the water electrolyzer satisfies the following relationship: TX B,P +TX W,P =TX B,P-1 +TX W,P-1 +ΔX B,P +ΔX W,P (13) Where ΔX B,P is the allowable expansion of the operating capacity of the energy storage device in the Pth decision cycle, ΔX W,P is the allowable expansion of the operating capacity of the water electrolyzer in the Pth decision cycle, TX B,P-1 TX W,P-1 are the total allowed operating capacities of the energy storage device and the water electrolyzer in the P-1th decision cycle, respectively; 2) The energy balance constraint condition satisfies the following relationship: GenPower P,t ×Eff grid +ImportPower P,t +Discharge P,t =Demand P,t +H2Prod new,P,t +H2Prod existing,P,t +Curtail P,t (14) Where, GenPower P,t is the amount of electricity generated by the wind turbine in time period t during the Pth decision cycle, Eff grid ImportPower is the AC to AC conversion efficiency of the power grid. P,t is the amount of electricity purchased from the main grid during time period t in the Pth decision cycle, Discharge P,t Demand is the discharge amount of the energy storage device in time period t during the Pth decision cycle. P,t is the power demand in time period t during the Pth decision cycle, H2Prod new,P,t H2Prod is the amount of hydrogen produced by the newly installed water electrolyzer in the Pth decision cycle during time period t, existing,P,t is the amount of hydrogen produced by the previously installed water electrolyzer in time period t during the Pth decision cycle, Curtail P,t is the amount of power curtailed in the Pth decision cycle during time period t due to grid instability or unmet demand; 3) The energy storage system constraints must satisfy the following relationship: InitialStorage P -FinalStorage P =0 (16) Where, ESS P,y,t+1 is the energy storage system status at the end of time period t+1 in the yth year of the Pth decision cycle, ESS P,y,t is the energy storage system state at the beginning of time period t in the yth year of the Pth decision cycle, σ B is the self-discharge rate of the energy storage device, η AC / DC is the efficiency of the rectifier, η B is the discharge efficiency of the energy storage device, Charge P,y,t is the charging power in time period t in year y during the Pth decision cycle, Discharge P,y,t is the discharge power in time period t in year y during the Pth decision cycle, InitialStorage P and FinalStorage P are the energy storage device power at the beginning and end of the Pth decision cycle, respectively, and Δt is the time increment; 4) Supply and demand constraints satisfy the following relationship: sup P,y,t ≤D P,y,t (17) In the formula, sup P,y,t is the power supply in time period t in year y during the Pth decision cycle, D P,y,t is the electricity demand in time period t in year y during the Pth decision cycle.

9. The microgrid system planning capacity configuration method based on double-layer optimization according to claim 8 is characterized in that: The capacity allocation plan for the decision cycle, including the target power production value for each facility during the decision cycle; Capacity allocation plans for the long-term planning period, including target power production values ​​for the microgrid system; Capacity allocation plan for the short-term planning period, including power production target values ​​and dispatch instructions for each facility.

10. The microgrid system planning capacity configuration method based on double-layer optimization according to claim 9 is characterized in that: Using the updated reward function of the decision cycle, the economic indicators of the capacity allocation plan in the long-term planning cycle are updated according to the following relationship: In the formula, maxNPV is the economic index of the capacity allocation plan in the long-term planning period, To find the expected value of the function, P finish is the decision cycle number of the capacity allocation plan that has been implemented, P plan The decision cycle number in which the capacity allocation plan has not yet been implemented. The updated reward function for the decision cycle where the capacity allocation scheme has been implemented, The reward function for decision cycles where the capacity allocation plan has not yet been implemented.

11. A microgrid system planning capacity configuration system based on double-layer optimization, wherein the microgrid system planning includes a long-term planning cycle and a short-term planning cycle, wherein: The facilities within the microgrid system include: wind turbines, energy storage equipment, and water electrolyzers; and are characterized by: The cycle setting module is used to set the length of the decision cycle; a long-term planning cycle includes multiple decision cycles, and a decision cycle includes multiple short-term planning cycles; The long-term planning scheme configuration module is used to construct the state space of the decision cycle based on the allowable operating capacity of each facility within the decision cycle, and to construct the action space of the decision cycle based on the allowable expansion of the allowable operating capacity of each facility within the decision cycle. The module uses a deep Q-network reinforcement learning method to determine the capacity configuration scheme for each decision cycle based on the state space and action space. The module then constructs the capacity configuration scheme for the long-term planning cycle based on the capacity configuration schemes of all decision cycles within the long-term planning cycle. The short-term planning scheme configuration module is used to construct the total allowable operating capacity of each facility in the current decision cycle based on the sum of the state space of the previous decision cycle and the action space of the current decision cycle; the total allowable operating capacity of each facility in the current decision cycle is used as the upper limit of the capacity of each short-term planning cycle in the current decision cycle, and the maximum annual operating profit in the current decision cycle is used as the target model; mathematical programming methods are used to determine the capacity configuration scheme of each short-term planning cycle in the current decision cycle according to the target model under the constraints of the target model and the upper limit of the configuration capacity of each short-term planning cycle in the current decision cycle; The long-term planning scheme update module is used to construct the actual value of power output of the decision cycle with the actual power output value of each short-term planning cycle within a decision cycle after the capacity allocation scheme of each short-term planning cycle is implemented; the reward function of the decision cycle is updated according to the actual power output value of the decision cycle; the economic index of the capacity allocation scheme of the long-term planning cycle is updated using the updated reward function of the decision cycle; when the updated economic index is less than the economic index before the update, the updated reward function of the decision cycle in which the capacity allocation scheme has been implemented is fed back to update the capacity allocation scheme of the long-term planning cycle.

12. The microgrid system planning and capacity configuration system based on double-layer optimization according to claim 11, characterized in that: Assume that the length of the decision cycle is N years, and the long-term planning cycle includes W decision cycles, then the length of the long-term planning cycle is N×W years; a decision cycle includes m short-term planning cycles, and the length of the short-term planning cycle is W / m years.

13. The microgrid system planning and capacity configuration system based on double-layer optimization according to claim 11, characterized in that: The state space of the decision cycle is constructed based on the allowed operating capacity of each facility within the decision cycle, satisfying the following relationship: S P-1 =[s T,P-1 s B,P-1 s W,P-1 ] T (1) Where S P-1 is the state space of the P-1th decision cycle, s T,P-1 、s B,P-1 、s W,P-1 They are the allowable operating capacity of the wind turbine, the allowable operating capacity of the energy storage device, and the allowable operating capacity of the water electrolyzer in the P-1th decision cycle.

14. The microgrid system planning and capacity configuration system based on double-layer optimization according to claim 13, characterized in that: The action space of the decision cycle is constructed based on the allowable expansion of the allowable operating capacity of each facility within the decision cycle, satisfying the following relationship: A P =[ΔX T,P ΔX B,P ΔX W,P ] T (2) Where A P is the action space of the Pth decision cycle, ΔX T,P , ΔX B,P , ΔX W,P They are respectively the allowable expansion amount of the allowable operating capacity of the wind turbine in the Pth decision cycle, the allowable expansion amount of the allowable operating capacity of the energy storage device, and the allowable expansion amount of the allowable operating capacity of the water electrolyzer.

15. The microgrid system planning and capacity configuration system based on double-layer optimization according to claim 14, characterized in that: Using a deep Q-network reinforcement learning method, the capacity allocation scheme for each decision cycle is determined based on the state space and action space, including: 1) According to the state space S of the P-1th decision cycle P-1 and action space A P-1 Set the initial Q value Q of the facility in state s and action a (s,a) , where s∈S P-1 , a∈A P-1 ; 2) The state space S in the P-1th decision cycle P-1 Next, select and execute the action space A of the Pth decision cycle P , to obtain the state space S of the Pth decision cycle P and reward function; The reward function of the Pth decision cycle satisfies the following relationship: Where, is the reward function of the facility in state s and action a in the Pth decision cycle, C total,P is the total installation cost of each facility in the Pth decision cycle, R oP,P,y is the total operating income of each facility in the yth year during the Pth decision cycle, r is the discount rate, and N is the number of years in a decision cycle; 3) Update the Q value according to the deep Q network reinforcement learning rules to satisfy the following relationship: Where Q (s,a),P is the Q value in the Pth decision cycle, γ is the discount factor, satisfying The discount factor is used to discount the value of future rewards to the current value; P (s′|s,a) To perform action a∈A P Then from the current state s∈S P-1 Transfer to the updated state s′∈S P The probability of max a′ Q (s′,a′) To make Q in the updated state s′ (s′,a′) The action a′∈A that reaches the maximum value P ; 4) In the updated state s′, repeat steps 2) and 3) to iteratively update the Q value until the predetermined number of iterations is reached; 5) In the Pth decision cycle, the action space corresponding to the maximum Q value in the state space is used as the capacity configured in the decision cycle, satisfying the following relationship: π(s)=max a Q (s,a) (6) Where π(s) is the capacity configured in the decision cycle.

16. The microgrid system planning and capacity configuration system based on double-layer optimization according to claim 15, characterized in that: The total allowed operating capacity of each facility in the current decision cycle is constructed by the sum of the state space of the previous decision cycle and the action space of the current decision cycle, satisfying the following relationship: Where, TX T,P TX B,P TX W,P are the total allowable operating capacity of wind turbines, the total allowable operating capacity of energy storage devices, and the total allowable operating capacity of water electrolyzers in the Pth decision cycle respectively.

17. The microgrid system planning and capacity configuration system based on double-layer optimization according to claim 16, characterized in that: The total allowed operating capacity of each facility in the current decision cycle is used as the upper limit of the capacity of each short-term planning cycle in the current decision cycle, and the maximum annual operating profit in the current decision cycle is used as the target model. Satisfies the following relationship: max∏ P,y =Rev P,y -C oP,P,y -C main,P,y (8) Where, π P,y is the annual operating profit based on electricity price and operating cost in the yth year of the Pth decision cycle, Rev P,y is the total annual income in the yth year of the Pth decision cycle, C oP,P,y is the total annual operating cost in the yth year of the Pth decision cycle, C main,P,y is the total annual maintenance cost in year y during the P-th decision cycle; The total annual revenue includes electricity sales and hydrogen sales, and satisfies the following relationship: Where EPrice is the market price of electricity, ESupplyP,y,t is the amount of electricity supplied to the market in period t in year y during the Pth decision cycle, HPrice is the market price of hydrogen, and HProd P,y,t is the hydrogen supplied to the market in time period t in year y during the Pth decision cycle, where T is the total number of time periods; The total annual operating cost satisfies the following relationship: Where C penalty is the penalty cost for not meeting electricity demand, D P,y,t is the electricity demand in period t of year y in the Pth decision cycle, E supply,P,y,t is the amount of electricity supplied in time period t in year y during the Pth decision cycle, C carbon is the carbon tax cost per unit of electricity, E purchase,P,y,t is the amount of electricity purchased from the main grid in time period t in year y during the P-th decision cycle, where T is the total number of time periods; The total annual maintenance cost satisfies the following relationship: C main,P,y =(TX j,P ×IC j,P ×MaintCoef j ) (11) Where, TX j,P is the total allowed operating capacity of facilities of type j in the Pth decision cycle, IC j,P is the unit investment cost of facility type j in the Pth decision cycle, MaintCoef j is the maintenance cost coefficient of facility type j.

18. The microgrid system planning and capacity configuration system based on double-layer optimization according to claim 17, characterized in that: The constraints of the target model include: 1) Constraints on the total operating capacity of each facility The total permissible operating capacity of the wind turbine generator satisfies the following relationship: TX T,P =ΔX T,P +TX T,P-1 (12) Where, TX T,P is the total allowed operating capacity of wind turbines in the Pth decision cycle, ΔX T,P is the allowable expansion of the wind turbine operating capacity in the Pth decision cycle, TX T,P-1 is the total allowed operating capacity of wind turbines in the P-1th decision cycle; The total allowable operating capacity of the energy storage device and the water electrolyzer satisfies the following relationship: TX B,P +TX W,P =TX B,P-1 +TX W,P-1 +ΔX B,P +ΔX W,P (13) Where ΔX B,P is the allowable expansion of the operating capacity of the energy storage device in the Pth decision cycle, ΔX W,P is the allowable expansion of the operating capacity of the water electrolyzer in the Pth decision cycle, TX B,P-1 TX W,P-1 are the total allowed operating capacities of the energy storage device and the water electrolyzer in the P-1th decision cycle, respectively; 2) The energy balance constraint condition satisfies the following relationship: GenPower P,t ×Eff grid +ImportPower P,t +Discharge P,t =Demand P,t +H2Prod new,P,t +H2Prod existing,P,t +Curtail P,t (14) Where, GenPower P,t is the amount of electricity generated by the wind turbine in time period t during the Pth decision cycle, Eff grid ImportPower is the AC to AC conversion efficiency of the power grid. P,t is the amount of electricity purchased from the main grid during time period t in the Pth decision cycle, Discharge P,t Demand is the discharge amount of the energy storage device in time period t during the Pth decision cycle. P,t is the power demand in time period t during the Pth decision cycle, H2Prod new,P,t H2Prod is the amount of hydrogen produced by the newly installed water electrolyzer in the Pth decision cycle during time period t, existing,P,t is the amount of hydrogen produced by the previously installed water electrolyzer in time period t during the Pth decision cycle, Curtail P,t is the amount of power curtailed in the Pth decision cycle during time period t due to grid instability or unmet demand; 3) The energy storage system constraints must satisfy the following relationship: InitialStorage P -FinalStorage P =0 (16) Where, ESS P,y,t+1 is the energy storage system status at the end of time period t+1 in the yth year of the Pth decision cycle, ESS P,y,t is the energy storage system state at the beginning of time period t in the yth year of the Pth decision cycle, σ B is the self-discharge rate of the energy storage device, η AC / DC is the efficiency of the rectifier, η B is the discharge efficiency of the energy storage device, Charge P,y,t is the charging power in time period t in year y during the Pth decision cycle, Discharge P,y,t is the discharge power in time period t in year y during the Pth decision cycle, InitialStorage P and FinalStorage P are the energy storage device power at the beginning and end of the Pth decision cycle, respectively, and Δt is the time increment; 4) Supply and demand constraints satisfy the following relationship: sup P,y,t ≤D P,y,t (17) In the formula, sup P,y,t is the power supply in time period t in year y during the Pth decision cycle, D P,y,t is the electricity demand in time period t in year y during the Pth decision cycle.

19. The microgrid system planning and capacity configuration system based on double-layer optimization according to claim 18, characterized in that: The capacity allocation plan for the decision cycle, including the target power production value for each facility during the decision cycle; Capacity allocation plans for the long-term planning period, including target power production values ​​for the microgrid system; Capacity allocation plan for the short-term planning period, including power production target values ​​and dispatch instructions for each facility.

20. The microgrid system planning and capacity configuration system based on double-layer optimization according to claim 19, characterized in that: Using the updated reward function of the decision cycle, the economic indicators of the capacity allocation plan in the long-term planning cycle are updated according to the following relationship: In the formula, maxNPV is the economic index of the capacity allocation plan in the long-term planning period, To find the expected value of the function, P finish is the decision cycle number of the capacity allocation plan that has been implemented, P plan The decision cycle number in which the capacity allocation plan has not yet been implemented. The updated reward function for the decision cycle where the capacity allocation scheme has been implemented, The reward function for decision cycles where the capacity allocation plan has not yet been implemented.

21. A terminal comprising a processor and a storage medium; characterized in that: The storage medium is used to store instructions; The processor is configured to operate according to the instructions to execute the steps of the method according to any one of claims 1 to 10.

22. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the steps of the method according to any one of claims 1 to 10 are implemented.

Citation Information

Patent Citations

  • Integrated energy system energy management method under demand response

    CN116663820A

  • Enterprise portfolio analysis using finite state Markov decision process

    US20060195373A1