Microgrid joint optimization scheduling method and system integrating heterogeneous flexibility resources
By establishing a photovoltaic load model and energy storage model, combining Markov decision-making process and deep learning methods, the scheduling problem of heterogeneous flexible resources in the microgrid is solved, dynamic regulation across time periods and efficient utilization of resources is achieved, and the economic benefits of the microgrid are improved.
Patent Information
- Application Number
- CN202510805109.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-17
- Publication Date
- 2025-08-12
- Estimated Expiration
- 2045-06-17
AI Technical Summary
The existing microgrid optimization scheduling methods are difficult to effectively deal with the uncertainty and difference in adjustable capabilities of heterogeneous flexible resources, resulting in the lack of dynamic adjustment capabilities across time periods and multiple cycles, and it is difficult to select appropriate optimization models to describe user response behavior.
Establish a photovoltaic load model, a recent energy storage rental model, an intraday virtual energy storage model and an intraday shared energy storage model, and combine the Markov decision-making process and an asynchronous semi-coupled optimization method of dual-delay depth deterministic strategy gradient to realize the optimization scheduling of microgrid operators.
It improves the economic efficiency of microgrid operators, improves the dynamic regulation capability of flexible loads, solves the uncertainty and differences of flexible resources, and realizes the full utilization of flexible resources and maximizes benefits.
Smart Images

Figure CN120341858B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of power system automation and relates to microgrid optimization scheduling technology, and specifically to a microgrid joint optimization scheduling method and system integrating heterogeneous flexibility resources. Background Art
[0002] With the acceleration of global energy transformation and the pursuit of dual carbon goals, the power system is facing a series of challenges. Traditional centralized power grids, due to limitations such as one-way energy flow, long-distance transmission losses, and insufficient flexibility, struggle to adapt to the large-scale integration of distributed energy resources (such as photovoltaic and wind power) and the new demands for user-side "source-load-storage" synergy. Microgrids, as small power systems integrating distributed generation, energy storage, power electronics, and loads, have become a key technological path for building new power systems and promoting low-carbon energy development due to their flexible networking, seamless off-grid and on-grid switching, and localized energy consumption.
[0003] In traditional microgrids, end users of electricity are typically viewed as passive consumers. Microgrid operators can directly determine operating strategies based on their total energy demand, typically aiming to minimize the cost of purchasing energy from the utility grid. With the rapid development of distributed energy resources, traditional energy consumers are gradually evolving into prosumers. Prosumers perform both the production and consumption roles of energy. By deploying large numbers of distributed photovoltaic systems, they can not only achieve their own energy supply but also supply excess power to the grid. Therefore, microgrid operators can not only purchase electricity from the external grid but also purchase excess distributed energy from prosumers. This operating model has accelerated the global transition to a low-carbon, distributed energy structure. However, load demand and photovoltaic output are susceptible to environmental factors, resulting in power uncertainty. Furthermore, external grid electricity prices fluctuate randomly with changes in the electricity market, posing challenges to the economical and efficient operation of microgrids.
[0004] With the accelerated deployment of demand-side flexibility resources (such as distributed energy storage, electric vehicle clusters, and smart temperature control devices), these resources are becoming an important means of effectively addressing fluctuations in external grid electricity prices and renewable energy output. However, the large number of flexible resources in microgrids and their diverse operating characteristics complicate optimal microgrid scheduling. Furthermore, the uncertainty of prosumer behavior can affect the participation of flexible resources in microgrid regulation, posing challenges to optimal microgrid operation.
[0005] Existing research primarily focuses on optimizing the scheduling of distributed energy storage or flexible loads, and currently, optimizing the scheduling of flexible loads is primarily achieved through demand response technology. Furthermore, in terms of solving optimization algorithms, existing research primarily relies on deterministic methods based on mathematical programming and search methods based on heuristic rules. Therefore, existing technologies have the following specific shortcomings:
[0006] 1. Most studies focus on the independent optimization of a single type of flexibility resource (such as distributed energy storage or flexible loads), lacking a system architecture for the dynamic coupling mechanism between different types of resources. As a result, the potential for multi-resource synergy has not been fully explored.
[0007] 2. Current research on flexible load optimization scheduling is mainly achieved through demand response technology. However, this regulation strategy is mostly limited to fixed time periods (such as only allowing load reduction at the current moment or shifting part of the load to adjacent time periods) and lacks dynamic regulation capabilities across time periods and multiple cycles.
[0008] 3. Since users are affected by multiple factors such as economic incentives and user habits, and the shared energy storage capacity and demand response resources among different users vary greatly, existing research has difficulty in selecting a suitable optimization model to accurately describe the response behavior of each user, which brings challenges to solving microgrid resource optimization. Summary of the Invention
[0009] Purpose of the invention: In order to solve the problem that existing optimization methods are difficult to deal with the uncertainty and differences in the adjustable capabilities of heterogeneous flexibility resources, a microgrid joint optimization scheduling method and system that integrates heterogeneous flexibility resources is provided, which can stimulate the adjustment potential of multiple types of flexibility resources and enhance the dynamic adjustment capability of flexible loads.
[0010] Technical Solution: To achieve the above objectives, the present invention provides a microgrid joint optimization scheduling method integrating heterogeneous flexibility resources, comprising the following steps:
[0011] S1: Establishing photovoltaic load model;
[0012] S2: Establish a day-ahead energy storage rental model;
[0013] S3: Establish a prosumer optimization model based on the PV load model and the day-ahead energy storage rental model;
[0014] S4: Establishing a virtual energy storage model within a day;
[0015] S5: Establish a daily shared energy storage model;
[0016] S6: Establish a microgrid operator optimization model based on the prosumer optimization model, the intraday virtual energy storage model, and the intraday shared energy storage model;
[0017] S7: Convert the microgrid operator optimization model into a well-established Markov decision process;
[0018] S8: Solve the Markov decision process of step S7, update and iterate the scheduling parameters to achieve optimal scheduling of the microgrid.
[0019] Furthermore, the establishment of the photovoltaic load model in step S1 includes:
[0020] The mathematical expressions for PV power and load requirements are as follows:
[0021]
[0022]
[0023] in, 、 are the PV panel power and load power of prosumer i in time t, respectively; 、 are the diffusion coefficients of PV panel power and load power, and are the Brownian motion processes used to describe the randomness of PV panels and loads, respectively.
[0024] Furthermore, the establishment of the day-ahead energy storage rental model in step S2 includes:
[0025] The day-ahead storage rental model includes the benefits and costs of renting storage capacity by prosumers;
[0026] The mathematical expression of the benefits of renting energy storage capacity by prosumers is as follows:
[0027]
[0028] in, is the energy storage capacity rented by prosumer i at time t, Is the highest rental price, is the trend of the rental price curve, is the integration variable;
[0029] The mathematical expression for the cost of renting energy storage capacity by prosumers is as follows:
[0030]
[0031] in, 、 and are the coefficients of the quadratic term, linear term, and constant term of the leasing energy storage cost of the i-th prosumer.
[0032] Furthermore, the establishment of the prosumer optimization model in step S3 includes:
[0033] The goal of prosumers in leasing idle energy storage capacity is to maximize their net profit. The mathematical expression of the optimization model for prosumers to lease energy storage capacity is as follows:
[0034]
[0035] in, Rent energy storage capacity for optimized prosumers, It is the maximum energy storage capacity rented by the prosumer.
[0036] Furthermore, the establishment of the intraday virtual energy storage model in step S4 includes:
[0037] The dispatchable load of the prosumer includes the transferable load and the curtailable load. For the constant temperature control load, the mathematical expression of its power consumption is as follows:
[0038]
[0039] in, is the power consumption of the thermostatically controlled load i, is the baseline power consumption of thermostatically controlled load i, is the regulating power of the constant temperature control load i participating in the response; the thermodynamic kinetic equation of the constant temperature control load is as follows:
[0040]
[0041] in, is the internal heat capacity of building i, and are the equivalent heat gain and thermal resistance of building i, is the indoor temperature of building i; is the coefficient of constant temperature control load i; is the set temperature of the thermostatically controlled load i; is the diffusion coefficient of the flexible load; is a Wiener process with random disturbances caused by prosumer behavior or environmental influences;
[0042] Define the auxiliary variable as and , and reconstruct the thermodynamic kinetic equations into a virtual energy storage model, whose mathematical expression is as follows:
[0043]
[0044] in, represents the remaining capacity of virtual energy storage unit i; represents the energy loss rate of virtual energy storage unit i; and Represent the minimum and maximum remaining capacity respectively; and Represent the minimum and maximum power constraints respectively; the initial remaining capacity Defined as ; The remaining capacity after termination is [ , ];
[0045] The virtual energy storage model includes the compensation cost of the microgrid operator, and its mathematical expression is as follows:
[0046]
[0047] in, and are the compensation coefficients of the first constant term and the second constant term respectively, and N is the number of constant temperature control loads.
[0048] Furthermore, the establishment of the intraday shared energy storage model in step S5 includes:
[0049] Microgrid operators use the leased energy storage capacity to store or release electricity to shared energy storage during the day-ahead phase. The mathematical expression of the shared energy storage model is as follows:
[0050]
[0051] in, yes The derivative of It is the surplus energy storage capacity leased by the microgrid operator to the prosumer. represents the energy storage power, and is the minimum and maximum power of energy storage, represents the energy dissipation rate of stored energy, It represents the efficiency of energy storage charging and discharging power. Its mathematical expression is as follows:
[0052]
[0053] in, and are the efficiency of charging and discharging power respectively.
[0054] Furthermore, the establishment of the microgrid operator optimization model in step S6 includes:
[0055] The mathematical expression for the cost / revenue gained from the exchange of electricity between the operator and the external grid area is as follows:
[0056]
[0057] in, is the real-time electricity price of the external grid, is the interaction power between the operator and the external grid;
[0058] The mathematical expression of the cost / revenue that operators receive by purchasing or selling electricity to prosumers is as follows:
[0059]
[0060] in, is the price at which the operator sells energy to the prosumer, is the interaction power between operators and prosumers, is the coefficient; I function is the indicator function;
[0061] The microgrid meets the power balance constraints during operation in each time slot:
[0062]
[0063] in, is the load demand of prosumer i;
[0064] The mathematical expression of the optimization model of the microgrid operator is as follows:
[0065]
[0066] The objective function consists of four parts: the benefits / costs generated by the electricity interaction between operators and prosumers , the benefits / costs of power interaction between operators and external grids , the cost of leasing energy storage capacity and compensation costs to motivate prosumers .
[0067] Furthermore, the Markov decision process in step S7 includes four parts: state space, action space, state transition probability, and reward function, which are as follows:
[0068] State space: The system state is divided into the day-ahead state and intraday stage status ; However, the ES capacity status in the day-ahead phase Unable to obtain in advance, its corresponding status is unobservable; in the intraday stage, the state is observable, and its mathematical expression is as follows:
[0069]
[0070] The intraday status information includes the load demand of prosumers, distributed photovoltaics, time-of-use electricity prices, real-time electricity prices, shared energy storage capacity, and virtual energy storage capacity.
[0071] Action Space: Operator actions are divided into day-ahead phase actions and intraday phase actions In the day-ahead phase, operators acquire shared energy storage capacity by developing leasing strategies. In the intraday phase, operators earn revenue by managing shared energy storage and virtual energy storage. The action space is defined as:
[0072]
[0073] in, and The value of is used to determine the day-ahead leasing strategy. They are the shared energy storage and virtual energy storage power dispatched by the operator during the day;
[0074] Transition probability function: The transition probability is a dynamic variable that the agent gradually estimates through interactive exploration with the environment.
[0075] Reward function: The system’s reward function is divided into the day-ahead reward and intraday rewards , the reward function is defined as:
[0076]
[0077] in, is the reward coefficient;
[0078] According to the above Markov decision process formula, the discounted cumulative reward within time slot t is defined as:
[0079]
[0080] in, Indicates time slot Rewards before or during the day, is the discount factor; the goal of the Markov decision process is to maximize the expected cumulative discounted return:
[0081]
[0082] in, is described as a mapping function that transforms the state Mapping to Actions , is the mathematical expectation; and Represents the status and action of the day before or within the day respectively.
[0083] Furthermore, in step S8, an asynchronous semi-coupled optimization method based on double-delayed deep deterministic policy gradient is used to solve the Markov decision process, as follows:
[0084] The training process for both the intraday and day-ahead phases is performed via an actor-critic process with double-delayed deep deterministic policy gradients, where the actor network learns a policy π and generates actions under a given state, and the critic network evaluates the learning quality of the actor network under policy π, i.e., the objective of the Markov decision process, which is to maximize the expected discounted return.
[0085] The critic network completes policy evaluation by constructing an action-value function, which quantifies the expected return of taking a specific action in a given state; the action-value function is defined as:
[0086]
[0087] in, are the current network parameters, and the value of Q is iteratively updated by utilizing the Bellman equation:
[0088]
[0089] in, is the target Q value, are the target network parameters, and are the next state and action respectively; is the reward for performing action a in state s; is the expected value;
[0090] Iteratively update the Q value through the double Q learning method:
[0091]
[0092] The parameters of the critic network are obtained by minimizing the loss function To update; the actor network learns to optimize the policy π by maximizing the expected discounted reward, and by using the policy gradient To update the parameters of the actor network; the policy gradient is defined as follows:
[0093] in, is the gradient of action a; is the network parameter gradient; is the network parameter The strategy for the next state S;
[0094] Based on the update iteration of the parameters of the above optimization method, the optimal scheduling problem of microgrid flexibility resources is solved.
[0095] The present invention also provides a microgrid joint optimization dispatching system integrating heterogeneous flexibility resources, comprising:
[0096] Model building module, used to build photovoltaic load model, day-ahead energy storage rental model, prosumer optimization model, intraday virtual energy storage model, intraday shared energy storage model and microgrid operator optimization model;
[0097] A decision process establishment module is used to establish a Markov decision process and convert the microgrid operator optimization model into the established Markov decision process;
[0098] Solving module, used to solve Markov decision processes.
[0099] Beneficial effects: Compared with the prior art, the present invention has the following advantages:
[0100] 1. An innovative two-stage microgrid demand-side resource scheduling strategy, day-ahead and intraday, is proposed. This strategy uses a curriculum-based reinforcement learning method for optimization and adjustment, fully utilizing the dispatchable potential of flexible resources and improving the economic efficiency of microgrid operators.
[0101] 2. A time-spanning virtual energy storage model is developed to describe the dynamic demand response process of prosumers’ flexible loads, overcoming the limitations of the static management of existing demand response models and thus further improving the economic benefits of microgrid operators.
[0102] 3. An asynchronous semi-coupled optimization method based on double-delay deep deterministic policy gradient is proposed, which solves the problem of difficulty in establishing an accurate model due to the uncertainty and variability of demand-side resource regulation capabilities, and realizes the full utilization of flexibility resources and maximizes the benefits of microgrid operators. BRIEF DESCRIPTION OF THE DRAWINGS
[0103] Figure 1 Schematic diagram of the process of the present invention;
[0104] Figure 2 This is a comparison chart of reward values. DETAILED DESCRIPTION
[0105] The present invention is further illustrated below with reference to the accompanying drawings and specific embodiments. It should be understood that these embodiments are only used to illustrate the present invention and are not used to limit the scope of the present invention. After reading the present invention, modifications of various equivalent forms of the present invention made by those skilled in the art all fall within the scope defined by the claims attached to this application.
[0106] Example 1:
[0107] like Figure 1 As shown, this embodiment provides a microgrid joint optimization scheduling method integrating heterogeneous flexibility resources, including the following steps:
[0108] S1: Establishing photovoltaic load model;
[0109] The microgrid system consists of N prosumers with different characteristics. Prosumers can reduce their net load by utilizing their rooftop photovoltaics. The mathematical expression of photovoltaic power and load demand is as follows:
[0110]
[0111]
[0112] in, 、 are the PV panel power and load power of prosumer i in time t, respectively; 、 are the diffusion coefficients of PV panel power and load power, and are the Brownian motion processes used to describe the randomness of PV panels and loads, respectively.
[0113] S2: Establish a day-ahead energy storage rental model;
[0114] During the day-ahead phase, prosumers can obtain a portion of the revenue by leasing idle energy storage capacity to microgrid operators.
[0115] The day-ahead storage rental model includes the benefits and costs of renting storage capacity by prosumers;
[0116] The mathematical expression of the benefits of renting energy storage capacity by prosumers is as follows:
[0117]
[0118] in, is the energy storage capacity rented by prosumer i at time t, Is the highest rental price, is the trend of the rental price curve, is the integration variable;
[0119] The mathematical expression for the cost of renting energy storage capacity by prosumers is as follows:
[0120]
[0121] in, 、 and are the coefficients of the quadratic term, linear term, and constant term of the leasing energy storage cost of the i-th prosumer.
[0122] S3: Establish a prosumer optimization model based on the PV load model and the day-ahead energy storage rental model;
[0123] The goal of prosumers in leasing idle energy storage capacity is to maximize their net profit. The mathematical expression of the optimization model for prosumers to lease energy storage capacity is as follows:
[0124]
[0125] in, Rent energy storage capacity for optimized prosumers, It is the maximum energy storage capacity rented by the prosumer.
[0126] S4: Establishing a virtual energy storage model within a day;
[0127] During the day, when the real-time electricity price is higher than the peak-valley price, prosumers can reduce the cost of microgrid operators by adjusting the dispatchable load. The dispatchable load of prosumers includes transferable load and curtailable load;
[0128] For constant temperature control load, the mathematical expression of power consumption is as follows:
[0129]
[0130] in, is the power consumption of the thermostatically controlled load i, is the baseline power consumption of thermostatically controlled load i, is the regulating power of the constant temperature control load i participating in the response; the thermodynamic kinetic equation of the constant temperature control load is as follows:
[0131]
[0132] in, is the internal heat capacity of building i, and are the equivalent heat gain and thermal resistance of building i, is the indoor temperature of building i; is the coefficient of constant temperature control load i; is the set temperature of the thermostatically controlled load i; is the diffusion coefficient of the flexible load; is a Wiener process with random disturbances caused by prosumer behavior or environmental influences;
[0133] Define the auxiliary variable as and , and reconstruct the thermodynamic kinetic equations into a virtual energy storage model, whose mathematical expression is as follows:
[0134]
[0135] in, represents the remaining capacity of virtual energy storage unit i; represents the energy loss rate of virtual energy storage unit i; and Represent the minimum and maximum remaining capacity respectively; and Represent the minimum and maximum power constraints respectively; the initial remaining capacity Defined as ; The remaining capacity after termination is [ , ]; The demand response model conversion process for other flexible loads follows a similar procedure and will not be described here.
[0136] Because the participation of prosumers' virtual energy storage in microgrid operator scheduling will lead to a decrease in power comfort, microgrid operators need to provide financial compensation to encourage the active participation of prosumers. The virtual energy storage model includes the compensation cost of the microgrid operator, and its mathematical expression is as follows:
[0137]
[0138] in, and are the compensation coefficients of the first constant term and the second constant term respectively, and N is the number of constant temperature control loads.
[0139] S5: Establish a daily shared energy storage model;
[0140] Microgrid operators use the leased energy storage capacity to store or release electricity to shared energy storage during the day-ahead phase. The mathematical expression of the shared energy storage model is as follows:
[0141]
[0142] in, yes The derivative of It is the surplus energy storage capacity leased by the microgrid operator to the prosumer. represents the energy storage power, and are the minimum and maximum energy storage powers, represents the energy dissipation rate of stored energy; It represents the efficiency of energy storage charging and discharging power. Its mathematical expression is as follows:
[0143]
[0144] in, and are the efficiency of charging and discharging power respectively.
[0145] S6: Establish a microgrid operator optimization model based on the prosumer optimization model, the intraday virtual energy storage model, and the intraday shared energy storage model;
[0146] Microgrid operators purchase electricity from the external grid and sell it to prosumers. When the operator has excess electricity, it can also sell electricity to the external grid. The mathematical expression of the cost / revenue obtained by the operator and the external grid area is as follows:
[0147]
[0148] in, is the real-time electricity price of the external grid, is the interactive power between the operator and the external grid, where a negative value indicates the power sold by the operator to the external grid and a positive value indicates the power purchased by the operator from the external grid;
[0149] Similarly, when the photovoltaic power generation of the prosumer exceeds the load demand, the operator can purchase electricity from the prosumer. The mathematical expression of the cost / revenue obtained by the operator by purchasing or selling electricity to the prosumer is as follows:
[0150]
[0151] in, is the price at which the operator sells energy to the prosumer, is the interaction power between operators and prosumers, is a coefficient whose value range is (0,1); I function is the indicator function;
[0152] To ensure the safety of system operation, the microgrid must meet the power balance constraints during operation in each time slot:
[0153]
[0154] in, is the load demand of prosumer i;
[0155] The mathematical expression of the optimization model of the microgrid operator is as follows:
[0156]
[0157] The objective function consists of four parts: the benefits / costs generated by the electricity interaction between operators and prosumers , the benefits / costs of power interaction between operators and external grids , the cost of leasing energy storage capacity and compensation costs to motivate prosumers .
[0158] S7: Convert the microgrid operator optimization model into a well-established Markov decision process;
[0159] The Markov decision process consists of four parts: state space, action space, state transition probability, and reward function, as follows:
[0160] State space: The system state is divided into the day-ahead state and intraday stage status ; However, the ES capacity status in the day-ahead phase Unable to obtain in advance, its corresponding status is unobservable; in the intraday stage, the state is observable, and its mathematical expression is as follows:
[0161]
[0162] The intraday status information includes the load demand of prosumers, distributed photovoltaics, time-of-use electricity prices, real-time electricity prices, shared energy storage capacity, and virtual energy storage capacity.
[0163] Action Space: Operator actions are divided into day-ahead phase actions and intraday phase actions In the day-ahead phase, operators acquire shared energy storage capacity by developing leasing strategies. In the intraday phase, operators earn revenue by managing shared energy storage and virtual energy storage. The action space is defined as:
[0164]
[0165] in, and The value of is used to determine the day-ahead leasing strategy. They are the shared energy storage and virtual energy storage power dispatched by the operator during the day;
[0166] Transition probability function: Transition probability is usually not a static parameter known in advance, but a dynamic variable that the agent gradually estimates through interactive exploration with the environment;
[0167] Reward function: The system’s reward function is divided into the day-ahead reward and intraday rewards In the day-ahead phase, the operator's goal is to minimize the rental cost of energy storage capacity while meeting the real-time capacity demand of energy storage; in the intraday phase, the operator's goal is to maximize its revenue. In summary, the reward function is defined as:
[0168]
[0169] in, is the reward coefficient;
[0170] According to the above Markov decision process formula, the discounted cumulative reward within time slot t is defined as:
[0171]
[0172] in, Indicates time slot Rewards before or during the day, is the discount factor; the goal of the Markov decision process is to maximize the expected cumulative discounted return:
[0173]
[0174] in, is described as a mapping function that transforms the state Mapping to Actions , is the mathematical expectation; and Represents the status and action of the day before or within the day respectively.
[0175] S8: Solve the Markov decision process of step S7, update and iterate the scheduling parameters to achieve optimal scheduling of the microgrid;
[0176] In this embodiment, an asynchronous semi-coupled optimization method based on double-delayed deep deterministic policy gradient is used to solve the Markov decision process, as follows:
[0177] The training process for both the intraday and day-ahead phases is performed via an actor-critic process with double-delayed deep deterministic policy gradients, where the actor network learns a policy π and generates actions under a given state, and the critic network evaluates the learning quality of the actor network under policy π, i.e., the objective of the Markov decision process, which is to maximize the expected discounted return.
[0178] The critic network completes policy evaluation by constructing an action-value function, which quantifies the expected return of taking a specific action in a given state; the action-value function is defined as:
[0179]
[0180] in, are the current network parameters, and the value of Q is iteratively updated by utilizing the Bellman equation:
[0181]
[0182] in, is the target Q value, are the target network parameters, and are the next state and action respectively; is the reward for performing action a in state s; is the expected value;
[0183] In order to suppress the overestimation bias of Q value and steadily train the learning strategy, the Q value is iteratively updated through the double Q learning method:
[0184]
[0185] This method utilizes two independent critic networks and ,When calculating, the minimum value of the two is taken to reduce the Q value estimation deviation and improve stability.
[0186] The loss function of the critic network is usually based on the time difference error, and its goal is to make the expected Q value as close as possible to the true expected return. The parameters of the critic network are obtained by minimizing the loss function. To update; the actor network learns to optimize the policy π by maximizing the expected discounted reward, and by using the policy gradient To update the parameters of the actor network; the policy gradient is defined as follows:
[0187]
[0188] in, is the gradient of action a; is the network parameter gradient; is the network parameter The strategy for the next state S;
[0189] Based on the update iteration of the parameters of the above optimization method, the optimal scheduling problem of microgrid flexibility resources is solved.
[0190] Example 2:
[0191] Based on the method of Example 1, this embodiment provides a microgrid joint optimization scheduling system integrating heterogeneous flexibility resources, including:
[0192] Model building module, used to build photovoltaic load model, day-ahead energy storage rental model, prosumer optimization model, intraday virtual energy storage model, intraday shared energy storage model and microgrid operator optimization model;
[0193] A decision process establishment module is used to establish a Markov decision process and convert the microgrid operator optimization model into the established Markov decision process;
[0194] Solving module, used to solve Markov decision processes.
[0195] Example 3:
[0196] In order to verify the effect of the method of the present invention, this embodiment is compared through simulation experiments, as follows:
[0197] In this example, the reward values obtained by the proposed method are compared with those of three existing methods in 10,000 training iterations. The three existing methods are TD3, DDPG and PSO. Figure 2 The reward value comparison chart shown.
[0198] like Figure 2 As shown in the figure, due to the exploration stage, the reward value of the proposed method is low in the initial stage, but as time goes by, the reward value gradually stabilizes. Specifically, the proposed method reaches stability after about 4,100 iterations, while the TD3 method requires about 5,200 iterations to reach stability. This shows that the proposed method combined with curriculum learning significantly speeds up the training process, reaching stability 19.23% faster than TD3 in terms of the number of iterations. In addition, the proposed method and TD3 eventually reach a reward value of about 58, while the reward values of DDPG and PSO are less than 40. Therefore, the proposed method not only has an advantage in convergence speed, but also performs significantly better than DDPG and PSO.
Claims
1. A microgrid joint optimization scheduling method integrating heterogeneous flexibility resources, characterized in that: The steps include: S1: Establishing photovoltaic load model; S2: Establish a day-ahead energy storage rental model; S3: Establish a prosumer optimization model based on the PV load model and the day-ahead energy storage rental model; S4: Establishing a virtual energy storage model within a day; S5: Establish a daily shared energy storage model; S6: Establish a microgrid operator optimization model based on the prosumer optimization model, the intraday virtual energy storage model, and the intraday shared energy storage model; S7: Convert the microgrid operator optimization model into a well-established Markov decision process; S8: Solve the Markov decision process of step S7, update and iterate the scheduling parameters to achieve optimal scheduling of the microgrid; The establishment of the intraday virtual energy storage model in step S4 includes: The dispatchable load of the prosumer includes the transferable load and the curtailable load. For the constant temperature control load, the mathematical expression of its power consumption is as follows: in, is the power consumption of the thermostatically controlled load i, is the baseline power consumption of thermostatically controlled load i, is the regulating power of the constant temperature control load i participating in the response; the thermodynamic kinetic equation of the constant temperature control load is as follows: Among them, C i is the internal heat capacity of building i, Q i and R i are the equivalent heat gain and thermal resistance of building i, θ i is the indoor temperature of building i; is the coefficient of constant temperature control load i; is the set temperature of the constant temperature control load i; σ vir is the diffusion coefficient of the flexible load; W vir is a Wiener process with random disturbances caused by prosumer behavior or environmental influences; Define the auxiliary variable as and The thermodynamic kinetic equations are reconstructed into a virtual energy storage model, and its mathematical expression is as follows: in, represents the remaining capacity of virtual energy storage unit i; represents the energy loss rate of virtual energy storage unit i; and Represent the minimum and maximum remaining capacity respectively; and Represent the minimum and maximum power constraints respectively; the initial remaining capacity Defined as The remaining capacity after termination is in the range of The virtual energy storage model includes the compensation cost of the microgrid operator, and its mathematical expression is as follows: Wherein, λ1 and λ2 are the compensation coefficients of the first constant term and the second constant term respectively, and N is the number of constant temperature control loads.
2. A microgrid joint optimization scheduling method integrating heterogeneous flexibility resources according to claim 1, characterized in that: The establishment of the photovoltaic load model in step S1 includes: The mathematical expressions for PV power and load requirements are as follows: in, are the PV panel power and load power of prosumer i in time t; ρ PV , ρ L are the diffusion coefficients of PV panel power and load power, W PV and W L are the Brownian motion processes used to describe the randomness of PV panels and loads, respectively.
3. A microgrid joint optimization scheduling method integrating heterogeneous flexibility resources according to claim 2, characterized in that: The establishment of the day-ahead energy storage rental model in step S2 includes: The day-ahead storage rental model includes the benefits and costs of renting storage capacity by prosumers; The mathematical expression of the benefits of renting energy storage capacity by prosumers is as follows: Among them, ξ i,t is the energy storage capacity rented by prosumer i at time t, α t is the maximum rental price, β t is the trend of the rental price curve, and x is the integration variable; The mathematical expression for the cost of renting energy storage capacity by prosumers is as follows: Among them, a i 、b i and c i are the coefficients of the quadratic term, linear term, and constant term of the leasing energy storage cost of the i-th prosumer.
4. A microgrid joint optimization scheduling method integrating heterogeneous flexibility resources according to claim 3, characterized in that: The establishment of the prosumer optimization model in step S3 includes: The goal of prosumers in leasing idle energy storage capacity is to maximize their net profit. The mathematical expression of the optimization model for prosumers to lease energy storage capacity is as follows: in, Rent energy storage capacity for optimized prosumers, It is the maximum energy storage capacity rented by the prosumer.
5. A microgrid joint optimization scheduling method integrating heterogeneous flexibility resources according to claim 4, characterized in that: The establishment of the intraday shared energy storage model in step S5 includes: Microgrid operators use the leased energy storage capacity to store or release electricity to shared energy storage during the day-ahead phase. The mathematical expression of the shared energy storage model is as follows: in, yes The derivative of It is the surplus energy storage capacity leased by the microgrid operator to the prosumer. represents the energy storage power, and are the minimum and maximum energy storage powers, γ ES represents the energy dissipation rate of stored energy, η ES It represents the efficiency of energy storage charging and discharging power. Its mathematical expression is as follows: Among them, η in and η out are the efficiency of charging and discharging power respectively.
6. A microgrid joint optimization scheduling method integrating heterogeneous flexibility resources according to claim 5, characterized in that: The establishment of the microgrid operator optimization model in step S6 includes: The mathematical expression for the cost / revenue gained from the exchange of electricity between the operator and the external grid area is as follows: in, is the real-time electricity price of the external grid, is the interaction power between the operator and the external grid; The mathematical expression of the cost / revenue that operators receive by purchasing or selling electricity to prosumers is as follows: in, is the price at which the operator sells energy to the prosumer, is the interaction power between operators and prosumers, is the coefficient; I function is the indicator function; The microgrid meets the power balance constraints during operation in each time slot: in, is the load demand of prosumer i; The mathematical expression of the optimization model of the microgrid operator is as follows: The objective function consists of four parts: the benefits / costs generated by the electricity interaction between operators and prosumers Benefits / costs of electricity interaction between operators and external grids Cost of leasing energy storage capacity and compensation costs to motivate prosumers 7. A microgrid joint optimization scheduling method integrating heterogeneous flexibility resources according to claim 6, characterized in that: The Markov decision process in step S7 includes four parts: state space, action space, state transition probability, and reward function, which are as follows: State space: The system state is divided into the day-ahead state and intraday stage status However, the ES capacity status in the day-ahead phase Unable to obtain in advance, its corresponding status is unobservable; in the intraday stage, the state is observable, and its mathematical expression is as follows: The intraday status information includes the load demand of prosumers, distributed photovoltaics, time-of-use electricity prices, real-time electricity prices, shared energy storage capacity, and virtual energy storage capacity. Action Space: Operator actions are divided into day-ahead phase actions and intraday phase actions In the day-ahead phase, operators acquire shared energy storage capacity by developing leasing strategies. In the intraday phase, operators earn revenue by managing shared energy storage and virtual energy storage. The action space is defined as: Among them, α t and β t The value of is used to determine the day-ahead leasing strategy. and They are the shared energy storage and virtual energy storage power dispatched by the operator during the day; Transition probability function: The transition probability is a dynamic variable that the agent gradually estimates through interactive exploration with the environment. Reward function: The system’s reward function is divided into the day-ahead reward and intraday rewards The reward function is defined as: Among them, ε is the reward coefficient; According to the Markov decision process formula, the discounted cumulative reward within time slot t is defined as: Among them, r t+k represents the reward before or within the time slot t+k, γ∈(0,1] is the discount coefficient; the goal of the Markov decision process is to maximize the expected cumulative discounted return: Among them, π:S→A is described as a mapping function that transforms the state s t Mapped to action a t , is the mathematical expectation; s t and a t Represents the status and action of the day before or within the day respectively.
8. A microgrid joint optimization scheduling method integrating heterogeneous flexibility resources according to claim 7, characterized in that: In step S8, an asynchronous semi-coupled optimization method based on double-delayed deep deterministic policy gradient is used to solve the Markov decision process, as follows: The training process for both the intraday and day-ahead phases is performed via an actor-critic process with double-delayed deep deterministic policy gradients, where the actor network learns a policy π and generates actions under a given state, and the critic network evaluates the learning quality of the actor network under policy π, i.e., the objective of the Markov decision process, which is to maximize the expected discounted return. The critic network completes policy evaluation by constructing an action-value function, which quantifies the expected return of taking a specific action in a given state; the action-value function is defined as: Where θ is the current network parameter, and the value of Q is iteratively updated using the Bellman equation: in, is the target Q value, θ' is the target network parameter, s' and a' are the next state and action respectively; r(s,a) is the reward for performing action a in state s; E is the expected value; Iteratively update the Q value through the double Q learning method: The parameters of the critic network are obtained by minimizing the loss function To update; the actor network learns to optimize the policy π by maximizing the expected discounted reward, and by using the policy gradient To update the parameters of the actor network; the policy gradient is defined as follows: in, is the gradient of action a; is the gradient of the network parameter φ; π φ (s) is the strategy of state S under network parameter φ; Based on the update iteration of the parameters of the above optimization method, the optimal scheduling problem of microgrid flexibility resources is solved.
9. A microgrid joint optimization scheduling system integrating heterogeneous flexibility resources according to the method of claim 1, characterized in that: include: Model building module, used to build photovoltaic load model, day-ahead energy storage rental model, prosumer optimization model, intraday virtual energy storage model, intraday shared energy storage model and microgrid operator optimization model; A decision process establishment module is used to establish a Markov decision process and convert the microgrid operator optimization model into the established Markov decision process; Solving module, used to solve Markov decision processes.
Citation Information
Patent Citations
Virtual energy storage-based collaborative day-ahead optimization scheduling method for interconnected micro-grid system
CN115021327A
Power dispatching optimization operation strategy considering uncertainty of new energy in urban environment
CN115811095A