A micro-grid double-layer scheduling method and system considering conditional risk
Patent Information
- Application Number
- CN202611079689.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-21
- Publication Date
- 2026-08-18
AI Technical Summary
[0007]本发明提供了一种考虑条件风险的微电网双层调度方法及系统,用于解决现有技术中经济信号与物理安全边界脱节的技术问题
[0033] The technical solution of this invention employs a zero-order conditional risk value-simultaneous perturbation stochastic approximation iterative algorithm to solve the objective function. By using simultaneous perturbation stochastic approximation, only two function evaluations are needed to construct the zero-order gradient approximation, thereby greatly reducing computational complexity. By setting the expected social cost and tail risk of conditional risk value in the objective function, and setting constraints on the physical feasible region and the expected unsupplied energy, this method of directly writing the tail risk into the leader's objective function can align economic signals with the physical security boundary, ultimately obtaining an economically and physically acceptable scheduling scheme.
Smart Images

Figure CN122600323A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of power system optimization dispatching technology, and in particular to a two-layer dispatching method and system for microgrids that takes into account conditional risks. Background Technology
[0002] With a high proportion of renewable energy being integrated into microgrids, the system faces the dual requirements of optimizing economic efficiency and protecting against extreme risks. Existing methods are mainly divided into four categories: (1) Robust Optimization and Distributed Robust Optimization (RO / DRO): This method describes the randomness of renewable energy output by constructing an uncertainty set and optimizes for the worst-case scenario. The drawbacks are: a large amount of redundancy needs to be reserved to ensure the safety boundary, resulting in significantly higher daily operating costs; and it is difficult to couple with a two-tier market structure.
[0003] (2) Deep Reinforcement Learning (DRL): This method uses algorithms such as SAC and TD3 to learn adaptive scheduling strategies through trial and error. Its drawbacks include: the decision-making process is a black box with insufficient interpretability; exploration behavior may lead to exceeding physical constraints; and there is a lack of posterior auditing mechanisms.
[0004] (3) Heuristic algorithms (such as PSO): They are simple to implement, but they are prone to getting trapped in local optima in high-dimensional decision spaces, have slow convergence speed, and lack systematic risk control mechanisms.
[0005] (4) Two-level scheduling method based on Stackelberg game: describes the hierarchical interaction between operators and followers, but existing technologies are mostly based on the assumption of smooth convex targets, which makes it difficult to handle non-smooth risk measures such as CVaR (Conditional Value at Risk); it relies on gradient available or mixed integer programming solvers, which have high computational complexity; and it lacks an explicit governance mechanism for EENS (Expected Energy Not Supplied).
[0006] However, these methods may have the following drawbacks: they can easily lead to economically acceptable but physically fragile scheduling schemes; risk indicators such as CVaR and EENS are often used as post-processing constraints and are not endogenized into the optimization objective; and non-smooth high-dimensional objectives are difficult to solve using traditional gradient methods. In particular, DRL-type methods cannot explain the physical source of policy advantage and are difficult to distinguish whether the performance improvement comes from risk transfer across time periods, accidental sampling, or over-conservatism. Summary of the Invention
[0007] This invention provides a two-layer scheduling method and system for microgrids that takes into account conditional risks, in order to solve the technical problem of the disconnect between economic signals and physical security boundaries in the prior art.
[0008] According to one aspect of the present invention, a two-layer scheduling method for microgrids considering conditional risks is provided, comprising: Acquire historical data on renewable energy and loads in microgrid systems; Modeling is performed based on historical data of the renewable energy sources and loads, and several renewable energy access scenarios are generated. Multiple preset scheduling algorithms are used to calculate the expected energy unsupplied value for each scenario, resulting in the expected energy unsupplied value distribution. Based on the expected energy unsupplied value distribution, an expected energy unsupplied threshold is determined. The expected energy unsupplied threshold is used to constrain the expected energy unsupplied value. The objective function is solved using the zero-order conditional value at risk-simultaneous perturbation stochastic approximation iterative algorithm. The objective function includes the expected social cost and the tail risk of the conditional value at risk. The constraints include the physical feasible region constraint and the expected unsupplied energy constraint. After the iterative calculation is completed, the optimal quota decision and the optimal risk threshold are output.
[0009] Optionally, modeling is performed based on the historical data of the renewable energy and load, and several renewable energy access scenarios are generated, including: The Gaussian Copula model was used to model the joint distribution of wind and solar power output. Several new energy intervention scenarios were generated by Monte Carlo sampling, and then K-means scenario reduction was used for clustering and compression. Typical scenarios were obtained, each of which included 24-hour load trajectories, available wind power, and available solar power.
[0010] Optionally, the step of calculating the expected energy shortage value for each scenario using multiple preset scheduling algorithms to obtain the expected energy shortage value distribution, and determining the expected energy shortage threshold based on the expected energy shortage value distribution, includes: Run the particle swarm optimization reference controller and the robust optimization reference controller respectively to obtain the expected power shortage sample set under each scenario; The expected power shortage sample set corresponding to the particle swarm optimization reference controller is merged with the expected power shortage sample set corresponding to the robust optimization reference controller to obtain the merged expected power shortage sample distribution. Calculate the 75th percentile of the merged expected power shortage sample distribution and set this percentile as the expected power shortage threshold.
[0011] Optionally, the objective function includes:
[0012] in, Let P_alloc be the objective function; P_alloc is the power quota decision matrix for M generating units over T scheduling periods. = For the expected social cost, N scen This represents the total number of scenes. The cost is at the scenario level; eta is the risk threshold variable. For the confidence level, w mean and w risk These are the expected social cost weighting coefficient and the conditional value-at-risk tail risk weighting coefficient, respectively, and they satisfy w mean + w risk = 1, C M and C V This is the normalization constant.
[0013] Optionally, the scene-level cost in the objective function An augmented form can be used:
[0014] The formula for cost pressure is:
[0015] in, This represents an augmented form of cost at the scenario level; Cost pressure; lambda_s is the soft penalty coefficient for the expected unsupplied energy EENS; represents the supply-demand imbalance in scenario xi during time period t; mean(imb) represents the average imbalance across all time periods; This is the cost coefficient; During the optimization process, the cost of the tail risk term of the conditional value at risk is converted from augmented cost. Substitute; It also includes reliability constraints:
[0016] in, EENS(xi) represents the expected energy not supplied value in each scenario, and epsilon_EENS is the hard constraint threshold of the expected energy not supplied value EENS.
[0017] Optionally, the step of employing the zero-order conditional value-at-risk-simultaneous-perturbation stochastic approximation iterative algorithm to solve the objective function includes: Generate synchronization disturbance vector Each dimension of the perturbation vector is independently sampled from a Bernoulli distribution; Calculate the positive perturbation evaluation value based on the perturbation vector. and negative disturbance assessment value ;P kFor the quota decision in the k-th iteration, eta k The risk threshold for the k-th iteration; Estimate the zeroth gradient based on the positive perturbation evaluation value and the negative perturbation evaluation value:
[0018] Among them, c k These are the gain sequence parameters; Update quota decisions:
[0019] The formula for the composite gradient is:
[0020] The gradient clipping formula is:
[0021] in, For the quota decision in the k+1th iteration; This is a composite gradient; the gradient clipping formula constrains the range of the composite gradient. To prevent excessively large updates to quota decisions in a single iteration, the quota adjustment for each unit in each time period is limited to ±100 MW; Proj P (.) represents the projection operation onto the physically feasible region; The zero-order gradient represents the marginal impact of the current quota decision on costs; w mean and w risk These are the expected social cost weighting coefficient and the conditional value at risk tail risk weighting coefficient, respectively. As a driver of expected costs, it guides quota decisions toward optimizing the direction of reducing average costs; This is the CVaR risk leverage term, i.e., when the current scenario cost L... k Exceeding the risk threshold At that time, the gradient is amplified and enhanced by a factor of 1 / (1- ); As a risk indicator, additional risk penalty gradients are applied only to scenarios where costs exceed a threshold; after each parameter update, the quota decision parameters need to be projected into the physical feasible region to ensure that all scheduling decisions satisfy the physical boundaries. Update risk thresholds:
[0022] in The learning rate is the risk threshold, and I is the indicator function; The risk threshold for the k-th iteration; Update gain sequence parameters: ,
[0023] a is the step gain scaling factor, a k c is the step size gain scaling factor for the k-th iteration; c is the perturbation step size scaling factor. k is the perturbation step size scaling factor for the k-th iteration; A is the stability constant, alpha_sp is the step size decay exponent, and gamma_sp is the perturbation decay exponent; Check if the iteration stopping condition is met; if so, stop the iteration.
[0024] Optionally, the projection operation Proj P(·) include: For each component P of the decision quota vector {i,t} If P {i,t} If P < 0, then let P {i,t} = 0; if P {i,t} >P i max Then let P {i,t} =P i max For CHP units, additionally check ramp-up constraints: if |P {CHP,t} -P {CHP,t-1} |>R {CHP} Then adjust P {CHP,t} To satisfy the climbing constraint, P {CHP,t} These are the power quota decision variables for the CHP units during time period t; The ramp rate limit for CHP of a combined heat and power (CHP) unit; where P {i,t} P represents the quota for unit i during time period t; i max This indicates the maximum quota for a single generator unit per time period.
[0025] Optionally, the constraints also include follower price response constraints: Define the follower's profit function:
[0026] in, For the profit of followers; P {alloc,i} Specifically refers to the 24-hour quota trajectory of unit i; c i For follower i, kappa is the physical elasticity parameter; The unconstrained optimal price is obtained through the first-order optimality condition, and the k-constrained truncated price is obtained by truncating the price by the upper bound:
[0027]
[0028] in, The optimal bid for follower i; arg max represents finding the bid that maximizes utility; U i c is the utility function of follower i; i The quote for follower i; The power quota that the leader allocates to the follower; Time-of-use pricing for the power grid; The absolute value of the difference between the follower's bid and the grid price; k is the price elasticity parameter; Let t be the electricity price during time period t.
[0029] Optionally, after outputting the optimal quota decision and the optimal risk threshold, the following method is further included: Calculate the cost, risk, and reliability corresponding to the output results, and output the final scheduling strategy, cost, electricity price, and conditional risk value.
[0030] According to another aspect of the present invention, a two-layer dispatch system for microgrids considering conditional risks is provided, comprising: The data acquisition unit is used to acquire historical data on renewable energy and loads in the microgrid system; The scenario generation unit is used to model based on the historical data of the renewable energy and load, and generate several renewable energy access scenarios; The threshold calibration unit is used to calculate the expected energy unsupplied value for each scenario using multiple preset scheduling algorithms, obtain the expected energy unsupplied value distribution, and determine the expected energy unsupplied threshold based on the expected energy unsupplied value distribution; the expected energy unsupplied threshold is used to constrain the expected energy unsupplied. The iterative computation unit is used to solve the objective function using the zero-order conditional value at risk-simultaneous perturbation stochastic approximation iterative algorithm; the objective function includes the expected social cost and the tail risk of the conditional value at risk; the constraints include the physical feasible region constraint and the expected energy unsupply constraint. The output unit is used to output the optimal quota decision and the optimal risk threshold after the iterative calculation is completed.
[0031] According to another aspect of the present invention, an electronic device is provided, the electronic device comprising: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor, which enables the at least one processor to perform the microgrid two-layer scheduling method considering conditional risks as described in any embodiment of the present invention.
[0032] According to another aspect of the present invention, a computer-readable storage medium is provided, the computer-readable storage medium storing computer instructions for causing a processor to execute and implement the microgrid two-layer scheduling method considering conditional risks as described in any embodiment of the present invention.
[0033] The technical solution of this invention employs a zero-order conditional risk value-simultaneous perturbation stochastic approximation iterative algorithm to solve the objective function. By using simultaneous perturbation stochastic approximation, only two function evaluations are needed to construct the zero-order gradient approximation, thereby greatly reducing computational complexity. By setting the expected social cost and tail risk of conditional risk value in the objective function, and setting constraints on the physical feasible region and the expected unsupplied energy, this method of directly writing the tail risk into the leader's objective function can align economic signals with the physical security boundary, ultimately obtaining an economically and physically acceptable scheduling scheme.
[0034] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of the present invention, nor is it intended to limit the scope of the invention. Other features of the invention will become readily apparent from the following description. Attached Figure Description
[0035] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0036] Figure 1 This is a flowchart of a two-layer scheduling method for microgrids that considers conditional risks, provided according to Embodiment 1 of the present invention; Figure 2 This is a schematic diagram illustrating the decoupling of the present invention from other scheduling methods in terms of 24-hour cost-fluctuation. Figure 3 This is a schematic diagram illustrating the application of the present invention and other scheduling methods in extreme scenarios. Figure 4 This is a CHP capacity sensitivity analysis diagram comparing the present invention with other scheduling methods; Figure 5 This is an experimental analysis diagram of κ ablation compared to other scheduling methods; Figure 6This is a schematic diagram of a two-layer dispatch system for microgrids that takes into account conditional risks, provided in Embodiment 2 of the present invention. Figure 7 This is a schematic diagram of the structure of an electronic device that implements the microgrid two-layer scheduling method that takes into account conditional risks, according to an embodiment of the present invention. Detailed Implementation
[0037] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.
[0038] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0039] Example 1 Figure 1 The flowchart illustrates a two-layer scheduling method for microgrids that considers conditional risks, as provided in Embodiment 1 of the present invention. Figure 1 As shown, the method includes: S101. Obtain historical data on renewable energy sources and loads in the microgrid system; Historical wind power, photovoltaic (PV) power, and load data are acquired; in this embodiment, data is acquired from a microgrid system. The microgrid system includes wind power, PV power, and combined heat and power (CHP) units, with a scheduling cycle of 24 time periods T, a time resolution that can be set to 1 hour, a number of units M, and a decision variable dimension of M. T.
[0040] S102. Based on the historical data of the renewable energy and load, a model is created, and several renewable energy access scenarios are generated; It can model the joint distribution of wind and solar power output and generate several renewable energy access scenarios. It generates scenarios containing N scenA set of scenes Xi = {xi1, xi2, ..., xi} of random scenes. {N_scen} Each scenario contains load trajectories and available renewable energy power for T time periods.
[0041] The correlation between wind and solar power output is complex, especially in tail scenarios such as extreme weather (e.g., no wind or solar power, or simultaneous low output due to strong winds and heavy rain). For example, on a cloudy day, solar power output may be low, while wind power output may be low (due to a stationary low-voltage system) or high (due to a passing low-voltage system). Simple correlation coefficients (such as Pearson's correlation coefficient) cannot characterize this nonlinear and asymmetric correlation, especially the probability of simultaneous extreme events.
[0042] Therefore, in a specific embodiment, the Gaussian Copula model (a Gaussian Copula model is a device that "strips out" the dependency structure of multiple random variables and describes it with a multivariate normal distribution) can be used to model the joint distribution of wind power and photovoltaic power output; several new energy intervention scenarios are generated by Monte Carlo sampling, and then K-means scenario reduction is used for clustering and compression; typical scenarios are obtained, and each typical scenario contains a 24-hour load trajectory, available wind power, and available photovoltaic power.
[0043] In this embodiment, the Gaussian Copula model can be used to model the joint distribution of wind power and photovoltaic power output; the K-means scenario reduction method is used to cluster and compress 100 representative profiles (scenarios) from 1000 Monte Carlo sampling scenarios; each scenario contains 24-hour load trajectories, available wind power, and available photovoltaic power.
[0044] S103. Calculate the expected energy unsupplied value for each scenario using multiple pre-set scheduling algorithms, obtain the expected energy unsupplied value distribution, and determine the expected energy unsupplied threshold based on the expected energy unsupplied value distribution.
[0045] The pre-built scheduling algorithms include those described in the background section, such as Robust Optimization and Distributed Robust Optimization (RO / DRO), Deep Reinforcement Learning (DRL), heuristic algorithms (such as PSO), and a two-level scheduling method based on Stackelberg game theory. For typical scenarios after clustering and compression, multiple algorithms can be used to calculate the expected unsupplied energy value for each scenario, thereby obtaining the expected unsupplied energy value distribution for multiple typical scenarios corresponding to each algorithm. The expected unsupplied energy value distributions of multiple algorithms are then merged, and the expected unsupplied energy value threshold is determined based on these distributions. For example, a quantile from 50% to 90% of the expected unsupplied energy value distribution can be selected as the expected unsupplied energy value threshold.
[0046] In one specific embodiment, a particle swarm optimization (PSO) reference controller and a robust optimization (RO) reference controller can be run separately to obtain a sample set of expected power shortages for each scenario; the sample set of expected power shortages corresponding to the PSO reference controller and the sample set of expected power shortages corresponding to the robust optimization reference controller are merged to obtain a merged sample distribution of expected power shortages; the 75th percentile of the merged sample distribution of expected power shortages is calculated and set as the threshold for the expected energy shortage value.
[0047] Of course, in practice, multiple scheduling algorithms can be used to determine the expected energy unsupplied value distributions for various scheduling algorithms, and the expected energy unsupplied value threshold can be determined by merging multiple expected energy unsupplied value distributions. The expected energy unsupplied value threshold can be used to constrain the expected energy unsupplied value.
[0048] S104. The zero-order conditional risk value-simultaneous perturbation stochastic approximation iterative algorithm is used to solve the objective function; the objective function includes the expected social cost and the tail risk of the conditional risk value; the constraints include the physical feasible region constraint and the expected energy unsupply constraint.
[0049] In this embodiment, the zero-order conditional value at risk-simultaneous perturbation stochastic approximation iterative algorithm CVaR-SPSA can be used to solve the objective function. This embodiment incorporates both the expected social cost and the tail risk of the conditional value at risk into the objective function, rather than using risk indicators such as CVaR and EENS as post-processing constraints, thus aligning the economic signals with the physical safety boundary.
[0050] Specifically, the composite objective function that includes expected social cost and tail risk of conditional value at risk (CVaR) is constructed as follows:
[0051] The objective function contains two terms: the first is the expected social cost, and the second is the tail risk term of the conditional value at risk. P_alloc is the power quota decision matrix for M generating units over T scheduling periods. = For the expected social cost, N scen This represents the total number of scenes. The cost is at the scenario level; eta is the risk threshold variable. For the confidence level, w mean and w risk These are the expected social cost weighting coefficient and the conditional value-at-risk tail risk weighting coefficient, respectively, and they satisfy w mean + w risk = 1, C M and CV This is the normalization constant.
[0052] The scenario-level cost in the objective function An augmented form can be used:
[0053] The formula for cost pressure is:
[0054] Where lambda_s is the soft penalty coefficient of the expected unsupplied energy value EENS; Representing scenario x i The supply-demand imbalance at time period t represents the power deficit; mean(imb) represents the average imbalance across all time periods. This is a cost coefficient, which can be set to 0.05. It is a conversion factor that transforms physical fluctuations into economic costs, reflecting the aging / degradation costs of equipment caused by frequent fluctuations. Time-based pressure weighting: 400 for off-peak hours (0-15 o'clock) and 2000 for peak hours (16-24 o'clock), indicating that the imbalance fluctuations during peak hours put greater pressure on the system; This represents stress cost, which quantifies the physical stress that the system experiences due to supply and demand imbalances and transforms it into an economic cost item, prompting the algorithm to learn a smooth scheduling strategy. During the optimization process, the cost of the tail risk term of the conditional value at risk is converted from augmented cost. Substitute; It also includes reliability constraints:
[0055] in, EENS(xi) represents the expected energy not supplied value in each scenario, and epsilon_EENS is the expected energy not supplied value EENS hard constraint threshold, which is related to the expected energy not supplied threshold.
[0056] In this embodiment, lambda_s and epsilon_EENS are pre-calibrated fixed parameters that remain unchanged during the optimization process. For example, lambda_s can be set to 0.02 MWh / MWh and epsilon_EENS to 4520.37 MWh / day.
[0057] In the aforementioned objective function, the CVaR tail risk is directly incorporated into the leader objective function, and a dual governance approach of soft penalties and hard constraints is adopted for EENS. Subsequent evaluations of the method presented in this invention compared to other scheduling methods show that the method presented in this invention reduces average cost by 14.8%, CVaR by 14.2%, and worst-case loss by 15.2% (relative to PSO). This method, which aligns economic signals with physical safety boundaries, can produce scheduling schemes that are acceptable both economically and physically.
[0058] In this embodiment, before proceeding to the iterative algorithm, the follower's profit function can be defined first:
[0059] Among them, P {alloc,i} Specifically refers to the 24-hour quota trajectory of unit i; c i For follower i, kappa is the physical elasticity parameter; The follower profit function uses a quadratic cost term. For price c i Implicit constraints: When the quote c i At lower levels, the linear revenue term Dominant, profits increase with increasing quotes; when quote c i Once the critical value is exceeded, the rate of decrease in secondary costs outpaces the rate of increase in linear revenue, leading to a reduction in profits. This causes followers to "voluntarily" limit their bids to a reasonable range, preventing them from indefinitely raising prices. This formula is directly related to the power quota decision matrix P_alloc in the objective function, where P_{alloc,i} is the 24-hour quota trajectory allocated to unit i after the objective function is optimized. Followers optimize their bids based on this quota.
[0060] C in the IPO i The unconstrained optimal price can be obtained through the first-order optimality condition, and the k-constrained truncated price can be obtained by truncating the price by the upper bound:
[0061]
[0062] in, The optimal offer for follower i (ultimately determining the selling price); arg max represents finding the offer that maximizes utility; U i c is the utility function of follower i (which can be understood as "profit" or "satisfaction"); i Let i be the quote for follower i (a variable that needs optimization); The power quota allocated by the leader (grid operator) to the follower (e.g., 100MW of electricity allocated to you); The electricity price is based on the time-of-use pricing of the power grid (e.g., 0.8 yuan / kWh during the day and 0.3 yuan / kWh at night). The absolute value of the difference between the follower's bid and the grid price (the degree of deviation); k is the price elasticity parameter (controlling the range of allowable deviation, for example, K=0.2 means that a deviation of 20% is allowed); Let t be the grid price during time period t. The follower price response constraint is to ensure that the follower's bid does not deviate too far from the grid price and must remain within a "neighborhood" of the grid price.
[0063] In a preferred embodiment, the physical resilience parameter kappa is calibrated according to the nominal operating capacity of the system, and the time-of-use electricity price c of the power grid is... grid A unified TOU electricity price series is adopted.
[0064] Traditional methods often involve unconstrained follower price responses, potentially leading to overreactions. This invention addresses this by using a quadratic profit function with an upper price bound to limit pricing within an interpretable range. This method reduces adverse pricing feedback, improves solution stability and structural interpretability, and minimizes the variance of the optimal objective values for multiple sub-subjects, indicating optimal solution stability.
[0065] This invention employs a zero-order conditional value at risk-simultaneous perturbation stochastic approximation iterative algorithm. The specific process of iteratively solving the objective function is as follows: S1041. Calculate the SPSA gain sequence parameters and perturbation amplitude parameters, and generate a random perturbation vector; When using the zero-order conditional value at risk-simultaneous perturbation stochastic approximation iterative algorithm, the parameters are first initialized. These initialized parameters include wind power installed capacity, photovoltaic installed capacity, CHP maximum scaling capacity, CHP ramping constraint, scheduling cycle, time resolution, over-limit safety threshold, over-limit penalty factor, EENS hard constraint threshold, EENS, conditional value at risk tail risk weight coefficient, and confidence level. Total number of scenes.
[0066] Hyperparameter settings: Initial gain a0, perturbation gain c0, stability offset A, step size decay exponent alpha_sp, perturbation decay exponent gamma_sp, risk threshold, learning rate lr eta Maximum number of iterations K, normalization constant C M C V Initialize quota decision P0 and risk threshold eta0.
[0067] For each iteration k = 1, 2, ..., K: Generate synchronization disturbance vector Each dimension of the perturbation vector is independently sampled from a Bernoulli distribution; Calculate the positive perturbation evaluation value based on the perturbation vector. and negative disturbance assessment value ;P k For the quota decision in the k-th iteration, eta k This is the risk threshold for the k-th iteration.
[0068] S1042. Under K constraints, assess cost risk internalization based on Conditional Value at Risk (CVaR) and Energy Shortage EENS. Calculate the supply-demand imbalance in time period t under scenario xi:
[0069] in, For the load of time period t under scenario xi, This represents the quota for the i-th unit during time period t.
[0070] Define overload:
[0071] Among them, L safe This is the preset over-limit safety threshold.
[0072] Define the expected energy not supplied:
[0073] in This is the scheduling time resolution.
[0074] Update risk thresholds:
[0075] in The learning rate is the risk threshold. Where I is the confidence level, and I is the indicator function; Update gain sequence parameters: ,
[0076] Formula explanation: The attenuation gain form recommended by Spall (1998) is adopted, with parameters set as a=1.0, c=2.0, A=3000, alpha_sp=0.602, gamma_sp=0.101, which satisfies the Robbins-Monro condition for SPSA convergence.
[0077] S1043. Construct the zeroth-order gradient approximation, update the parameters, and project them onto the feasible region; Estimating the zeroth gradient:
[0078] Among them, ck These are the gain sequence parameters; Update quota decisions:
[0079] The formula for the composite gradient is:
[0080] The gradient clipping formula is:
[0081] Among them, Proj P (.) represents the projection operation onto the physically feasible region; For the quota decision in the k+1th iteration; The composite gradient; the composite gradient obtained from the gradient clipping formula. This indicates a comprehensive optimization approach that considers both economic efficiency and extreme risks to achieve a balance between the two objectives of "pursuing economy in normal times and safety in extreme scenarios." To prevent excessively large updates to quota decisions in a single iteration, avoid "chaos" or "divergence" in the optimization process, limit the quota adjustment of each unit to no more than ±100 MW in each time period, and ensure the physical feasibility and numerical stability of scheduling decisions; The zero-order gradient represents the marginal impact of the current quota decision on costs; (i.e., the rate of change of costs during quota fine-tuning); w mean and w risk These are the expected social cost weighting coefficient and the conditional value at risk tail risk weighting coefficient, respectively. As a driver of expected costs, it guides quota decisions toward optimizing the direction of reducing average costs; This is the CVaR risk leverage term, i.e., when the current scenario cost L... k Exceeding the risk threshold At that time, the gradient is amplified and enhanced by a factor of 1 / (1- ); As a risk indicator, additional risk penalty gradients are applied only to scenarios where costs exceed a threshold; after each parameter update, the quota decision parameters need to be projected into the physical feasible region to ensure that all scheduling decisions satisfy the physical boundaries. In each iteration update, the quota decision needs to be projected onto the physically feasible region as follows:
[0082] Among them, P alloc Let P be the power quota decision matrix for M generating units over T scheduling periods, and let P be the iterative decision vector. k The power quota component corresponding to unit i, P kIncludes quotas for all units at all times. This is the upper limit of the capacity of unit i. For the ramp rate limit of CHP of combined heat and power units; P {CHP,t} These are the power quota decision variables for the CHP units during time period t; This refers to the ramp rate limit for the CHP (Concentrated Heat and Power) unit. The quota decision after gradient update may exceed the physical feasible region, therefore it needs to be mapped back to the feasible region via a projection operation, Proj. P(·) Specifically: for each component P of the decision vector {i,t} If P {i,t} If P < 0, then let P {i,t} = 0; if P {i,t} >P i max Then let P {i,t} =P i max For CHP units, additionally check ramp-up constraints: if |P {CHP,t} -P {CHP,t-1} |>R {CHP} Then adjust P {CHP,t} This satisfies the climbing constraint. Where P {i,t} P represents the quota for unit i during time period t; i max This indicates the maximum quota for a single generator unit per time period.
[0083] In the iterative algorithm of this embodiment, an initial quota P0 is first given. Then, in each iteration, the gradient g_hat_k of the objective function with respect to the quota is estimated by simultaneous perturbation stochastic approximation (SPSA). Next, the quota is updated along the negative gradient direction to obtain P. k - a k ·g_hat_k, and finally Proj through projection operations. P (·) Map the updated values to the physical feasible region to obtain P that satisfies the box constraint and the ramp constraint. {CHP,t} Specifically, the projection operation first checks the non-negativity constraints and capacity limits, then checks the ramping constraints. If |P {CHP,t} - P {CHP,t-1} |>R {CHP} Then P {CHP,t} Restricted to P {CHP,t-1} ±R {CHP} Therefore P {CHP,t} Essentially, it is the result of the combined effect of gradient optimization and physical constraint projection.
[0084] S1044. Determine whether the iteration stopping condition is met. If yes, output the optimal quota decision and the optimal risk threshold. Otherwise, proceed to the next iteration.
[0085] Check the stopping criteria; if they are met, terminate the iteration. The stopping criteria are reaching the maximum number of iterations K, or the change in the decision vector falling below a preset convergence threshold: i.e. .
[0086] Traditional gradient descent methods are not robust to non-smooth targets. This invention employs simultaneous perturbation stochastic approximation, requiring only two function evaluations to construct a zero-order gradient approximation. The computational complexity of this invention is reduced from O(2^n) to O(2), supports joint updates of 72-dimensional decision variables, and has significantly fewer convergence iterations than reinforcement learning.
[0087] S105. After the iterative calculation is completed, output the optimal quota decision and the optimal risk threshold.
[0088] After iterative calculations are completed, this invention can output the optimal quota decision and the optimal risk threshold. The optimal quota decision P and the optimal risk threshold eta are output.
[0089] Perform probability-intensity decomposition on the optimization results: E[vio] = P(vio>0) E[vio|vio>0] Where E[vio] is the expected value of overload, which is the average overload in all scenarios; P(vio>0) represents the probability of overload, indicating what proportion of scenarios will experience overload; E[vio | vio>0] is the conditional expectation (intensity) of the average severity of overload in scenarios where overload occurs.
[0090] Identify the extreme scenarios that cause the comparison method to generate the highest over-limit penalty, compare the scheduling trajectory, pressure price index and over-limit penalty distribution of the method of this invention and the comparison method, and verify whether there are cross-time period risk transfer characteristics.
[0091] The pressure price indicator is defined as follows:
[0092] Among them, c pen For exceeding the limit penalty factor; The overload amplitude during time period t under scenario xi.
[0093] In one embodiment, the method further includes: calculating the cost, risk, and reliability corresponding to the output results, and outputting the final scheduling strategy, cost, electricity price, and conditional risk value.
[0094] The costs are calculated at the scenario level and include system operating costs, overload penalty costs, expected energy shortage penalty costs, and stress costs. System operating costs include wind and solar curtailment penalty costs, fuel and maintenance costs for combined heat and power (CHP) units, electricity purchase costs from the grid, and follower settlement costs. Overload penalty costs are the product of the overload extent and the overload penalty factor. Expected energy shortage penalty costs include both soft and hard penalty terms. Stress costs characterize the degree of fluctuation in supply-demand imbalance. Expected social costs are the average of all scenario-level costs. Electricity prices are optimized using a follower profit function. Followers calculate their unconstrained optimal bid based on their allocated power quotas and physical elasticity parameters, and the final bid is truncated by the upper bound of the grid's time-of-use pricing.
[0095] This method can be jointly compared with other scheduling methods, all of which share the same stress scenario pool: a unified pool of 100 24-hour net load and renewable output trajectories, a unified time resolution, and a unified source of historical data. At the physical level, all controllers share the same feasible region P, the same overload threshold Lsaf, and the same upper bound εEENS, with a unified simulation backend responsible for detecting and recording various constraint violations. At the post-processing level, all overload frequencies, overload amplitudes, condition severity, daily average stress standard deviation, and EENS statistics are calculated by the same post-processing script in the same result format.
[0096] For control methods involving randomness (including SPSA, SAC, TD3, etc.), under a unified scene pool and physical constraints, repeated experiments with multiple random seeds are conducted, and the mean and dispersion of key indicators are reported to reflect the convergence behavior and robustness of the algorithm, rather than drawing conclusions based solely on the results of a single run.
[0097] The comparison results of the method of this invention with the PSO heuristic baseline, the SAC / TD3 reinforcement learning baseline, and the RO robust optimization baseline are shown in Table 1: Table 1. Main Baseline Test Results
[0098] Performance analysis: Compared with PSO, the present invention reduces average cost by 14.8%, worst-case loss by 15.2%, and CVaR by 14.2%; compared with RO, the average cost reduces by 19.1%, worst-case loss by 1.9%, and CVaR by 6.9%.
[0099] The cost of penalties for exceeding limits decreased by 27.9% relative to PSO and by 39.9% relative to RO; the frequency of exceeding limits decreased by 13.8% relative to PSO and by 32.8% relative to RO.
[0100] The inference latency is on the same order of magnitude as the reinforcement learning baseline (approximately 0.1ms), meeting the real-time scheduling requirements; the number of convergence iterations is significantly less than that of reinforcement learning methods; the variance of the optimal objective values for multiple sub-subsidiaries is minimized, indicating optimal solution stability.
[0101] The schematic diagram of the decoupling of cost and fluctuation between this invention and other scheduling methods in 24 hours is shown below. Figure 2 As shown, Figure 2 The two views on the left represent the mean and volatility of the out-of-limit amount, while the two diagrams on the right represent the frequency and intensity of out-of-limit occurrences. The method of this invention reduces the mean out-of-limit amount by 27.9% and the standard deviation by 24.4% while keeping the volatility of out-of-limit amounts to a minimum, significantly outperforming comparative methods such as SAC, TD3, and RL. This demonstrates the dual advantages of CVaR endogenization and zero-order stochastic approximation in risk suppression and volatility control.
[0102] The slice diagram of this invention and other scheduling methods in extreme scenarios is shown below. Figure 3 As shown, Figure 3 The data reflects the residual net load, CHP scheduling trajectory, shadow price curve, and over-limit penalty bar chart. Figure 3 In the study, the physical action trajectory and economic pressure map under the selected worst-case scenario show that, under the worst-case scenario, the CHP scheduling trajectory of the method of this invention closely follows the changes in the residual load of the scenario, and the shadow price and the over-limit penalty are significantly lower than those of the SAC method. This proves that the k-constrained follower response mechanism can effectively reduce adverse pricing feedback and reduce economic pressure.
[0103] The CHP capacity sensitivity analysis diagram of this invention and other scheduling methods is shown in the figure below. Figure 4 As shown, Figure 4 The image shows the transition boundary from the high-voltage zone to the plateau zone for a 600MW power plant. From... Figure 4 As can be seen, when the upper limit of CHP capacity reaches the pressure boundary of 600MW, the average imbalance and average cost of the method of the present invention tend to stabilize and remain at a low level, indicating that the method has robust adaptability to CHP capacity constraints and will not blindly pursue cost reduction due to capacity increase, leading to reliability deterioration.
[0104] Experimental analysis of κ ablation between this invention and other scheduling methods is shown in the figure below. Figure 5 As shown. From Figure 5 As can be seen, after introducing the physical elasticity parameter k, the method of this invention improves the final optimal target value by 14.8%, reduces the average loss in the later stage by 10.0%, reduces the 95th percentile loss in the later stage by 7.3%, and reduces the number of evaluations required to reach 95% progress by 4.7%, which verifies the comprehensive benefits of the k-constraint mechanism in improving solution efficiency, reducing loss and accelerating convergence.
[0105] Wherein, CVaR is the average system scheduling cost (loss) under the worst-case α% scenario.
[0106] The post-processing methods described above can support the reproduction of summary tables and mechanism diagrams, avoid reliance on implicit pre-processing procedures, and provide post-verification evidence consistent with risk transfer across time periods.
[0107] Example 3 Figure 6 This is a schematic diagram of a two-layer dispatch system for microgrids that considers conditional risks, provided in Embodiment 2 of the present invention. Figure 6 As shown, the device includes: Data acquisition unit 601 is used to acquire historical data on renewable energy and loads in the microgrid system; The scenario generation unit 602 is used to model based on the historical data of the renewable energy and load, and generate several renewable energy access scenarios; The threshold calibration unit 603 is used to calculate the expected energy unsupplied value for each scenario using multiple preset scheduling algorithms, obtain the expected energy unsupplied value distribution, and determine the expected energy unsupplied threshold based on the expected energy unsupplied value distribution; the expected energy unsupplied threshold is used to constrain the expected energy unsupplied. The iterative calculation unit 604 is used to solve the objective function using the zero-order conditional value at risk-simultaneous perturbation stochastic approximation iterative algorithm; the objective function includes the expected social cost and the tail risk of the conditional value at risk; the constraints include the physical feasible region constraint and the expected energy unsupply constraint. Output unit 605 is used to output the optimal quota decision and the optimal risk threshold after the iterative calculation is completed.
[0108] The microgrid two-layer scheduling system considering conditional risks provided in the embodiments of the present invention can execute the microgrid two-layer scheduling method considering conditional risks provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of the execution method.
[0109] Example 3 Figure 7 A schematic diagram of an electronic device 10, which can be used to implement embodiments of the present invention, is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices (e.g., helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.
[0110] like Figure 7 As shown, the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) 12 or a random access memory (RAM) 13, communicatively connected to the at least one processor 11. The memory stores computer programs executable by the at least one processor. The processor 11 can perform various appropriate actions and processes based on the computer program stored in the ROM 12 or loaded from storage unit 18 into the RAM 13. The RAM 13 can also store various programs and data required for the operation of the electronic device 10. The processor 11, ROM 12, and RAM 13 are interconnected via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.
[0111] Multiple components in electronic device 10 are connected to I / O interface 15, including: input unit 16, such as keyboard, mouse, etc.; output unit 17, such as various types of displays, speakers, etc.; storage unit 18, such as disk, optical disk, etc.; and communication unit 19, such as network card, modem, wireless transceiver, etc. Communication unit 19 allows electronic device 10 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0112] Processor 11 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. Processor 11 performs the various methods and processes described above, such as a two-level scheduling method for microgrids that takes conditional risks into account.
[0113] In some embodiments, a condition-risk-considered microgrid two-tier scheduling method can be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program can be loaded and / or installed on electronic device 10 via ROM 12 and / or communication unit 19. When the computer program is loaded into RAM 13 and executed by processor 11, one or more steps of the condition-risk-considered microgrid two-tier scheduling method described above can be performed. Alternatively, in other embodiments, processor 11 can be configured to perform a condition-risk-considered microgrid two-tier scheduling method by any other suitable means (e.g., by means of firmware).
[0114] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0115] Computer programs used to implement the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be performed. The computer programs may be executed entirely on a machine, partially on a machine, or as a standalone software package, partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0116] In the context of this invention, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination thereof. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.
[0117] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0118] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or middleware components (e.g., application servers), or frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.
[0119] A computing system can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system to address the shortcomings of traditional physical hosts and VPS services, such as high management difficulty and weak business scalability.
[0120] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this invention can be achieved, and this is not limited herein.
[0121] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.
Claims
1. A two-layer dispatch method for microgrids considering conditional risks, characterized in that, include: Acquire historical data on renewable energy and loads in microgrid systems; Modeling is performed based on historical data of the renewable energy sources and loads, and several renewable energy access scenarios are generated. Multiple preset scheduling algorithms are used to calculate the expected energy unsupplied value for each scenario, resulting in the expected energy unsupplied value distribution. Based on the expected energy unsupplied value distribution, an expected energy unsupplied threshold is determined. The expected energy unsupplied threshold is used to constrain the expected energy unsupplied value. The objective function is solved using the zero-order conditional value at risk-simultaneous perturbation stochastic approximation iterative algorithm. The objective function includes the expected social cost and the tail risk of the conditional value at risk. The constraints include the physical feasible region constraint and the expected unsupplied energy constraint. After the iterative calculation is completed, the optimal quota decision and the optimal risk threshold are output.
2. The microgrid two-layer dispatch method considering conditional risks according to claim 1, characterized in that, Based on historical data of the renewable energy sources and loads, a model is created, and several renewable energy access scenarios are generated, including: The Gaussian Copula model was used to model the joint distribution of wind and solar power output. Several new energy intervention scenarios were generated by Monte Carlo sampling, and then K-means scenario reduction was used for clustering and compression. Typical scenarios were obtained, each of which included 24-hour load trajectories, available wind power, and available solar power.
3. The microgrid two-layer dispatch method considering conditional risks according to claim 1, characterized in that, The process involves using multiple pre-defined scheduling algorithms to calculate the expected energy shortage value for each scenario, obtaining the expected energy shortage value distribution, and determining the expected energy shortage threshold based on the expected energy shortage value distribution, including: Run the particle swarm optimization reference controller and the robust optimization reference controller respectively to obtain the expected power shortage sample set under each scenario; The expected power shortage sample set corresponding to the particle swarm optimization reference controller is merged with the expected power shortage sample set corresponding to the robust optimization reference controller to obtain the merged expected power shortage sample distribution. Calculate the 75th percentile of the merged expected power shortage sample distribution and set this percentile as the expected power shortage threshold.
4. The microgrid two-layer dispatch method considering conditional risks according to claim 1, characterized in that, The objective function includes: in, Let P_alloc be the objective function; P_alloc is the power quota decision matrix for M generating units over T scheduling periods. = For the expected social cost, N scen This represents the total number of scenes. The cost is at the scenario level; eta is the risk threshold variable. For the confidence level, w mean and w risk These are the expected social cost weighting coefficient and the conditional value-at-risk tail risk weighting coefficient, respectively, and they satisfy w mean + w risk = 1, C M and C V This is the normalization constant.
5. The microgrid two-layer dispatch method considering conditional risks according to claim 4, characterized in that, The scenario-level cost in the objective function An augmented form can be used: The formula for cost pressure is: in, This represents an augmented form of cost at the scenario level; Cost pressure; lambda_s is the soft penalty coefficient for the expected unsupplied energy EENS; represents the supply-demand imbalance in scenario xi during time period t; mean(imb) represents the average imbalance across all time periods; This is the cost coefficient; During the optimization process, the cost of the tail risk term of the conditional value at risk is converted from augmented cost. Substitute; It also includes reliability constraints: in, EENS(xi) represents the expected energy not supplied value in each scenario, and epsilon_EENS is the hard constraint threshold of the expected energy not supplied value EENS.
6. The microgrid two-layer dispatch method considering conditional risks according to claim 4, characterized in that, The method employs a zero-order conditional value-at-risk-simultaneous-perturbation stochastic approximation iterative algorithm to solve the objective function, including: Generate synchronization disturbance vector Each dimension of the perturbation vector is independently sampled from a Bernoulli distribution; Calculate the positive perturbation evaluation value based on the perturbation vector. and negative disturbance assessment value ;P k For the quota decision in the k-th iteration, eta k The risk threshold for the k-th iteration; Estimate the zeroth gradient based on the positive perturbation evaluation value and the negative perturbation evaluation value: Among them, c k These are the gain sequence parameters; Update quota decisions: The formula for the composite gradient is: The gradient clipping formula is: in, For the quota decision in the k+1th iteration; This is a composite gradient; the gradient clipping formula constrains the range of the composite gradient. To prevent excessively large updates to quota decisions in a single iteration, the quota adjustment for each unit in each time period is limited to ±100 MW; Proj P (.) represents the projection operation onto the physically feasible region; The zero-order gradient represents the marginal impact of the current quota decision on costs; w mean and w risk These are the expected social cost weighting coefficient and the conditional value at risk tail risk weighting coefficient, respectively. As a driver of expected costs, it guides quota decisions toward optimizing the direction of reducing average costs; This is the CVaR risk leverage term, i.e., when the current scenario cost L... k Exceeding the risk threshold At that time, the gradient is amplified and enhanced by a factor of 1 / (1- ); As a risk indicator, additional risk penalty gradients are applied only to scenarios where costs exceed a threshold; after each parameter update, the quota decision parameters need to be projected into the physical feasible region to ensure that all scheduling decisions satisfy the physical boundaries. Update risk thresholds: in The learning rate is the risk threshold, and I is the indicator function; The risk threshold for the k-th iteration; Update gain sequence parameters: , a is the step gain scaling factor, a k c is the step size gain scaling factor for the k-th iteration; c is the perturbation step size scaling factor. k is the perturbation step size scaling factor for the k-th iteration; A is the stability constant, alpha_sp is the step size decay exponent, and gamma_sp is the perturbation decay exponent; Check if the iteration stopping condition is met; if so, stop the iteration.
7. The microgrid two-layer dispatch method considering conditional risks according to claim 6, characterized in that, The projection operation Proj P(·) include: For each component P of the decision quota vector {i,t} If P {i,t} If P < 0, then let P {i,t} = 0; if P {i,t} >P i max Then let P {i,t} =P i max For CHP units, additionally check ramp-up constraints: if |P {CHP,t} -P {CHP,t-1} |>R {CHP} Then adjust P {CHP,t} To satisfy the ramping constraint, P{CHP,t} is the power quota decision variable of the CHP unit in time period t; The ramp rate limit for CHP of a combined heat and power (CHP) unit; where P {i,t} P represents the quota for unit i during time period t; i max This indicates the maximum quota for a single generator unit per time period.
8. The microgrid two-layer dispatch method considering conditional risks according to claim 4, characterized in that, The constraints also include follower price response constraints: Define the follower's profit function: in, For the profit of followers; P {alloc,i} Specifically refers to the 24-hour quota trajectory of unit i; c i For follower i, kappa is the physical elasticity parameter; The unconstrained optimal price is obtained through the first-order optimality condition, and the k-constrained truncated price is obtained by truncating the price by the upper bound: in, The optimal bid for follower i; arg max represents finding the bid that maximizes utility; U i c is the utility function of follower i; i The quote for follower i; The power quota that the leader allocates to the follower; Time-of-use pricing for the power grid; The absolute value of the difference between the follower's bid and the grid price; k is the price elasticity parameter; Let t be the electricity price during time period t.
9. The two-layer dispatch method for microgrids considering conditional risks according to claim 1, characterized in that, Following the output of the optimal quota decision and the optimal risk threshold, the following is also included: Calculate the cost, risk, and reliability corresponding to the output results, and output the final scheduling strategy, cost, electricity price, and conditional risk value.
10. A two-layer dispatch system for microgrids considering conditional risks, characterized in that, include: The data acquisition unit is used to acquire historical data on renewable energy and loads in the microgrid system; The scenario generation unit is used to model based on the historical data of the renewable energy and load, and generate several renewable energy access scenarios; The threshold calibration unit is used to calculate the expected energy unsupplied value for each scenario using multiple preset scheduling algorithms, obtain the expected energy unsupplied value distribution, and determine the expected energy unsupplied threshold based on the expected energy unsupplied value distribution; the expected energy unsupplied threshold is used to constrain the expected energy unsupplied. The iterative computation unit is used to solve the objective function using the zero-order conditional value at risk-simultaneous perturbation stochastic approximation iterative algorithm; the objective function includes the expected social cost and the tail risk of the conditional value at risk; the constraints include the physical feasible region constraint and the expected energy unsupply constraint. The output unit is used to output the optimal quota decision and the optimal risk threshold after the iterative calculation is completed.