A new energy power grid forward scheduling method and device

By establishing a two-layer robust optimization model and a constrained reinforcement learning method, the problem of balancing the solution speed and policy robustness of the forward scheduling model under the uncertainty of new energy in the new energy power grid is solved. This achieves rapid response and security of power grid scheduling and is applicable to forward scheduling of new energy power grids.

CN120879603BActive Publication Date: 2026-07-07WUHAN UNIV
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
WUHAN UNIV
Filing Date
2025-07-10
Publication Date
2026-07-07

AI Technical Summary

Technical Problem

In new energy power grids, existing technologies struggle to balance the solution speed of forward scheduling models with the robustness and security of strategies under the strong uncertainties of new energy, especially in extreme scenarios where it is difficult to guarantee the security and rapid response of strategies.

Method used

Uncertainty is represented by chance constraints, and a two-layer robust optimization model is established. The chance constraints are decomposed into deterministic current flow and uncertain current flow, and the cumulative probability density function of slack variables is used to characterize the boundary. Combined with constraint reinforcement learning methods, a risk-avoiding agent based on SAC algorithm and Lagrangian function is constructed. Training is accelerated by imitation learning technology, and an offline simulation environment is constructed for pre-training.

Benefits of technology

It achieves a balance between solution speed, strategy robustness, and security in new energy power grids, improves the ability to handle extreme scenarios, and ensures the safety and economy of power grid dispatch.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120879603B_ABST
    Figure CN120879603B_ABST
Patent Text Reader

Abstract

The application provides a new energy power grid forward scheduling method and device, and relates to the technical field of power system daily economic scheduling. The application firstly constructs an opportunity constraint optimal power flow model of forward scheduling, analyzes the influence of new energy uncertainty on the opportunity constraint, and establishes a constraint Markov decision process of forward scheduling, then uses a risk evaluator network fitting the probability distribution of a risk function and an actor network considering the performance of extreme scenarios to enhance the processing capacity of the agent for the forward scheduling scene of extreme ramping events containing new energy, and finally uses the imitation learning technology to accelerate the training of the agent in the offline simulation environment of the power grid forward scheduling. The application can balance the solution speed, strategy robustness and safety of the double-layer robust optimization model of forward scheduling.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of intraday economic dispatching technology in power systems, specifically to a forward-looking dispatching method and apparatus for new energy power grids. Background Technology

[0002] With the continued advancement of the "dual carbon" target, it is projected that by 2060, the proportion of electricity generated by new energy sources will exceed 55%, and the installed capacity of wind and solar power will reach over 5 billion kilowatts, gradually becoming the main source of electricity supply. However, the inherent uncertainty of new energy sources makes accurate forecasting difficult. During severe weather events such as cold waves and typhoons, the maximum deviation in ultra-short-term forecasts can reach as high as 40%, forcing power dispatching agencies to frequently and significantly revise their daily dispatch plans. In the future, the proportion of electricity generated by new energy sources will continue to rise. The traditional dispatching model, relying on manual experience, will struggle to quickly formulate effective strategies for coordinating and optimizing power generation, grid, load, and storage. The resulting problems of new energy curtailment and load limiting will cause significant losses to the national economy.

[0003] Forward-looking power system dispatch, also known as proactive intraday rolling dispatch, analyzes and extrapolates uncertain scenarios within a few hours' time window. Based on system operating status and continuously updated ultra-short-term forecasts, it quantifies potential future power deficits and adaptively adjusts active power plans for various adjustable resources within the system. This allows for real-time online provision of safe, robust, economical, and reliable dispatch schemes within the future dispatch window. Under the strong uncertainty of renewable energy sources, forward-looking power dispatch can be viewed as an optimal power flow problem with chance constraints. Essentially, it is a nonlinear mixed-integer programming problem with extremely high computational complexity.

[0004] Model-driven methods for solving optimal power flow problems with opportunity constraints mainly include: analytical transformation of opportunity constraints, simulation methods, and methods based on Gaussian mixture models and power flow analysis. Analytical transformation of opportunity constraints relies on the known probability density function of renewable energy uncertainties to derive the distribution offset factor, but it is difficult to accurately transform under arbitrary distributions of renewable energy output. Simulation methods and Gaussian mixture models are suitable for arbitrarily distributed renewable energy uncertainties, constructing opportunity constraint boundaries through preset control variables, but they limit the operating space of opportunity-constrained optimal power flow, affecting the economic efficiency of the results. Traditional power flow analysis, based on a large amount of historical or simulated data, derives the system operating boundary margin required to mitigate renewable energy uncertainties, offering high accuracy but also a heavy computational burden, making it unsuitable for real-time power system operation and dispatch.

[0005] Compared to model-driven methods, reinforcement learning holds great promise for look-ahead scheduling problems due to its fast search and solution capabilities. However, conventional reinforcement learning methods fail to adequately consider constraints in optimization scheduling, severely limiting the reliability of policies applied in look-ahead scheduling. Constraint-based reinforcement learning improves the usability of reinforcement learning agent output policies in practical applications by incorporating safety constraints during agent training and using risk functions to evaluate the safety of agent actions. Among them, Lagrange-based constraint-based reinforcement learning algorithms effectively improve policy safety by constructing risk functions related to safety constraints and restricting the risk function terms of the agent's output policy within a safety threshold, and are widely used in the power system field. However, the agent considers the cumulative risk function values ​​of the short and long term, only guaranteeing that the average performance of the agent's policy meets the safety threshold limit, and failing to guarantee that all actions do not violate safety constraints. For extreme scenarios that occur infrequently in the training set (i.e., the long-tail effect of scenario distribution), the agent cannot guarantee the safety of its policy, and it is difficult to balance solution speed and policy robustness. Summary of the Invention

[0006] The purpose of this invention is to provide a forward scheduling method and apparatus for new energy power grids, which can solve the problem that it is difficult to balance the solution speed and policy robustness of the current forward scheduling model with opportunity constraints, and can balance the solution speed, policy robustness and security of the forward scheduling model.

[0007] To achieve the above objectives, in a first aspect, the present invention provides a method for forward-looking dispatching of a new energy power grid, comprising:

[0008] Considering extreme ramping events in new energy sources, uncertainty is represented by chance constraints, and a two-layer robust optimization model for look-ahead scheduling is established.

[0009] Opportunity constraints are decomposed into deterministic power flows based on day-ahead scheduling plans and uncertain power flows dominated by new energy power fluctuations. The cumulative probability density function of slack variables is used to characterize the boundary of uncertain power flows, thus transforming opportunity constraints into deterministic constraints.

[0010] A constrained Markov decision process for a two-layer robust optimization model is established. The state space, action space, reward function, risk function, and risk threshold of the agent are established. The constrained reinforcement learning method is used to solve the constrained Markov decision process.

[0011] We construct a risk-avoidance-based constraint reinforcement learning method for scheduling agents, using the SAC algorithm and Lagrange function as the algorithmic framework. We establish an evaluator network based on a truncated Gaussian distribution and construct an actor network that considers the worst-case scenario.

[0012] An offline simulation environment is constructed, and the actor network is pre-trained using imitation learning technology. At fixed time steps, a small batch of trajectories is randomly sampled from the experience pool to calculate the gradient information of the actor network and the evaluator network, and to update the network parameters.

[0013] According to the forward-looking dispatching method for new energy power grids provided by the present invention, in the two-layer robust optimization model, the first-layer max function is used to determine the scenario with the largest prediction deviation, and the second-layer min function is used to minimize the dispatching cost and curtailment cost of the scenario with the largest prediction deviation, as expressed in:

[0014] (1)

[0015] in, (2)

[0016] (3)

[0017] (4)

[0018] (5)

[0019] (6)

[0020] (7)

[0021] (8)

[0022] (9)

[0023] In the formula, C g and C C These represent the cost of power generation and the cost of curtailing renewable energy, respectively. , and These are respectively the set of generator sets, the set of new energy generator sets, and the set of nodes; , and These are the reference operating points of the generator set, new energy unit, and load at time t for the i-th node, respectively; This represents the power transfer distribution factor corresponding to the i-th node; Adjust the generator output of the i-th node at time t; This indicates the generator output adjustment of the i-th node during the k-th time period; and These represent the maximum uphill and downhill climbing capabilities of the generator set at time t for the i-th node, respectively. , and Let be the power deviation, actual power deviation, and curtailment value of the new energy source at time t for the i-th node; This represents the upper limit of new energy sources that the system can accept at time t for node i; Indicates the first l The power limit allowed to be transmitted on each line; E represents the expected value, Pr represents the probability, and α represents the confidence level of the probability; and Let represent the maximum and minimum allowable power generation of the i-th generator unit at time t, respectively.

[0024] According to the prospective dispatch method for new energy power grids provided by the present invention, the opportunity constraint is:

[0025] (10)

[0026] In the formula, The system running point after rolling correction, i.e. ; This is the factor by which the generating unit shares the impact of power fluctuations from new energy sources, i.e. .

[0027] According to the forward-looking dispatching method for new energy power grids provided by the present invention, the deterministic constraint is:

[0028] (11)

[0029] In the formula, and These are the system power flow operating point and the generator set output operating point, respectively. , , , These are the positive and negative relaxation values ​​for line power flow constraints and generator capacity constraints, respectively. The uncertainty boundary of the relaxation quantity; , Let these represent the maximum and minimum allowable generating power of the i-th generator set, respectively. This represents the adjustment amount of the power generation of the i-th generator set.

[0030] According to the present invention, a forward-looking dispatching method for a new energy power grid includes a state space comprising: a system operating point based on a day-ahead plan, a unit ramp mileage, upper and lower limits of unit ramp capacity, upper and lower limits of unit capacity, upper and lower limits of line capacity, a new energy power deviation value, and a new energy regulation upper limit value; the action space is to respond to new energy power fluctuations by adjusting the output of generator units.

[0031] The reward function is:

[0032] (14)

[0033] The risk function is:

[0034] (15)

[0035] In the formula, c1 and c2 represent unit cost coefficients;

[0036] The risk threshold d is set based on the allowable slack of the constraint;

[0037] Representing a constrained Markov decision process as a tuple ,in For state set, For action sets, For reward function, For the penalty function, For state transition function, This represents the initial state distribution.

[0038] According to the present invention, a forward-looking dispatching method for a new energy power grid is provided, wherein the method of solving a constrained Markov decision process using a constrained reinforcement learning method includes:

[0039] Represent the agent's policy as a parameter-dependent neural network. The long-term discount reward value and the long-term discount penalty value are represented as follows:

[0040] (16)

[0041] (17)

[0042] in, It is a discount factor. Representation Strategy The experience trajectory of the agent, i.e. ; This represents all future time steps generated by policy π starting from the current moment. The expected value of the corresponding random variable; Let be a single-step reward function, describing the state s of the agent at time t. t Execute action a t And transition to state s t+1 Instant rewards from environmental feedback; Let be the single-step cost function, describing the agent's state s at time t. t Execute action a t And transition to state s t+1 The immediate cost of environmental feedback;

[0043] The strategy of secure reinforcement learning updates the parameters of the actor neural network. Therefore, the mathematical form of its update process is as follows:

[0044] (18).

[0045] According to the present invention, a forward-looking scheduling method for a new energy power grid is provided, wherein the construction of a scheduling agent based on a risk-avoidance constrained reinforcement learning method with SAC algorithm and Lagrange function as the algorithm framework, and the establishment of an evaluator network based on truncated Gaussian distribution, include:

[0046] Based on the SAC algorithm, a reward evaluator and a risk evaluator are constructed, and an adaptive safety weight is used to balance the reward value and safety. The optimization objective is:

[0047] (19)

[0048] In the formula, To optimize the optimal strategy, For Lagrange multipliers, These are the optimal Lagrange multipliers; L represents the Lagrange function; Let Q be the payoff function under strategy π. The cost Q function under strategy π; Let H0 represent the policy function, which describes the conditional probability distribution of the agent choosing action a in state s; H0 is the baseline entropy.

[0049] Update the original strategy through alternating iterations. and dual variables , To achieve strategy optimization, the loss function is:

[0050] (20)

[0051] In the formula, Represents a projection onto the dual space;

[0052] The neural network parameters for the reward evaluator and risk evaluator of the SAC algorithm are as follows: and Both are updated in the same way; taking the reward evaluator as an example, the loss function is:

[0053] (twenty one)

[0054] In the formula, Indicates based on parameters The Q-function of the return; Indicates based on target network parameters The Q-function of the return;

[0055] The loss function of the actuator is expressed as:

[0056] (twenty two)

[0057] Let the reward evaluator represent the network that estimates the long-term reward of an action, and the risk evaluator represent the network that estimates the long-term risk of an action; the strategy... The probability distribution of the risk function is denoted as a Gaussian distribution. Approximate the fit using a truncated Gaussian distribution. , is represented as:

[0058] (twenty three)

[0059] Based on the Bellman optimality equation, the Q-function and value function are estimated, namely:

[0060] (twenty four)

[0061] (25)

[0062] In the formula, Let represent the policy function, which describes the conditional probability distribution of the agent choosing action a' in state s'; Indicating in strategy p π Under the condition that the agent performs action a in state s, the cost squared C 2 The expected value of the condition; This represents the cutoff threshold for cost distribution.

[0063] To learn the risk assessor, its loss function is estimated by calculating the second-order Wasserstein distance, so that the estimated risk distribution can better match the actual risk distribution; the loss function of the risk assessor is:

[0064] (26)

[0065] in, The calculation function representing the longitudinal error of the distribution shape; The objective cost Q-function represents the teacher's policy μ, and describes the agent's state s at time t. t Next, execute action a t At that time, the expected long-term cumulative discount cost under teacher strategy μ; The Q-function represents the current cost corresponding to the teacher's policy μ, and describes the agent's state s at time t. t Next, execute action a t Real-time cost-value estimation under teacher strategy μ.

[0066] According to the present invention, a forward-looking dispatching method for a new energy power grid includes constructing an actuator network that considers the worst-case scenario, comprising:

[0067] Describing the confidence level of not exceeding the opportunity constraint as the agent's risk aversion level, we define agent policy security based on CVaR, where the risk value of the agent policy is:

[0068] (27)

[0069] In the formula, Represents random variable C π α-quantile (inverse cumulative distribution function);

[0070] An evaluator network based on a truncated Gaussian distribution estimates the CVaR value at each time step and replaces the long-term expected risk value with a new risk metric. The formula for calculating the risk metric of the agent's actions in each iteration is:

[0071] (28)

[0072] in, and Let represent the probability density function and cumulative distribution function of the standard normal distribution, respectively;

[0073] The agent's policy optimization must meet a risk threshold, namely: Following the policy update formula of the SAC algorithm, and taking policy safety into account, the KL divergence is written as:

[0074] (29)

[0075] In the formula, ; It is the partition function of the standardized distribution, which affects the parameters of the neural network. No effect; the loss function of the actioner network is written as:

[0076] (30)

[0077] Based on risk metrics dual variables The update formula is revised as follows:

[0078] (31).

[0079] According to the present invention, a forward-looking scheduling method for a new energy power grid includes constructing an offline simulation environment, pre-training an actor network using imitation learning technology, and randomly sampling a small batch of trajectories from the experience pool at fixed time steps to calculate the gradient information of the actor network and the evaluator network, and updating the network parameters, including:

[0080] Based on the PandaPower power grid offline simulation environment, day-ahead scheduling plans for various scenarios are constructed. Mathematical optimization methods are used to solve a two-layer robust optimization model to obtain prior knowledge. After obtaining the prior knowledge trajectory, the state and actions are reconstructed into features and labels, and a loss function is designed as follows:

[0081] (32)

[0082] In the formula, The training set represents the prior knowledge; This indicates the number of samples taken from the training set; This represents the desired action output by the action network;

[0083] At fixed time steps, a small batch of experience trajectories is randomly selected from the trajectory experience pool of the agent in the stored scenario. This is used to calculate the gradient information of the agent's policy and update the neural network parameters of the agent. The neural network parameters of the target network are also softly updated in the main network. At the same time, the learning rate of the agent is reduced in the later stage of learning so that it can converge to a stable optimal policy.

[0084] Secondly, the present invention provides a new energy power grid forward dispatching device, comprising:

[0085] The modeling unit is used to consider extreme ramping events of new energy sources, and uncertainty is represented by chance constraint form to establish a two-layer robust optimization model for look-ahead scheduling.

[0086] The constraint transformation unit is used to decompose opportunity constraints into deterministic power flow based on day-ahead scheduling plans and uncertain power flow dominated by new energy power fluctuations, and to use the cumulative probability density function of slack variables to characterize the boundary of uncertain power flow, thereby transforming opportunity constraints into deterministic constraints.

[0087] The solution unit is used to establish the constrained Markov decision process of the two-layer robust optimization model, establish the state space, action space, reward function, risk function and risk threshold of the agent, and solve the constrained Markov decision process using the constrained reinforcement learning method.

[0088] The building unit is used to construct a risk-avoidance constrained reinforcement learning method to schedule an agent based on the SAC algorithm and the Lagrange function as the algorithm framework, establish an evaluator network based on a truncated Gaussian distribution, and construct an actor network that takes into account the worst scenarios.

[0089] The training and update unit is used to build an offline simulation environment and pre-train the actor network using imitation learning techniques. Every fixed time step, a small batch of trajectories is randomly sampled from the experience pool to calculate the gradient information of the actor network and the evaluator network and update the network parameters.

[0090] Compared with the prior art, the present invention has at least the following technical effects:

[0091] This invention provides a forward-looking dispatch method and apparatus for a new energy power grid. First, it constructs an opportunity-constrained optimal power flow model for forward-looking dispatch, analyzes the impact of new energy uncertainty on opportunity constraints, and establishes a constrained Markov decision process for forward-looking dispatch. Then, it utilizes a risk evaluator network that fits the probability distribution of the risk function and an actor network that considers performance in extreme scenarios to enhance the agent's ability to handle forward-looking dispatch scenarios involving extreme ramping events of new energy sources. Finally, it accelerates the training of the agent using imitation learning technology in an offline simulation environment for forward-looking dispatch of the power grid. This invention can balance the solution speed, policy robustness, and security of the two-layer robust optimization model for forward-looking dispatch. Attached Figure Description

[0092] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0093] In the attached diagram:

[0094] Figure 1 This is a flowchart of the forward-looking dispatching method for new energy power grids according to the present invention;

[0095] Figure 2 This is a logic block diagram of the forward scheduling method for new energy power grids according to the present invention. Detailed Implementation

[0096] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.

[0097] The following detailed description of some embodiments of the present invention will be provided in conjunction with the accompanying drawings. Unless otherwise specified, the following embodiments and features can be combined with each other.

[0098] Please see Figure 1 and Figure 2 This invention provides a forward scheduling method for high-proportion renewable energy power grids based on risk-avoidance constrained reinforcement learning, comprising the following steps:

[0099] Step 1: Considering extreme ramping events in new energy sources, uncertainty is represented by chance constraints, and a two-layer robust optimization model for look-ahead scheduling is established.

[0100] Specifically, in extreme climbing scenarios for renewable energy sources, such as sudden increases in wind speed in the morning and evening and drastic changes in sunlight intensity at midday, the output of renewable energy often far exceeds the confidence interval of short-term forecasts, thus triggering the risk of power curtailment within the forward-looking window. Considering the limited ramp-up reserves that the system can reserve, the forward-looking scheduling strategy needs to ensure that the system operation and power curtailment costs are reduced under extreme climbing events. This invention considers extreme climbing events of renewable energy sources and represents forward-looking scheduling as a two-layer robust optimization model. The first-layer max function finds the scenario with the largest prediction deviation, and the second-layer min function minimizes the scheduling cost and power curtailment cost of such scenarios, as follows:

[0101] (1)

[0102] in, (2)

[0103] (3)

[0104] (4)

[0105] (5)

[0106] (6)

[0107] (7)

[0108] (8)

[0109] (9)

[0110] in, C g and C C These represent the cost of power generation and the cost of curtailing renewable energy, respectively. , and These are respectively the set of generator units, the set of new energy units, and the set of nodes. , and These are the baseline operating points of the generator, new energy unit, and load at time t, respectively, for the i-th node. This represents the power transmission distribution factor corresponding to the i-th node. Adjust the generator output of the i-th node at time t. This represents the generator output adjustment of the i-th node in the k-th time period. and These represent the maximum uphill and downhill climbing capabilities of the generator set at time t for the i-th node. , and Let represent the power deviation, actual power deviation, and curtailment value of the new energy source at time t for the i-th node. This represents the upper limit of renewable energy that the system can accept at time t for node i; the overall upper limit of acceptable renewable energy is the maximum downhill ramp capability of all generating units, and renewable energy output is curtailed according to a certain proportion, i.e. . Indicates the first l The power limit allowed to be transmitted on each line, where E represents the expected value, Pr represents the probability, and α represents the confidence level of the probability. and Let represent the maximum and minimum allowable power generation of the i-th generator unit at time t, respectively.

[0111] Step 2: The opportunity constraint is decomposed into a deterministic power flow based on the day-ahead scheduling plan and an uncertain power flow dominated by the power fluctuation of new energy sources. The cumulative probability density function of the slack variables is used to characterize the boundary of the uncertain power flow, thus transforming the opportunity constraint into a deterministic constraint and realizing the deduction and analysis of uncertainty.

[0112] Considering that forward-looking scheduling involves rolling adjustments based on day-ahead plans, the opportunity constraint is written as follows:

[0113] (10)

[0114] in, The system operating point after rolling correction, i.e. . This is the factor by which the generator shares the impact of power fluctuations from new energy sources, i.e. Opportunity constraints can be decomposed into deterministic power flow based on day-ahead scheduling plans and uncertain power flow caused by fluctuations in renewable energy power. Considering that the system's operational risks are mainly caused by the uncertain power flow component, the cumulative probability density function of slack variables is used to characterize the boundary of the uncertain power flow, thus transforming opportunity constraints into deterministic constraints.

[0115] (11)

[0116] in, and These are the system power flow operating point and the generator output operating point, respectively. , , , These represent the positive and negative relaxation values ​​for line power flow constraints and generator capacity constraints, respectively. For the uncertainty boundary of the relaxation quantity, it can be determined according to Calculated. , These represent the maximum and minimum allowable power generation (upper and lower limits) of the i-th generator set, respectively. This represents the adjustment amount of the power generation of the i-th generator set.

[0117] Step 3: Establish the constrained Markov decision process of the two-layer robust optimization model of look-ahead scheduling, establish the state space, action space, reward function, risk function and risk threshold of the agent, and solve the constrained Markov decision process using the constrained reinforcement learning method.

[0118] Specifically, for intraday look-ahead scheduling, the agent needs to limit the risk of constraint violation within a risk threshold. This can be solved as a constrained Markov process decision process, and the following important elements are defined:

[0119] 1) Status Corresponding to the constraints of the chance-constrained optimal power flow problem, the agent's state space should include: the system operating point based on the day-ahead plan, the unit ramp-up mileage, the upper and lower limits of the unit ramp-up capability, the upper and lower limits of the unit capacity, the upper and lower limits of the line capacity, the renewable energy power deviation value, and the renewable energy regulation upper limit value. The agent's observation state space can be represented as:

[0120] (12)

[0121] 2) Actions The action space of forward scheduling is to adjust generator output to cope with fluctuations in renewable energy power. Therefore, the action space of the agent can be represented as:

[0122] (13)

[0123] 3) Reward function r Based on the objective function of the real-time optimal power flow under opportunity constraints, the reward function should be related to the generator operating cost and the cost of power curtailment, and is expressed as follows:

[0124] (14)

[0125] 4) Risk Function cThe risk function is used to evaluate the violation of system constraints after the agent performs an action. It is related to the constraint expression and is expressed as:

[0126] (15)

[0127] 5) Risk threshold d Constraint-based reinforcement learning requires setting thresholds based on the desired operating state of the system. In opportunity-constrained optimal power flow problems, uncertainties in renewable energy sources mainly lead to insufficient reserves and excessive power flow across lines. This invention sets risk thresholds based on the allowable slack in constraints.

[0128] Representing a constrained Markov decision process as a tuple ,in For state set, For action sets, For reward function, For the penalty function, For state transition function, This represents the initial state distribution.

[0129] This constrained Markov decision process is solved using constrained reinforcement learning, and the agent's policy is represented as a parameter-dependent neural network. The long-term discount reward value and the long-term discount penalty value can be expressed as follows:

[0130] (16)

[0131] (17)

[0132] in, It is a discount factor. Representation Strategy The experience trajectory of the agent, i.e. The strategy of secure reinforcement learning can be achieved by updating the parameters of the actor neural network. Therefore, the mathematical form of its update process can be expressed as:

[0133] (18)

[0134] Step 4: Construct a risk-avoidance-based constraint reinforcement learning method to schedule agents using the SAC algorithm and the soft actor critic-Lagrangian function (SAC-Lagrangian) as the algorithmic framework. The probability distribution of the risk function under the policy is denoted as a Gaussian distribution. An evaluator network based on the truncated Gaussian distribution is established, and an actor network that takes into account the worst-case scenario is constructed. By defining agent policy security based on conditional risk value, the agent policy is made robust.

[0135] Specifically, SAC-Lagrangian builds upon SAC by constructing two evaluator networks and utilizing an automatically adaptive safety weight to balance reward value and safety. The optimization objective is:

[0136] (19)

[0137] SAC-Lagrangian is a primitive-dual optimization algorithm that updates the primitive policy through alternating iterations. and dual variables , To achieve strategy optimization, the loss function is as follows:

[0138] (20)

[0139] in, This represents a projection onto the dual space.

[0140] The neural network parameters for the SAC-Lagrangian Reward-critic and Safety-critic are as follows: and Both are updated in the same way. Taking Reward-critic as an example, the loss function is:

[0141] (twenty one)

[0142] The loss function of the actor can be expressed as:

[0143] (twenty two)

[0144] Then, a reward evaluator represents the network that estimates the long-term reward of an action, and a risk evaluator represents the network that estimates the long-term risk of an action. The policy... The probability distribution of the risk function is denoted as a Gaussian distribution. Considering that the uncertainty of new energy sources manifests as a bounded probability distribution, using an unbounded Gaussian distribution to calculate the conditional risk value would overestimate the risk magnitude. Therefore, a truncated Gaussian distribution is used to approximate the risk. , is represented as:

[0145] (twenty three)

[0146] Based on the Bellman optimality equation, the Q-function and value function are estimated, namely:

[0147] (twenty four)

[0148] (25)

[0149] To learn the Safety-critic, its loss function is estimated by calculating the second-order Wasserstein distance (bulldozer distance, a distance metric used to measure the difference between two probability distributions), making the estimated risk distribution more closely resemble the actual risk distribution. The loss function of the Safety-critic is:

[0150] (26)

[0151] in, This represents the calculation function for solving the longitudinal error of the distribution shape.

[0152] To ensure the robustness of the agent's policy, an actor network is constructed that considers the worst-case scenario. The confidence level of not exceeding the opportunity constraint is described as the agent's risk aversion level. Agent policy security is defined based on CvaR (Conditional Value at Risk), where the risk value of the agent policy is:

[0153] (27)

[0154] A critic network based on a truncated Gaussian distribution estimates the CVaR value at each time step and uses a new risk metric to replace the long-term expected risk value, providing guidance for the agent's gradient updates. In each iteration, the risk metric for the agent's actions can be calculated as:

[0155] (28)

[0156] in, and Let represent the probability density function and cumulative distribution function of the standard normal distribution, respectively.

[0157] The agent's policy optimization must meet a risk threshold, namely: Following the policy update formula of SAC, and taking policy safety into account, the KL divergence (Kullback-Leibler Divergence) can be written in the following form:

[0158] (29)

[0159] in, . It is the partition function of the standardized distribution, which affects the parameters of the neural network. No impact. The loss function of the actioner network can be...

[0160] (30)

[0161] Based on risk metrics The dual variable can be The update formula can be modified as follows:

[0162] (31)

[0163] Step 5: Construct an offline simulation environment and pre-train the actor network using imitation learning techniques. At fixed time steps, randomly sample a small batch of trajectories from the experience pool to calculate the gradient information of the Actor and Critic networks and update the neural network parameters. For off-policy SAC-type algorithms, the parameters of the Target Network are also synchronized to the Main Network periodically using a soft update method. In the later stages of training, gradually reduce the learning rate to ensure that the policy converges to a stable, optimal policy.

[0164] Based on the PandaPower power grid offline simulation environment, day-ahead scheduling plans for various scenarios are constructed. Mathematical optimization methods are used to solve the look-ahead scheduling model to obtain prior knowledge. After obtaining the prior trajectory, the state and actions are reconstructed into "features" and "labels," and the following loss function is designed:

[0165] (32)

[0166] in, The training set represents the prior knowledge. This indicates the number of samples taken from the training set. This indicates the expected action output by the actor network.

[0167] At fixed time steps, a small batch of experience trajectories is randomly retrieved from the agent's trajectory experience pool stored in the scene. These trajectories are used to calculate the gradient information of the agent's policy and update the agent's neural network parameters. Additionally, the neural network parameters of the target network are softly updated in the main network. Furthermore, the learning rate needs to be reduced in the later stages of learning to allow the agent to converge to a stable, optimal policy.

[0168] Based on the same inventive concept, another embodiment of the present invention provides a new energy grid forward dispatching device, which corresponds to the method of the aforementioned embodiment, and the device includes:

[0169] The modeling unit is used to consider extreme ramping events of new energy sources, and uncertainty is represented by chance constraint form to establish a two-layer robust optimization model for look-ahead scheduling.

[0170] The constraint transformation unit is used to decompose opportunity constraints into deterministic power flow based on day-ahead scheduling plans and uncertain power flow dominated by new energy power fluctuations, and to use the cumulative probability density function of slack variables to characterize the boundary of uncertain power flow, thereby transforming opportunity constraints into deterministic constraints.

[0171] The solution unit is used to establish the constrained Markov decision process of the two-layer robust optimization model, establish the state space, action space, reward function, risk function and risk threshold of the agent, and solve the constrained Markov decision process using the constrained reinforcement learning method.

[0172] The building unit is used to construct a risk-avoidance constrained reinforcement learning method to schedule an agent based on the SAC algorithm and the Lagrange function as the algorithm framework, establish an evaluator network based on a truncated Gaussian distribution, and construct an actor network that takes into account the worst scenarios.

[0173] The training and update unit is used to build an offline simulation environment and pre-train the actor network using imitation learning techniques. Every fixed time step, a small batch of trajectories is randomly sampled from the experience pool to calculate the gradient information of the actor network and the evaluator network and update the network parameters.

[0174] Other embodiments of the invention will readily occur to those skilled in the art upon consideration of the specification and practice of the embodiments disclosed herein. This invention is intended to cover any variations, uses, or adaptations of the invention that follow the general principles of the invention and include common knowledge or customary techniques in the art not disclosed herein. It should be understood that the invention is not limited to the precise structures described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of the invention is limited only by the appended claims.

Claims

1. A forward-looking dispatching method for new energy power grids, characterized in that, include: Considering extreme ramping events in new energy sources, uncertainty is represented by chance constraints, and a two-layer robust optimization model for look-ahead scheduling is established. Opportunity constraints are decomposed into deterministic power flows based on day-ahead scheduling plans and uncertain power flows dominated by new energy power fluctuations. The cumulative probability density function of slack variables is used to characterize the boundary of uncertain power flows, thus transforming opportunity constraints into deterministic constraints. The constrained Markov decision process of the two-layer robust optimization model is established, and the state space, action space, reward function, risk function and risk threshold of the agent are established. The constrained reinforcement learning method is used to solve the constrained Markov decision process. We construct a risk-avoidance-based constraint reinforcement learning method for scheduling agents, using the SAC algorithm and Lagrange function as the algorithmic framework. We establish an evaluator network based on a truncated Gaussian distribution and construct an actor network that considers the worst-case scenario. An offline simulation environment is constructed, and the actor network is pre-trained using imitation learning technology. At fixed time steps, a small batch of trajectories is randomly sampled from the experience pool to calculate the gradient information of the actor network and the evaluator network, and to update the network parameters. In the two-layer robust optimization model, the first-layer max function determines the scenario with the largest prediction deviation, and the second-layer min function minimizes the scheduling cost and power curtailment cost of the scenario with the largest prediction deviation. The expression is as follows: (1) in, (2) (3) (4) (5) (6) (7) (8) (9) In the formula, C g and C C These represent the cost of power generation and the cost of curtailing renewable energy, respectively. , and These are respectively the set of generator sets, the set of new energy generator sets, and the set of nodes; , and These are the reference operating points of the generator set, new energy unit, and load at time t for the i-th node, respectively; This represents the power transfer distribution factor corresponding to the i-th node; Adjust the generator output of the i-th node at time t; This indicates the generator output adjustment of the i-th node during the k-th time period; and These represent the maximum uphill and downhill climbing capabilities of the generator set at time t for the i-th node, respectively. , and Let be the power deviation, actual power deviation, and curtailment value of the new energy source at time t for the i-th node; This represents the upper limit of new energy sources that the system can accept at time t for node i; Indicates the first l The power limit allowed to be transmitted on each line; E represents the expected value, Pr represents the probability, and α represents the confidence level of the probability; and Let represent the maximum and minimum allowable power generation of the i-th generator unit at time t, respectively.

2. The forward-looking dispatching method for new energy power grids according to claim 1, characterized in that, The opportunity constraint is: (10) In the formula, The system operating point after rolling correction, i.e. ; This is the factor by which the generating unit shares the impact of power fluctuations from new energy sources, i.e. .

3. The forward-looking dispatching method for new energy power grids according to claim 2, characterized in that, The deterministic constraint is: (11) In the formula, and These are the system power flow operating point and the generator set output operating point, respectively. , , , These are the positive and negative relaxation values ​​for line power flow constraints and generator capacity constraints, respectively. The uncertainty boundary of the relaxation quantity; , Let these represent the maximum and minimum allowable generating power of the i-th generator set, respectively. This represents the adjustment amount of the power generation of the i-th generator set.

4. The forward-looking dispatching method for new energy power grids according to claim 3, characterized in that, The state space includes: system operating point based on the day-ahead plan, unit ramp mileage, upper and lower limits of unit ramp capacity, upper and lower limits of unit capacity, upper and lower limits of line capacity, new energy power deviation value, and new energy regulation upper limit value; the action space is to respond to new energy power fluctuations by adjusting the generator output. The reward function is: (14) The risk function is: (15) In the formula, c1 and c2 represent unit cost coefficients; The risk threshold d is set based on the allowable slack of the constraint; Representing a constrained Markov decision process as a tuple ,in For state set, For action sets, For reward function, For the penalty function, For state transition function, This represents the initial state distribution.

5. The forward-looking dispatching method for new energy power grids according to claim 4, characterized in that, The method of using constrained reinforcement learning to solve constrained Markov decision processes includes: Represent the agent's policy as a parameter-dependent neural network. The long-term discount reward value and the long-term discount penalty value are represented as follows: (16) (17) in, It is a discount factor. Representation Strategy The experience trajectory of the agent, i.e. ; This represents all future time steps generated by policy π starting from the current moment. The expected value of the corresponding random variable; Let be a single-step reward function, describing the state s of the agent at time t. t Execute action a t And transition to state s t+1 Instant rewards from environmental feedback; Let be the single-step cost function, describing the agent's state s at time t. t Execute action a t And transition to state s t+1 The immediate cost of environmental feedback; The strategy of secure reinforcement learning updates the parameters of the actor neural network. Therefore, the mathematical form of its update process is as follows: (18)。 6. The forward-looking dispatching method for new energy power grids according to claim 5, characterized in that, The construction of a risk-avoidance-based constrained reinforcement learning method for scheduling agents, using the SAC algorithm and Lagrange function as the algorithmic framework, and the establishment of an evaluator network based on a truncated Gaussian distribution, includes: Based on the SAC algorithm, a reward evaluator and a risk evaluator are constructed, and an adaptive safety weight is used to balance the reward value and safety. The optimization objective is: (19) In the formula, To optimize the optimal strategy, For Lagrange multipliers, The optimal Lagrange multiplier; L represents the Lagrange function; Let Q be the payoff function under strategy π. The cost Q function under strategy π; Let H0 represent the policy function, which describes the conditional probability distribution of the agent choosing action a in state s; H0 is the baseline entropy. Update the original strategy through alternating iterations. and dual variables , To achieve strategy optimization, the loss function is: (20) In the formula, Represents a projection onto the dual space; The neural network parameters for the reward evaluator and risk evaluator of the SAC algorithm are as follows: and Both are updated in the same way; taking the reward evaluator as an example, the loss function is: (21) In the formula, Indicates based on parameters The Q-function of the return; Indicates based on target network parameters The Q-function of the return; The loss function of the actuator is expressed as: (22) Let the reward evaluator represent the network that estimates the long-term reward of an action, and the risk evaluator represent the network that estimates the long-term risk of an action; the strategy... The probability distribution of the risk function is denoted as a Gaussian distribution. Approximate the fit using a truncated Gaussian distribution. , is represented as: (23) Based on the Bellman optimality equation, the Q-function and value function are estimated, namely: (24) (25) In the formula, Let represent the policy function, which describes the conditional probability distribution of the agent choosing action a' in state s'; Indicating in strategy p π Under the condition that the agent performs action a in state s, the cost squared C 2 The expected value of the condition; The cutoff threshold representing the cost distribution; To learn the risk assessor, its loss function is estimated by calculating the second-order Wasserstein distance, so that the estimated risk distribution can better match the actual risk distribution; the loss function of the risk assessor is: (26) in, The calculation function representing the longitudinal error of the distribution shape; The objective cost Q-function represents the teacher's policy μ, and describes the agent's state s at time t. t Next, execute action a t At that time, the expected long-term cumulative discount cost under teacher strategy μ; The Q-function represents the current cost corresponding to the teacher's policy μ, and describes the agent's state s at time t. t Next, execute action a t Real-time cost-value estimation under teacher strategy μ.

7. The forward-looking dispatching method for new energy power grids according to claim 6, characterized in that, The construction of the actuator network, which considers the worst-case scenarios, includes: Describing the confidence level of not exceeding the opportunity constraint as the agent's risk aversion level, we define agent policy security based on CVaR, where the risk value of the agent policy is: (27) In the formula, Represents random variable C π α-quantile; An evaluator network based on a truncated Gaussian distribution estimates the CVaR value at each time step and replaces the long-term expected risk value with a new risk metric. The formula for calculating the risk metric of the agent's actions in each iteration is: (28) in, and Let represent the probability density function and cumulative distribution function of the standard normal distribution, respectively; The agent's policy optimization must meet a risk threshold, namely: Following the policy update formula of the SAC algorithm, and taking policy safety into account, the KL divergence is written as: (29) In the formula, ; It is the partition function of the standardized distribution, which affects the parameters of the neural network. No effect; the loss function of the actioner network is written as: (30) Based on risk metrics dual variables The update formula is revised as follows: (31)。 8. The forward-looking dispatch method for new energy power grids according to claim 7, characterized in that, The offline simulation environment is constructed, and the actor network is pre-trained using imitation learning technology. At fixed time steps, a small batch of trajectories is randomly sampled from the experience pool to calculate the gradient information of the actor network and the evaluator network, and to update the network parameters, including: Based on the PandaPower power grid offline simulation environment, day-ahead scheduling plans for various scenarios are constructed. Mathematical optimization methods are used to solve a two-layer robust optimization model to obtain prior knowledge. After obtaining the prior knowledge trajectory, the state and actions are reconstructed into features and labels, and a loss function is designed as follows: (32) In the formula, The training set represents the prior knowledge; This indicates the number of samples taken from the training set; This represents the desired action output by the action network; At fixed time steps, a small batch of experience trajectories is randomly selected from the trajectory experience pool of the agent in the stored scenario. This is used to calculate the gradient information of the agent's policy and update the neural network parameters of the agent. The neural network parameters of the target network are also softly updated in the main network. At the same time, the learning rate of the agent is reduced in the later stage of learning so that it can converge to a stable optimal policy.

9. A forward-looking dispatching device for a new energy power grid, characterized in that, The apparatus for implementing the forward-looking dispatching method for new energy power grids as described in any one of claims 1-8 includes: The modeling unit is used to consider extreme ramping events of new energy sources, and uncertainty is represented by chance constraints to establish a two-layer robust optimization model for look-ahead scheduling. The constraint transformation unit is used to decompose opportunity constraints into deterministic power flow based on day-ahead scheduling plans and uncertain power flow dominated by new energy power fluctuations, and to use the cumulative probability density function of slack variables to characterize the boundary of uncertain power flow, thereby transforming opportunity constraints into deterministic constraints. The solution unit is used to establish the constrained Markov decision process of the two-layer robust optimization model, establish the state space, action space, reward function, risk function and risk threshold of the agent, and solve the constrained Markov decision process using the constrained reinforcement learning method. The building unit is used to construct a risk-avoidance constrained reinforcement learning method to schedule an agent based on the SAC algorithm and the Lagrange function as the algorithm framework, establish an evaluator network based on a truncated Gaussian distribution, and construct an actor network that takes into account the worst scenarios. The training and update unit is used to build an offline simulation environment and pre-train the actor network using imitation learning techniques. Every fixed time step, a small batch of trajectories is randomly sampled from the experience pool to calculate the gradient information of the actor network and the evaluator network and update the network parameters.

Citation Information

Patent Citations

  • CN118396367A