Real-time scheduling method and system for high-proportion new energy power system under source-storage cooperation
Through the multi-agent deep deterministic strategy gradient method guided by the static security domain, coordinated scheduling of source and storage is realized, the power imbalance caused by fluctuations in new energy output is solved, and the system safety and economy are improved.
Patent Information
- Application Number
- CN202510373705.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-27
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-03-27
AI Technical Summary
The traditional deterministic scheduling model based on the calculation and scheduling plan online matching based on typical operating modes is difficult to cope with the power imbalance and wind and light abandonment problems caused by the violent fluctuations in new energy output, affecting the safety and economic operation of the system.
A multi-agent deep deterministic strategy gradient method (SSRG-MADDPG) guided by static security domain is proposed. Through source storage collaborative scheduling and combined with static security domain model, the real-time scheduling of new energy power systems is optimized to improve the security and economical decision-making.
Effectively respond to source load uncertainty, reduce system operation costs, improve system safety and decision-making efficiency, and avoid the risk of overloading trends.
Smart Images

Figure CN119995046A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a real-time dispatching method for a high-proportion new energy power system under source-storage coordination, and belongs to the field of power grid operation analysis and dispatching. Background Art
[0002] With the gradual advancement of the construction of new power systems, the proportion of renewable energy output has gradually increased, and the uncertainty of the power system has become more prominent. The traditional deterministic dispatching mode based on online matching of typical operation mode calculation and dispatching plan is difficult to cope with the sharp fluctuations in renewable energy output in a short period of time. The resulting power imbalance and wind and solar abandonment problems have seriously affected the safety and economic operation of the system. Therefore, it is necessary to further revise the day-ahead dispatching plan in combination with the ultra-short-term forecast value of source and load in the real-time dispatching stage to eliminate the risk of system flow exceeding the limit as much as possible.
[0003] Real-time dispatch of power systems is essentially an uncertainty optimization problem, that is, given the load and new energy forecast values, considering the uncertainty of source and load, solving the output plan of various types of generators and flexible resources such as energy storage to meet various safety constraints of the system and minimize the system operation cost. This type of problem has the characteristics of non-convexity, nonlinearity, and high dimension, and is difficult to solve. At present, the solution methods for system source and load uncertainty mainly include stochastic optimization (SO) taking into account random variables, robust optimization (RO) taking into account the range of uncertain variables, and methods based on artificial intelligence (AI).
[0004] The stochastic optimization method describes the uncertainty of variables by assuming that random variables obey a certain probability distribution. The robust optimization algorithm is a method that uses interval disturbance information to solve the optimal strategy under the worst disturbance conditions. Compared with the stochastic optimization algorithm, robust optimization does not need to know the specific probability distribution form of the random variables, but only focuses on the fluctuation boundary values of the uncertain variables. It cannot be ignored that the above algorithms all need to simplify the nonlinear power flow equations and constraints of the power system. This process may lose key system information and cannot accurately reflect the actual operation of the system. To this end, some scholars have tried to solve the optimization scheduling model based on heuristic algorithms such as genetic algorithms and particle swarm algorithms. However, this type of algorithm has low solution efficiency and is prone to falling into local optimal solutions when solving high-dimensional problems. Summary of the invention
[0005] Based on the above analysis, the present invention proposes a Steady-state Security Region Guided Multi-agent Deep Deterministic Policy Gradient (SSRG-MADDPG), which uses source-storage collaborative scheduling to cope with the impact of source-load uncertainty on the system in the real-time scheduling stage, thereby improving the safety and economy of model decision-making. Specifically, based on the day-ahead unit output plan, energy storage resources are considered, and the corresponding controlled thermal power units are selected to perform online solution of the real-time scheduling strategy of the new energy power system under source-storage collaboration. Considering the growing scale of the power grid, the huge order of magnitude of the system state quantity and controllable resources, and the single-agent reinforcement learning facing the dimensionality curse problem, the real-time scheduling problem is modeled as a decentralized partially observed Markov decision process (DPOMDP), and a new energy power system based on SSRG-MADDPG is proposed. Each controlled unit is modeled as an agent, and the corresponding observation space and action space are designed. Considering the system operation goals, the cost of the controlled units and energy storage scheduling, the model reward function is designed. Furthermore, combined with the unit's day-ahead output plan, the ultra-short-term forecast value of source and load, and the prior distribution of source and load forecast error, the system operating point at the next moment is sampled, and the system operating state at the next moment is predicted based on the solved static safety domain boundary model. In the multi-agent model training stage, the state prediction information is used as the model observation to provide a reference for model decision solving. In the multi-agent model deployment and application stage, the state prediction information is used as a hard constraint on the model action output to shield unnecessary actions, thereby reducing the system operating cost and improving the system safety.
[0006] The object of the present invention is to provide a real-time dispatching method for a high-proportion new energy power system under source-storage coordination, comprising:
[0007] Collect and organize source-load operation scenario data and perform pre-processing;
[0008] Combined with the source-load day-ahead forecast scenario data, solve the day-ahead dispatch plan of the unit;
[0009] Markov decision process modeling of real-time scheduling processes;
[0010] Solve the static security domain model;
[0011] Combining the solved static safety domain model, the agent observation space, action space and reward function in Markov decision process modeling, and the solved unit day-ahead dispatch plan, a multi-agent deep deterministic policy gradient model guided by the static safety domain is constructed. The input of the policy gradient model is the system state at each time step under each operating scenario: including new energy output, load value, unit day-ahead output plan value, and line load rate; the output is the action after the observed state, which is used to guide the dispatch of the controlled unit output; the model is trained using source-load operation scenario data, and the trained policy gradient model is used to realize real-time dispatch of the new energy power system.
[0012] Furthermore, collecting and organizing source-load operation scenario data includes: setting an N-node power system, node 0 is the reference node, and N G Thermal power units, N R Renewable energy units, N ESS Energy storage devices, N L load nodes; a total of N S The group source-load operation scenario, the single day-ahead real scenario data is represented by S rq = {P r,rq ,P l,rq}, of which new energy output data Load data The time interval of the day-ahead scenario is 15 minutes, so a single day-ahead scenario has 96 steps in total; the data of a single day-ahead forecast scenario is represented by S rq,f = {P r,rq,f ,P l,rq,f}, The real scene data in a single day can be expressed as S rn = {P r,rn ,P l,rn}, The time interval of the intraday scenario is 5 minutes, so there are 288 steps in the time series of a single intraday scenario; the prediction scenario data of a single day is represented by S rn,f = {P r,rn,f ,P l,rn,f}, Preprocessing includes supplementing missing values and removing outliers in source-load operation scenario data.
[0013] Furthermore, the solution of the unit day-ahead dispatch plan includes:
[0014] Set a single day-ahead forecast scenario data scenario S rq,f The corresponding unit day-ahead dispatch plan is denoted as P g,rq , N GRepresents the number of thermal power units, approximates the system AC power flow equation to a second-order cone form, and takes minimizing the system power generation cost as the goal. The mathematical model of the unit day-ahead scheduling plan is as follows:
[0015] (1) Objective function
[0016]
[0017] Where, T DA is the step length of the day-ahead scheduling plan. The time interval of the day-ahead scheduling plan is 15 minutes, that is, T DA =96, N G is the number of generators in the system, P G,i,t is the output value of the ith generator at time t, a i , b i 、c i is the cost coefficient of the ith generator;
[0018] (2) Power balance constraint at node j
[0019]
[0020] Where P G,k,t , Q G,k,t are the active and reactive outputs of generator set k at time t, respectively, k∈j means that generator set k is connected to the grid through node j; P b,ij,t and Q b,ij,t are the active and reactive powers flowing from node i to node j on branch i→j at time t; P b,je,t and Q b,je,t are the active and reactive powers flowing from node j to node e on branch j→e at time t; P L,j,t , Q L,j,t are the active and reactive loads of node j at time t, respectively; r ij and x ij are the resistance and reactance of branch i→j respectively, I ij,t is the square of the current magnitude flowing through branch i→j at time t;
[0021] (3) Voltage amplitude constraints at both ends of lines i and j
[0022]
[0023] Where U i,t , U j,t are the voltage amplitudes of nodes i and j at time t respectively;
[0024] (4) Power constraints of lines i and j
[0025]
[0026] (5) Line end node voltage phase angle constraint
[0027]
[0028] In the formula, and They represent variables that consider the voltage amplitude of nodes i and j to be constant, rather than optimization variables, and are taken as 1 here; θ i,t ,θ j,t are the phase angles of nodes i and j at time t respectively;
[0029] (6) Generator output constraints
[0030]
[0031] In the formula, is the minimum and maximum output value of the i-th generator set, is the day-ahead planned value of the i-th generator unit at time t;
[0032] (7) Generator ramp rate constraint
[0033]
[0034] In the formula, is the maximum ramp power of the ith generator set per unit time period;
[0035] (8) Branch transmission power constraints
[0036]
[0037] In the formula, is the maximum transmission power of lines i and j;
[0038] Therefore, the mathematical model of the unit day-ahead scheduling plan is written as:
[0039]
[0040] Finally, based on the above optimization model of Gurobi optimization solver and combined with the day-ahead source-load forecast value in the scenario data, the day-ahead output plan of the unit is solved.
[0041] Furthermore, the process of Markov decision process modeling is as follows:
[0042] Agent i: Agent i is deployed at the node where the controlled generator set is located or the node where the energy storage device is located, so as to correct the output plan of the node device in the real-time scheduling stage in combination with the intraday forecast value of the source and load;
[0043] Status t:Indicators reflecting the operating status of the system at time t, including the actual output value of the thermal power unit at time t and the planned output value at time t+1, the actual value of the new energy unit and each node load at time t and the intraday forecast value at time t+1, the real-time SoC value of the energy storage equipment and the load rate of each line;
[0044] Observation value o i,t : The system state value that each agent can observe is set as the state information of the node where the agent is located and its K-order neighbor nodes;
[0045] Action a i,t :The scheduling actions that the intelligent agent can take; using discrete action space, for the controlled thermal power units, the climbing power of each unit per unit time is Normalized to [-1,1] and then N Pr Equally divided, -1 means the unit reduces output 1 means the unit increases output For energy storage equipment, the maximum charge and discharge power per unit time of each unit is Normalized to [-1,1] and then N Pr Equally divided, -1 means the charging power of the energy storage device is 1 means the discharge power of the energy storage device is Assume there are n agents in total. At time t, the environment receives the joint action A of the agents. t =[a 1,t ,a 2,t ,...,a n,t ];
[0046] Reward: The environment performs the joint action A of the agent t Finally, the reward is calculated through the set reward function and fed back to the multi-agent system.
[0047] Furthermore, the reward function is designed as follows:
[0048] 1) If the following conditions exist, the scenario will be terminated and a penalty r will be imposed. d :a) The system power flow does not converge; b) There are one or more lines with a load rate ρ>2; c) There are two or more lines with a load rate ρ>1; d) The balancing machine output exceeds the limit, r d The calculation formula is as follows:
[0049]
[0050] In the formula, r d,base is the baseline penalty for not completing the round, T d is the total step length of the round, T down is the current running step of the system, r d,minIt is the minimum penalty given for an incomplete round, to avoid the incomplete penalty given in the later stage being too small, causing the agent to tend to not act;
[0051] 2) If none of the above conditions exists, calculate the penalty r for overloaded and overloaded lines. L,t :
[0052]
[0053] In the formula, ρ l,t is the load rate of line l at time t, are the penalty coefficients for heavy-loaded lines and overloaded lines, respectively. If the line load rate exceeds the first threshold, the line is an overloaded line. If the load rate is between the second threshold and the first threshold, the line is a heavy-loaded line. When the system has overloaded or overloaded lines, a certain penalty should be given to guide the intelligent agent to make corresponding decisions and optimize the system operation status. Br is the total number of system lines;
[0054] 3) Give system survival and operation status rewards r B,t :
[0055]
[0056] In the formula, r b is the basic reward for system survival, λ r,b is the reward for the system running step length t, which is used to guide the system to extend the running step length as much as possible, λ l,1 , l,2 , l,3 is the incentive factor for reducing the maximum load rate of the system line, ρ l,max is the maximum load rate of the system;
[0057] 4) Calculate the operating costs of the controlled generator sets and energy storage equipment:
[0058]
[0059] Where N ctg is the number of controlled units, N stg is the number of energy storage devices in the system, is the adjustment cost coefficient of generator g, ΔP ctg,g,t is the absolute value of the output regulation of the generator g, I(·) is the indicative function, and the number of times the energy storage device operates is counted. is the action cost coefficient of energy storage device s;
[0060] 5) Give the system a reward for running the entire scenario smoothly C,t ;
[0061] Therefore, the system reward at time t is:
[0062]
[0063] Then the cumulative reward value of agent i at any time k is:
[0064]
[0065] Where γ is the reward discount factor.
[0066] Furthermore, the static safety domain of the new energy power system is defined as the set of points that satisfy the power flow equation and the operational safety constraints, expressed as formula (16), where the new energy units and thermal power units are uniformly modeled as generator nodes, with the installed capacity as the maximum output upper limit;
[0067]
[0068] In the formula, according to different constraints, it is divided into different types of static security domains, Ω U The static voltage safety domain that satisfies the voltage constraint; Ω P The static generator active output safety domain that satisfies the generator active output constraint; Ω Q The static generator reactive output safety domain that satisfies the generator reactive output constraint; To meet the branch thermal stability safety domain of the system branch transmission power; Ω ss For the entire static security domain, Ω U ,Ω P ,Ω Q , Intersection composition; y SR is the system node voltage, phase vector; x SR is the node power injection vector; φ(x SR ,y SR )=0 is the AC system power flow equation; N is the system bus node set; N G is the set of system generator nodes; are the upper and lower limits of the voltage at the system node i respectively; They are the upper and lower limits of active output of system generator node i respectively; are the upper and lower limits of reactive power output of system generator node i; P b,ij , -P b,ij are the forward and reverse transmission power limits of the connection nodes i and j in the system, respectively;
[0069] In the case of reactive power approximation local balance, only the static safety domain under active power injection is considered, that is:
[0070]
[0071] In the formula, x SR =[xP ,x Q ]; x P Active power injection for each node; x Q = C is the reactive power injection of each node, C is a constant; P G,i is the active output of the i-th generator; P L,j is the active load value of the jth node.
[0072] Furthermore, the overall processing process of the multi-agent deep deterministic policy gradient model guided by the static safety domain is as follows:
[0073] The intraday forecast scenario and real scenario are randomly selected from the source-load operation scenario set. At T=1, the actual data at the first time point in the intraday real scenario and the data at the first time point in the unit's day-ahead dispatch plan are loaded to perform AC power flow calculation to obtain the current state of the system. Then, the system state value at time T, the source-load intraday forecast value at time T+1, and the system state judgment result at time T+1 are provided to the intelligent agent as observation data. The intelligent agent makes decisions based on the observation data and outputs decision actions. After the action legitimacy is verified, the decision actions are provided to the power system computing environment for power flow calculation to obtain the system state at time T+1. At this time, the reward value is calculated based on the intelligent agent's actions and the current system state information, which serves as feedback information to guide the intelligent agent to evolve its strategy.
[0074] Furthermore, the judgment result of the system state at time T+1 is obtained by solving the static safety domain boundary, that is, combining the unit's day-ahead output value at time T+1 and the source load intra-day forecast value at time T+1, according to the source load forecast error, combined with Latin hypercube sampling, the possible source load intra-day output value at the next moment is extracted, and the day-ahead output value of the unit together constitutes the system operation point set at time T+1, and the static safety domain boundary is used for online judgment, and the result is provided to the intelligent agent as observation information; specifically, the ultra-short-term forecast value of the source load at the next moment is P R,f , P L,f , the current output value of the controlled unit in the system P g,sk , the planned output value P of the next moment of the non-controlled unit in the system g,fsk ; Assume that the distribution of source-load ultra-short-term prediction error conforms to the normal distribution, based on LHS sampling N err Group error P R,err , P L,err , construct the next moment N err The observation x of the static safety domain model is constructed by combining the unit output with the source-load operation scenario value. p = {P g,sk ,P g,fsk ,P R,f +P R,err ,P L,f +P L,err}, the static security domain model is used to determine the security distribution of the source-load operation scenario at the next moment, and the static security domain model is set to determine N err Group Scene N' err The samples are safe. Provided to the agent as observation features at time T;
[0075] when , it indicates that the system is safe at the next moment and no additional adjustment is required. Otherwise, it reminds the agent to adjust the controlled resources. S is the set threshold.
[0076] Furthermore, the training process of the multi-agent deep deterministic policy gradient model guided by the static safety domain is as follows:
[0077] First, initialize the power grid computing environment, training data set, agent parameters, experience pool, and set the number of training rounds to episode = m;
[0078] Then, within the set number of rounds, a scene is randomly selected each time. At T = 0, the initial state x is calculated by combining the new energy output, load value, and unit planned output value. Then, the static safety domain model is used to predict the system state safety information at time T + 1, and the observation value o is formed together with x. s ;
[0079] Then, in the scenario step, for each agent i, from the global observation o s Constructing its own observation i , then for the observed value o i , select action a i =μ θ,i (o i )+ε,ε~N(0,σ 2 I), μ θ,i (o i ) represents the policy network, ε~N(0,σ 2 I) represents the random perturbation of the policy network; the actions are discretized in combination with the Gumbel distribution, and the compliance of the actions is checked. Then, the joint actions of each agent are executed in the power grid computing environment. a=[a1,a2,...,a n ], get instant rewards r , the next state of the system is x′, and the next state information is predicted based on the static safety domain model, which together with x′ constitutes the observation value o′ s The round ends with signal d, (x, a, x′, r, d) is stored in the experience pool D, and the system state x←x′ is updated; for the update of each agent, a batch of data B={(x j ,a j ,x′j ,r j ,d j )},x j ,a j ,x′ j ,r j ,d j They are the state, action, updated state after executing the action, reward value and round terminator of the jth sample respectively; finally, the data is updated until the number of training rounds M is reached or the reward converges.
[0080] The present invention also provides a real-time dispatching system for a high-proportion new energy power system under source-storage coordination, comprising:
[0081] A processor and a memory, the memory is used to store program instructions, and the processor is used to call the stored instructions in the memory to execute a real-time scheduling method for a high-proportion new energy power system under source-storage coordination as described in the above technical solution.
[0082] Compared with the prior art, the present invention has the following beneficial effects:
[0083] Compared with existing reinforcement learning methods, the proposed method integrates the safety domain model into the system training and deployment application process, improving the intelligent agent's state perception and prediction ability of the system uncertainty. Especially in the deployment decision-making process, the proposed model improves system security and reduces system operation costs. Compared with the model-driven method, the proposed method improves system security and model online solution efficiency, and the operation cost is similar. BRIEF DESCRIPTION OF THE DRAWINGS
[0084] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the structures shown in these drawings without paying creative work.
[0085] Figure 1 It is a schematic diagram of the process of the present invention.
[0086] Figure 2 The real-time dispatching technical framework of high-proportion new energy power system based on SSRG-MADDPG designed for the present invention.
[0087] Figure 3 Pseudo code for the SSRG-MADDPG algorithm. DETAILED DESCRIPTION
[0088] The technical solution of the present invention is further described in detail below in conjunction with the accompanying drawings and embodiments.
[0089] like Figure 1 As shown, the present invention provides a real-time dispatching method for a high-proportion new energy power system under source-storage coordination, comprising the following steps:
[0090] 1. Collect and organize source-load operation scenario data;
[0091] Assume an N-node power system, node 0 is the reference node, and N G Thermal power units, N R Renewable energy units, N ESS Energy storage devices, N L A total of N load nodes are collected. S The group source-load operation scenario, the single day-ahead real scenario data can be expressed as S rq = {P r,rq ,P l,rq}, of which new energy output data Load data The time interval of the day-ahead scenario is 15 minutes, so a single day-ahead scenario has 96 steps in total. The data of a single day-ahead forecast scenario can be expressed as S rq,f = {P r,rq,f ,P l,rq,f}, The real scene data in a single day can be expressed as S rn = {P r,rn ,P l,rn}, The time interval of the intraday scenario is 5 minutes, so there are 288 steps in the time series of a single intraday scenario. The forecast scenario data of a single day can be expressed as S rn,f = {P r,rn,f ,P l,rn,f}, The subscript r in the above symbols represents new energy, l represents load, rq represents day-ahead, rn represents intraday, and f represents forecast. The above scenario data are collected and sorted, and data processing such as missing value supplementation and outlier removal is performed.
[0092] 2. Solve the day-ahead dispatch plan of the unit;
[0093] Combined with N S S rq,f Day-ahead forecast scenario, solving the output plan of conventional thermal power units, single scenario S rq,f The corresponding unit day-ahead dispatch plan is denoted as P g,rq , The mathematical model for solving the unit day-ahead scheduling plan is as follows:
[0094] The system AC power flow equation is approximated as a second-order cone form. In this stage, the participation of energy storage resources is not considered, and the goal is to minimize the system power generation cost. The mathematical model is as follows:
[0095] (1) Objective function
[0096]
[0097] Where, T DA is the step length of the day-ahead scheduling plan. The time interval of the day-ahead scheduling plan is 15 minutes, that is, T DA =96, N G is the number of generators in the system, P G,i,t is the output value of the ith generator at time t, a i 、b i 、c i is the cost coefficient of the ith generator.
[0098] (2) Power balance constraint at node j
[0099]
[0100] Where P G,k,t , Q G,k,t are the active and reactive outputs of generator set k at time t, respectively, k∈j means that generator set k is connected to the grid through node j; P b,ij,t and Q b,ij,t are the active and reactive powers flowing from node i to node j on branch i→j at time t; P b,je,t and Q b,je,t are the active and reactive powers flowing from node j to node e on branch j→e at time t; P L,j,t , Q L,j,t are the active and reactive loads of node j at time t, respectively; r ij and x ij are the resistance and reactance of branch i→j respectively, I ij,t is the square of the current magnitude flowing through branch i→j at time t.
[0101] (3) Voltage amplitude constraints at both ends of lines i and j
[0102]
[0103] Where U i,t , U j,t are the voltage amplitudes of nodes i and j at time t respectively.
[0104] (4) Power constraints of lines i and j
[0105]
[0106] (5) Line end node voltage phase angle constraint
[0107]
[0108] In the formula, and They represent variables that consider the voltage amplitude of nodes i and j to be constant, rather than optimization variables, and are taken as 1 here; θ i,t ,θ j,t are the phase angles of nodes i and j at time t respectively.
[0109] (6) Generator output constraints
[0110]
[0111] In the formula, is the minimum and maximum output value of the i-th generator set, is the day-ahead planned value for the ith generator at time t.
[0112] (7) Generator ramp rate constraint
[0113]
[0114] In the formula, It is the maximum ramp power of the ith generator set per unit time period.
[0115] (8) Branch transmission power constraints
[0116]
[0117] In the formula, is the maximum transmission power of lines i and j.
[0118] Therefore, the mathematical model of the unit's day-ahead scheduling plan can be written as:
[0119] min formula (1) (9)
[0120] st type (2)-(8)
[0121] Finally, the above optimization model is solved based on the Gurobi optimization solver, that is, the day-ahead forecast value of the source and load in the scenario data is combined, and the formula (9) is solved based on the Gurobi optimization solver to achieve the solution of the day-ahead output plan of the unit.
[0122] 3. Markov decision process modeling of real-time scheduling process;
[0123] Combined with Multi Agent Reinforcement Learning (MARL), the real-time dispatching problem of high-proportion renewable energy power system under source-storage coordination is solved. Each controlled unit and energy storage device is modeled as a single agent to achieve the coordinated dispatching of multiple agents in the region. In MARL, each agent follows the basic paradigm of reinforcement learning and learns its own strategy based on the decentralized partial observation Markov decision process DPOMDP. The main components of this process are:
[0124] (1) Agent i: Agent i is deployed at the node where the controlled generator set is located or the node where the energy storage device is located to correct the output plan of the node device in the real-time scheduling stage based on the intraday forecast value of the source and load.
[0125] (2) State s t : An indicator reflecting the operating status of the system at time t, including the actual output value of the thermal power unit at time t and the actual output value at time t+1
[0126] The planned output value at the moment, the actual value of the new energy unit and each node load at the moment t and the intraday forecast value at the moment t+1, the real-time SoC value of the energy storage equipment and the load rate of each line, etc.
[0127] (3) Observation value o i,t : The system state value that each agent can observe is set as the state information of the node where the agent is located and its K-order neighbor nodes;
[0128] (4) Action a i,t :The scheduling actions that the agent can take; using discrete action space, for the controlled thermal power units, the climbing power per unit time of each unit is Normalized to [-1,1] and then N Pr Equally divided, -1 means the unit reduces output 1 means the unit increases output For energy storage equipment, the maximum charge and discharge power per unit time of each unit is Normalized to [-1,1] and then N Pr Equally divided, -1 means the charging power of the energy storage device is 1 means the discharge power of the energy storage device is Assume there are n agents in total. At time t, the environment receives the joint action A of the agents. t =[a 1,t ,a 2,t ,...,a n,t ].
[0129] (5) Reward: The environment executes the joint action A of the agent t Finally, the reward is calculated through the set reward function and fed back to the multi-agent system.
[0130] The goal of setting the reward function is to allow the agent to balance the system power distribution, eliminate the risk of system power flow exceeding the limit, and avoid power flow non-convergence according to the real-time operating status of the system and the system's intraday forecast value. Therefore, in view of the above goals, the reward function is designed as follows:
[0131] 1) If the following conditions exist, the scenario will be terminated and a penalty r will be imposed. d :a) The system power flow does not converge; b) There are one or more lines with a load rate of ρ>2; c) There are two or more lines with a load rate of ρ>1; d) The output of the balancing machine exceeds the limit. d The calculation formula is as follows:
[0132]
[0133] In the formula, r d,base is the baseline penalty for not completing the round, T d is the total step length of the round, T down is the current running step of the system. As the running step of the system increases, the penalty for the system termination scenario becomes smaller and smaller. d,min It is the minimum penalty given for an incomplete round, to avoid the incomplete penalty given in the later stage being too small, causing the agent to tend to not act.
[0134] 2) If none of the above conditions exists, calculate the penalty r for overloaded and overloaded lines. L,t :
[0135]
[0136] In the formula, ρ l,t is the load rate of line l at time t, are the penalty coefficients for heavy-loaded lines and overloaded lines, respectively. If the line load rate exceeds 1, the line is an overloaded line, and if the load rate is between 0.95-1, the line is a heavy-loaded line. When the system has overloaded or overloaded lines, a certain penalty needs to be given to guide the intelligent agent to make corresponding decisions and optimize the system operation status. Br is the total number of system lines.
[0137] 3) To further avoid system operation risks, optimize system operation indicators, and guide the system to survive as many steps as possible, give the system survival and operation status rewards, r B,t :
[0138]
[0139] In the formula, r b is the basic reward for system survival, λ r,bis the reward for the system running step length t, which is used to guide the system to extend the running step length as much as possible, λ l,1 , l,2 , l,3 is the incentive factor for reducing the maximum load rate of the system line, ρ l,max is the maximum load rate of the system.
[0140] 4) Calculate the operating costs of the controlled generator sets and energy storage equipment:
[0141]
[0142] Where N ctg is the number of controlled units, N stg is the number of energy storage devices in the system, is the adjustment cost coefficient of generator g, ΔP ctg,g,t is the absolute value of the output regulation of the generator g, I(·) is the indicative function, and the number of times the energy storage device operates is counted. is the action cost coefficient of energy storage device s.
[0143] 5) Give the system a reward for running the entire scenario smoothly C,t .
[0144] Therefore, the system reward at time t is:
[0145]
[0146] Then the cumulative reward value of agent i at any time k is:
[0147]
[0148] Where γ is the reward discount factor.
[0149] 4. Solution of static security domain model
[0150] The static safety domain of the renewable energy power system is defined as the set of points that satisfy the power flow equation and operational safety constraints, which can be expressed as formula (16), where the renewable energy units and thermal power units are uniformly modeled as generator nodes, with the installed capacity as the maximum output upper limit.
[0151]
[0152] In the formula, according to different constraints, it is divided into different types of static security domains, Ω U The static voltage safety domain that satisfies the voltage constraint; Ω P The static generator active output safety domain that satisfies the generator active output constraint; Ω Q The static generator reactive output safety domain that satisfies the generator reactive output constraint; To meet the branch thermal stability safety domain of the system branch transmission power; Ω ss For the entire static security domain, Ω U ,Ω P ,Ω Q , Intersection composition; y SR is the system node voltage, phase vector; x SR is the node power injection vector; φ(x SR ,y SR )=0 is the AC system power flow equation; N is the system bus node set; N G is the set of system generator nodes; are the upper and lower limits of the voltage at the system node i respectively; They are the upper and lower limits of active output of system generator node i respectively; are the upper and lower limits of reactive power output of system generator node i; P b,ij , -P b,ij are the forward and reverse transmission power limits of the connection nodes i and j in the system respectively.
[0153] In the case of reactive power approximation local balance, only the static safety domain under active power injection can be considered, that is:
[0154]
[0155] In the formula, x SR =[x P ,x Q ]; x P Active power injection for each node; x Q = C is the reactive power injection of each node, C is a constant; P G,i is the active output of the i-th generator; P L,j is the active load value of the jth node.
[0156] The content of the static safety domain research under active power injection is the set of operating points where the system unit output and node load demand meet various safety constraints under a specific reactive power configuration.
[0157] For the static safety domain under active power injection, under the given network topology and system component parameters, it is uniquely determined, connected, independent of the operating state, and has no internal voids. Therefore, if a certain operating point under active power injection is within the static safety domain of the system, when it slowly changes its active injection power in any direction in a quasi-steady state, it will definitely encounter the safety domain boundary. The area surrounded by these boundaries is the static safety domain. The solution of the static safety domain model SSR is a prior art and is not described in this invention. The solution result is used for the formulation of subsequent scheduling strategies.
[0158] 5. Training, debugging and application of multi-agent deep deterministic policy gradient models;
[0159] Aiming at the real-time dispatch problem of high-proportion renewable energy power system, a Steady-state Security Region Guided Multi-agent DeepDeterministic Policy Gradient (SSRG-MADDPG) method based on static security region guidance is proposed. Figure 2 First, randomly select the intraday forecast scenario S in the scenario set rn,f With real scene S rn , at T = 1, load S rn The actual data at the first time point in the day-ahead dispatch plan and the data at the first time point in the day-ahead dispatch plan are used to calculate the AC power flow and obtain the current state of the system. Then, the system state value at time T, the source-load short-term forecast value at time T+1, and the system state judgment result at time T+1 are provided to the intelligent agent as observation data. The intelligent agent makes decisions based on the observation data and outputs decision actions. After the action legitimacy is verified, the decision actions are provided to the power system computing environment for power flow calculation to obtain the system state at time T+1. At this time, the reward value is calculated based on the intelligent agent action and the current system state information, which serves as feedback information to guide the intelligent agent to evolve its strategy.
[0160] Attached Figure 2 In the above, the judgment result of the system state at time T+1 is obtained by solving the boundary of the static safety domain of the system. That is, combined with the unit's day-ahead output value at time T+1, the source-load ultra-short-term prediction value at time T+1, and considering the source-load prediction error, the Latin Hypercube sampling (LHS) is used to extract the possible source-load ultra-short-term output value at the next moment, which together with the unit's day-ahead output value constitutes the system operation point set at time T+1, and the SSR makes an online judgment, and the result is provided to the intelligent agent as observation information. Specifically, let the system source-load ultra-short-term prediction value at the next moment be P R,f , P L,f , the current output value of the controlled unit in the system P g,sk , the planned output value of the non-controlled units in the system at the next moment, P g,fsk Assuming that the distribution of source-load ultra-short-term forecast error conforms to the normal distribution, according to the LHS sampling N err Group error P R,err , P L,err , construct the next moment N err The group source and load operation scenario values are combined with the unit output to construct the observation x of the safety domain model p = {P g,sk ,Pg,fsk ,P R,f +P R,err ,P L,f +P L,err}, the static security domain model SSR is used to determine the security distribution of the source-load operation scenario at the next moment, and the static security domain model SSR is set to determine N err Group Scene N' err The samples are safe. Provided as observations to the agent as observation features at time T.
[0161] After the model training is completed, in the deployment and application stage, the judgment results of SSR are used as hard constraints for the action output of the intelligent agent, shielding some unnecessary actions to reduce the action cost of the model and improve the economic efficiency of system operation. , it indicates that the system is safe at the next moment and no additional adjustment is required. Otherwise, it reminds the agent that it needs to adjust the controlled resources. S is the set threshold.
[0162] The specific algorithm flow of SSRG-MADDPG is as follows Figure 3 shown.
[0163] The input of the multi-agent deep deterministic policy gradient model is the system state at each time step in each operating scenario (each scenario is 288 time steps, 96 on the day before, and will be expanded to 288): including new energy output, load value, unit output plan value on the day before, line load rate, etc.
[0164] The output is the action after observing the state, that is, how to dispatch the output of the controlled unit;
[0165] First, initialize the power grid computing environment, training data set, agent parameters, experience pool, and set the number of training rounds episode = m. Then, within the set number of rounds, randomly select a scene each time, and at T = 0, combine the new energy output, load value, and unit planned output value to calculate the initial state x, and then combine the SSR model to predict the system state safety information at time T + 1, and together with x, form the observation value o s Then, in this scenario step, for each agent i, from the global observation o s Constructing its own observation i , then for the observed value o i , select action a i =μ θ,i (o i )+ε,ε~N(0,σ 2 I), μ θ,i (o i ) represents the policy network, ε~N(0,σ 2I) represents applying random perturbations to the policy network for action exploration, discretizing the actions in combination with the Gumbel distribution, and performing action compliance verification, and then executing the joint actions of each agent in the power grid computing environment a=[a1,a2,...,a n ], get instant rewards r , the next state of the system is x', and the next state information is predicted based on SSR, which together with x' constitutes the observation value o' s The round ends with signal d. Store (x, a, x', r, d) into the experience pool D and update the system state x←x'. For the update of each agent's strategy, randomly extract a batch of data B from D = {(x j ,a j ,x' j ,r j ,d j )},x j ,a j ,x' j ,r j ,d j are the state, action, updated state after executing the action, reward value and round terminator of the jth sample respectively. Represents the Q network. Update the judgement parameters: Update the actor parameters: Then, update the target network parameters Among them, φ i ,θ i are the parameters of the online actor and judge, φ target,i ,θ target,i are the parameters of the target actor and the critic, and ρ is the weight of the updated parameters. This continues until the number of training rounds M is reached or the reward converges.
[0166] After the model training is completed, in order to further verify the effectiveness of the model's real-time scheduling strategy, the trained multi-agent system is deployed in the power system computing environment to apply the model.
[0167] The effect of the present invention is described below by a specific example:
[0168] The power system computing environment is built based on the IEEE-118 node system, and the power flow calculation part is implemented using PYPOWER. 9 wind turbines and 6 photovoltaics, a total of 15 new energy units, are connected to the IEEE-118 node system. Taking into account the topological connection information of the IEEE118 nodes, 9 thermal power units are selected as controlled generators based on the day-ahead scheduling solution. Five energy storage devices are installed in the corresponding new energy unit access nodes to jointly participate in the real-time scheduling of the system.
[0169] For the constructed system operation scenario dataset, the day-ahead unit output plan is firstly arranged based on the proposed QCQP model. Then, the SSRG-MADDPG model is trained in combination with the constructed training scenario set data. The reward function coefficients of the agent are set as shown in Table S1.
[0170] Table S1 Reward function correlation coefficient setting
[0171]
[0172] After the model training is completed, in order to further verify the effectiveness of the model's real-time dispatch strategy, the trained multi-agent system is deployed in the power system computing environment, and the performance of the model is verified and analyzed on the test set. The proposed SSRG-MADDPG algorithm is compared with the unit day-ahead plan, MADDPG and the model-driven mixed integer second-order cone optimization model (MISOCP). Among them, the unit day-ahead plan model means that all generators are output according to the day-ahead dispatch plan, energy storage does not participate in the real-time dispatch link, and the controlled units do not make additional adjustments. The MADDPG model means that its actions are not subject to hard constraints in the safety domain during the model deployment stage. The performance of each model in 50 test scenarios is shown in Table S2.
[0173] In Table S2, when the scenario has completed 90% of the process, that is, after running 260 steps, it is judged as a winner. It can be seen that in the face of ultra-short-term uncertainty of new energy, if the real-time scheduling strategy is not formulated, the winning rate in 50 test scenarios is 0, and the average system operation step length is only 54.51 steps. In most of these scenarios, continuous line over-limits occurred when new energy was greatly developed. In some scenarios, due to the fluctuation of new energy, the balancing machine adjustment capacity was insufficient and the balancing machine output exceeded the limit. Therefore, it is necessary to solve and formulate real-time scheduling strategies during the intraday stage.
[0174] Table S2 Performance analysis of each model in the test scenario
[0175]
[0176] Table S2 compares and analyzes the performance of SSRG-MADDPG, MADDPG, MISOCP and day-ahead unit planning in terms of training time, average solution time for a single scenario, average operation step length, winning rate, average maximum load factor of the line and average operation cost.
[0177] As can be seen from Table S2, in terms of solution efficiency, the data-driven model is significantly better than the model-driven model. The data-driven model can solve a single scenario in less than 1 minute on average, while the model-driven model takes more than 2 minutes, and the solution efficiency is increased by 3 times. In terms of average running step length and winning rate, the MISOCP model is slightly better than
[0178] However, under the guidance of SSR, the SSRG-MADDPG model outperforms the comparison model in terms of running step length and winning rate, and the running cost is greatly reduced compared with the MADDPG model, which is basically the same as the MISOCP model.
[0179] After analysis, it was found that under the action of SSR, the model avoided a large number of unnecessary actions of the intelligent agent, which not only reduced the operating cost, but also reserved sufficient flexible and adjustable resources for the system so that the intelligent agent could respond to the real "crisis" of the system.
[0180] The energy storage and controllable units can be fully coordinated and dispatched at all times. Therefore, the SSRG-MADDPG model significantly improves the safety and economy of the system.
[0181] On the other hand, an embodiment of the present invention further provides a real-time dispatching system for a high-proportion new energy power system under source-storage coordination, including:
[0182] A processor and a memory, the memory is used to store program instructions, and the processor is used to call the stored instructions in the memory to execute a real-time scheduling method for a high-proportion new energy power system under source-storage coordination as described in the above technical solution.
[0183] The above is only a preferred embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any technician familiar with the technical field can make equivalent replacements or changes according to the technical solution and inventive concept of the present invention within the scope disclosed by the present invention, which shall fall within the protection scope of the present invention.
Claims
1. A real-time dispatching method for a high-proportion new energy power system under source-storage coordination, characterized in that: include: Collect and organize source-load operation scenario data and perform pre-processing; Combined with the source-load day-ahead forecast scenario data, solve the day-ahead dispatch plan of the unit; Markov decision process modeling of real-time scheduling processes; Solve the static security domain model; Combining the solved static safety domain model, the agent observation space, action space and reward function in Markov decision process modeling, and the solved unit day-ahead dispatch plan, a multi-agent deep deterministic policy gradient model guided by the static safety domain is constructed. The input of the policy gradient model is the system state at each time step under each operating scenario: including new energy output, load value, unit day-ahead output plan value, and line load rate; the output is the action after the observed state, which is used to guide the dispatch of the controlled unit output; the model is trained using source-load operation scenario data, and the trained policy gradient model is used to realize real-time dispatch of the new energy power system.
2. A real-time dispatching method for a high-proportion new energy power system under source-storage coordination as claimed in claim 1, characterized in that: Collecting and organizing source-load operation scenario data includes: setting an N-node power system, with node 0 as the reference node and N G Thermal power units, N R Renewable energy units, N ESS Energy storage devices, N L load nodes; a total of N S The group source-load operation scenario, the single day-ahead real scenario data is represented by S rq = {P r,rq ,P l,rq }, of which new energy output data Load data The time interval of the day-ahead scenario is 15 minutes, so a single day-ahead scenario has 96 steps in total; the data of a single day-ahead forecast scenario is represented by S rq,f = {P r,rq,f ,P l,rq,f }, The real scene data in a single day can be expressed as S rn = {P r,rn ,P l,rn }, The time interval of the intraday scenario is 5 minutes, so there are 288 steps in the time series of a single intraday scenario; the prediction scenario data of a single day is represented by S rn,f = {P r,rn,f ,P l,rn,f }, Preprocessing includes supplementing missing values and removing outliers in source-load operation scenario data.
3. The real-time dispatching method for a high-proportion new energy power system under source-storage coordination according to claim 1, characterized in that: The solution of the unit day-ahead dispatch plan includes: Set a single day-ahead forecast scenario data scenario S rq,f The corresponding unit day-ahead dispatch plan is recorded as N G Represents the number of thermal power units, approximates the system AC power flow equation to a second-order cone form, and takes minimizing the system power generation cost as the goal. The mathematical model of the unit day-ahead scheduling plan is as follows: (1) Objective function Where, T DA is the step length of the day-ahead scheduling plan. The time interval of the day-ahead scheduling plan is 15 minutes, that is, T DA =96, N G is the number of generators in the system, P G,i,t is the output value of the ith generator at time t, a i , b i 、c i is the cost coefficient of the ith generator; (2) Power balance constraint at node j Where P G,k,t , Q G,k,t are the active and reactive outputs of generator set k at time t, respectively, k∈j means that generator set k is connected to the grid through node j; P b,ij,t and Q b,ij,t are the active and reactive powers flowing from node i to node j on branch i→j at time t; P b,je,t and Q b,je,t are the active and reactive powers flowing from node j to node e on branch j→e at time t; P L,j,t , Q L,j,t are the active and reactive loads of node j at time t, respectively; r ij and x ij are the resistance and reactance of branch i→j respectively, I ij,t is the square of the current magnitude flowing through branch i→j at time t; (3) Voltage amplitude constraints at both ends of lines i and j Where U i,t , U j,t are the voltage amplitudes of nodes i and j at time t respectively; (4) Power constraints of lines i and j (5) Line end node voltage phase angle constraint In the formula, and They represent variables that consider the voltage amplitude of nodes i and j to be constant, rather than optimization variables, and are taken as 1 here; θ i,t ,θ j,t are the phase angles of nodes i and j at time t respectively; (6) Generator output constraints In the formula, is the minimum and maximum output value of the i-th generator set, is the day-ahead planned value of the i-th generator unit at time t; (7) Generator ramp rate constraint In the formula, is the maximum ramp power of the ith generator set per unit time period; (8) Branch transmission power constraints In the formula, is the maximum transmission power of lines i and j; Therefore, the mathematical model of the unit day-ahead scheduling plan is written as: Finally, based on the above optimization model of Gurobi optimization solver and combined with the day-ahead source-load forecast value in the scenario data, the day-ahead output plan of the unit is solved.
4. The real-time dispatching method for a high-proportion new energy power system under source-storage coordination according to claim 1, characterized in that: The process of modeling the Markov decision process is as follows: Agent i: Agent i is deployed at the node where the controlled generator set is located or the node where the energy storage device is located, so as to correct the output plan of the node device in the real-time scheduling stage in combination with the intraday forecast value of the source and load; Status t :Indicators reflecting the operating status of the system at time t, including the actual output value of the thermal power unit at time t and the planned output value at time t+1, the actual value of the new energy unit and each node load at time t and the intraday forecast value at time t+1, the real-time SoC value of the energy storage equipment and the load rate of each line; Observed value o i,t : The system state value that each agent can observe is set as the state information of the node where the agent is located and its K-order neighbor nodes; Action a i,t :The scheduling actions that the agent can take; using discrete action space, for the controlled thermal power units, the climbing power per unit time of each unit is Normalized to [-1,1] and then N Pr Equally divided, -1 means the unit reduces output 1 means the unit increases output For energy storage equipment, the maximum charge and discharge power per unit time of each unit is Normalized to [-1,1] and then N Pr Equally divided, -1 means the charging power of the energy storage device is 1 means the discharge power of the energy storage device is Assume there are n agents in total. At time t, the environment receives the joint action A of the agents. t =[a 1,t ,a 2,t ,...,a n,t ]; Reward: The environment performs the joint action A of the agent t Finally, the reward is calculated through the set reward function and fed back to the multi-agent system.
5. The real-time dispatching method for a high-proportion new energy power system under source-storage coordination according to claim 4 is characterized in that: The reward function is designed as follows: 1) If the following conditions exist, the scenario will be terminated and a penalty r will be imposed. d :a) The system power flow does not converge; b) There are one or more lines with a load rate ρ>2; c) There are two or more lines with a load rate ρ>1; d) The balancing machine output exceeds the limit, r d The calculation formula is as follows: In the formula, r d,base is the baseline penalty for not completing the round, T d is the total step length of the round, T down is the current running step of the system, r d,min It is the minimum penalty given for an incomplete round, to avoid the incomplete penalty given in the later stage being too small, causing the agent to tend to not act; 2) If none of the above conditions exists, calculate the penalty r for overloaded and overloaded lines. L,t : In the formula, ρ l,t is the load rate of line l at time t, are the penalty coefficients for heavy-loaded lines and overloaded lines, respectively. If the line load rate exceeds the first threshold, the line is an overloaded line. If the load rate is between the second threshold and the first threshold, the line is a heavy-loaded line. When the system has overloaded or overloaded lines, a certain penalty should be given to guide the intelligent agent to make corresponding decisions and optimize the system operation status. Br is the total number of system lines; 3) Give system survival and operation status rewards r B,t : In the formula, r b is the basic reward for system survival, λ r,b is the reward for the system running step length t, which is used to guide the system to extend the running step length as much as possible, λ l,1 , l,2 , l,3 is the incentive factor for reducing the maximum load rate of the system line, ρ l,max is the maximum load rate of the system; 4) Calculate the operating costs of the controlled generator sets and energy storage equipment: Where N ctg is the number of controlled units, N stg is the number of energy storage devices in the system, is the adjustment cost coefficient of generator g, ΔP ctg,g,t is the absolute value of the output regulation of the generator g, I(·) is the indicative function, and the number of times the energy storage device operates is counted. is the action cost coefficient of energy storage device s; 5) Give the system a reward for running the entire scenario smoothly C,t ; Therefore, the system reward at time t is: Then the cumulative reward value of agent i at any time k is: Where γ is the reward discount factor.
6. The real-time dispatching method for a high-proportion new energy power system under source-storage coordination according to claim 1, characterized in that: The static safety domain of the new energy power system is defined as the set of points that satisfy the power flow equation and operation safety constraints, expressed as formula (16), where the new energy units and thermal power units are uniformly modeled as generator nodes, with the installed capacity as the maximum output upper limit; In the formula, according to different constraints, it is divided into different types of static security domains, Ω U To meet the static voltage safety domain of voltage constraints; Ω P The static generator active output safety domain that satisfies the generator active output constraint; Ω Q The static generator reactive output safety domain that satisfies the generator reactive output constraint; To meet the branch thermal stability safety domain of the system branch transmission power; Ω ss For the entire static security domain, Ω U ,Ω P ,Ω Q , Intersection composition; y SR is the system node voltage, phase vector; x SR is the node power injection vector; φ(x SR ,y SR )=0 is the AC system power flow equation; N is the system bus node set; N G is the set of system generator nodes; are the upper and lower limits of the voltage at the system node i respectively; They are the upper and lower limits of active output of system generator node i respectively; are the upper and lower limits of reactive power output of system generator node i; P b,ij , -P b,ij are the forward and reverse transmission power limits of the connection nodes i and j in the system, respectively; In the case of reactive power approximation local balance, only the static safety domain under active power injection is considered, that is: In the formula, x SR =[x P ,x Q ]; x P Active power injection for each node; x Q = C is the reactive power injection of each node, C is a constant; P G,i is the active output of the i-th generator; P L,j is the active load value of the jth node.
7. The real-time dispatching method for a high-proportion new energy power system under source-storage coordination according to claim 1, characterized in that: The overall process of the multi-agent deep deterministic policy gradient model guided by static safety domain is as follows: The intraday forecast scenario and real scenario are randomly selected from the source-load operation scenario set. At T=1, the actual data at the first time point in the intraday real scenario and the data at the first time point in the unit's day-ahead dispatch plan are loaded to perform AC power flow calculation to obtain the current state of the system. Then, the system state value at time T, the source-load intraday forecast value at time T+1, and the system state judgment result at time T+1 are provided to the intelligent agent as observation data. The intelligent agent makes decisions based on the observation data and outputs decision actions. After the action legitimacy is verified, the decision actions are provided to the power system computing environment for power flow calculation to obtain the system state at time T+1. At this time, the reward value is calculated based on the intelligent agent's actions and the current system state information, which serves as feedback information to guide the intelligent agent to evolve its strategy.
8. The real-time dispatching method for a high-proportion new energy power system under source-storage coordination according to claim 1, characterized in that: The judgment result of the system state at time T+1 is obtained by solving the static safety domain boundary, that is, combining the unit's day-ahead output value at time T+1 and the source load intra-day forecast value at time T+1, according to the source load forecast error, combined with Latin hypercube sampling, the possible source load intra-day output value at the next moment is extracted, and the day-ahead output value of the unit together constitutes the system operation point set at time T+1, and the static safety domain boundary is used for online judgment, and the result is provided to the intelligent agent as observation information; specifically, the ultra-short-term forecast value of the source load at the next moment is P R,f , P L,f , the current output value of the controlled unit in the system P g,sk , the planned output value P of the next moment of the non-controlled unit in the system g,fsk ; Assuming that the source-load ultra-short-term forecast error distribution conforms to the normal distribution, according to the LHS sampling N err Group error P R,err , P L,err , construct the next moment N err The observation x of the static safety domain model is constructed by combining the unit output with the source-load operation scenario value. p = {P g,sk ,P g,fsk ,P R,f +P R,err ,P L,f +P L,err }, the static security domain model is used to determine the security distribution of the source-load operation scenario at the next moment, and the static security domain model is set to determine N err Group Scene N' err The samples are safe. Provided to the agent as observation features at time T; when , it indicates that the system is safe at the next moment and no additional adjustment is required. Otherwise, it reminds the agent to adjust the controlled resources. S is the set threshold.
9. The real-time dispatching method for a high-proportion new energy power system under source-storage coordination according to claim 1, characterized in that: The training process of the multi-agent deep deterministic policy gradient model guided by static safety domain is as follows: First, initialize the power grid computing environment, training data set, agent parameters, experience pool, and set the number of training rounds to episode = m; Then, within the set number of rounds, a scene is randomly selected each time. At T = 0, the initial state x is calculated by combining the new energy output, load value, and unit planned output value. Then, the static safety domain model is used to predict the system state safety information at time T + 1, and the observation value o is formed together with x. s ; Then, in the scenario step, for each agent i, from the global observation o s Constructing its own observation i , then for the observed value o i , select action a i =μ θ,i (o i )+ε,ε~N(0,σ 2 I), μ θ,i (o i ) represents the policy network, ε~N(0,σ 2 I) represents the random perturbation of the policy network; the actions are discretized in combination with the Gumbel distribution, and the compliance of the actions is checked. Then, the joint actions of each agent are executed in the power grid computing environment. a=[a1,a2,...,a n ], get instant rewards r , the next state of the system is x', and the next state information is predicted based on the static safety domain model, which together with x' constitutes the observation value o' s The round ends with signal d, (x, a, x', r, d) is stored in the experience pool D, and the system state x←x' is updated; for the update of each agent, a batch of data B = {(x j ,a j ,x' j ,r j ,d j )},x j ,a j ,x' j ,r j ,d j They are the state, action, updated state after executing the action, reward value and round terminator of the jth sample respectively; finally, the data is updated until the number of training rounds M is reached or the reward converges.
10. A real-time dispatching system for a high-proportion new energy power system under source-storage coordination, characterized in that: include: A processor and a memory, the memory is used to store program instructions, and the processor is used to call the stored instructions in the memory to execute a real-time scheduling method for a high-proportion new energy power system under source-storage coordination as described in any one of claims 1-9.
Citation Information
Patent Citations
Power system transient stability prevention and emergency coordination control auxiliary decision-making method
CN113591379A
N-1 security constraint-considered risk scheduling method based on near-end strategy optimization algorithm
CN114142530A
Day-before-day coordinated optimization scheduling method considering source load uncertainty
CN115189401A
Deep reinforcement learning-based day-ahead-intra-day combined dispatching method for regional power grid
CN115441437A
Power grid active scheduling intelligent decision-making method and system based on Lagrange relaxation
CN117254468A