Power grid intra-day scheduling optimization method considering photovoltaic output uncertainty in severe weather
Through the A2C algorithm framework and knowledge migration technology, combined with the CGAN scenario generation method, the grid scheduling agent is optimized, which solves the challenge of photovoltaic output uncertainty on grid scheduling and realizes the safe and economic operation of the power grid in bad weather.
Patent Information
- Application Number
- CN202510109948.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-23
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2045-01-23
AI Technical Summary
Photovoltaic output is greatly affected by seasons and weather, resulting in a large deviation between the prediction results and the real value recently. A single deep reinforcement learning algorithm is difficult to fully adapt to the complex scheduling needs in high uncertain scenarios.
The A2C algorithm framework and knowledge transfer technology are adopted, combined with the CGAN scenario generation method, training samples are generated, scheduling agents are optimized, and scheduling agents are adapted to multiple source load scenarios, and specific intraday scheduling agents are improved through transfer learning.
It has achieved operation in a relatively economical way while ensuring the safety of the power grid, and improved the intelligence and adaptability of power grid scheduling, and adapted to the grid operation needs of high proportion of new energy access.
Smart Images

Figure CN119944652A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of electric power systems, and more specifically, relates to a method for optimizing intraday dispatching of power grids taking into account the uncertainty of photovoltaic output under severe weather conditions. The method uses an A2C algorithm framework and knowledge transfer technology to solve and optimize the power grid dispatching problem with the uncertainty of photovoltaic output under severe weather conditions. Background Art
[0002] Photovoltaic output is greatly affected by seasons and weather, and there is a large deviation between the day-ahead forecast and the actual value. Therefore, in a large-scale photovoltaic access power system, the accuracy of photovoltaic forecast data will have a great impact on the stable operation of the power grid. Therefore, it is very important to consider the photovoltaic output scenarios under various circumstances the next day.
[0003] Existing research has widely used reinforcement learning algorithms to solve and optimize the dispatching problems of power grids in power systems. However, the output of photovoltaic units is greatly affected by the weather and is difficult to accurately predict. It is difficult to fully adapt to the complex dispatching needs under such high uncertainty scenarios by relying on a single deep reinforcement learning algorithm. To this end, this paper proposes a grid intraday dispatch optimization method that considers the uncertainty of photovoltaic output under severe weather conditions, aiming to further improve the intelligence and adaptability of grid dispatching and provide support for the operation of grids with a high proportion of new energy access. Summary of the invention
[0004] The present invention establishes a safe and economical dispatch model for a power grid with access to photovoltaic new energy based on the intraday optimization dispatch mode, and proposes an intraday dispatch optimization method for a power grid that takes into account the uncertainty of photovoltaic output under severe weather conditions, which can enable the power grid to operate in a more economical manner while ensuring safety.
[0005] To achieve the above object, the present invention adopts the following technical solution:
[0006] A method for optimizing grid intraday dispatch considering uncertainty of photovoltaic output under severe weather conditions includes the following steps:
[0007] Step 1: Determine the sequential decision-making process of grid daily dispatch and establish a mathematical model;
[0008] Step 2: Determine the optimization objectives and constraints of the daily dispatch of the power grid and establish its learning optimization model;
[0009] Step 3: Use the CGAN scenario generation method to generate training samples, and use the A2C algorithm framework to optimize the scheduling agent that can adapt to various source-load scenarios;
[0010] Step 4: Use transfer learning to obtain the intraday scheduling agent for specific source-load scenarios.
[0011] The sequential decision-making process of the grid’s intraday dispatch described in step 1 has the following specific model:
[0012] The sequential decision model first needs to construct the observed power grid operation information into a state matrix at the tth moment, which is recorded as S t The state matrix mainly includes the current time t, the observation information ob of the power grid operation state t , and the load forecast value and new energy unit output forecast value for the next time t+1 t , so the state matrix can be written as follows:
[0013] S t =[t,ob t ,fore t ],t∈(0,1,2,...,T)
[0014] The time interval is 15 minutes, so T = 95. The agent receives a state matrix S t After being taken as input, the agent strategy network π will output a unit action adjustment A based on the information in the current state matrix t Assume that the system contains N ther thermal power unit, then the action adjustment amount of the nth thermal power unit at the tth time is recorded as The adjustment action matrix A of all units t It can be written in the form of the following matrix:
[0015]
[0016] The adjustment amount of the unit needs to be mapped to the current unit output value. The mapping formula is as follows:
[0017]
[0018] Where P t ther represents the output value of all thermal power generating units in the system at time t, They represent the allowable adjustable lower limit and adjustable upper limit of the output of all thermal power units in the system at time t respectively.
[0019] As mentioned above, the state matrix S at time t t Input into the strategy π and decide the adjustment amount A t After that, the power grid adjusts the corresponding thermal power units and executes them, and enters the next decision-making process. The decision-making process can be described by mathematical symbols:
[0020] S t →π→A t →S t+1 →π→A t+1 …→ST
[0021] The constraints of the intraday dispatch model of the power grid described in step 2 include line flow power constraints, power balance between generation and consumption, upper and lower limits of thermal power unit ramp power constraints, and upper and lower limits of unit output power constraints, etc., as follows:
[0022] The line power flow constraint is:
[0023]
[0024] in, is the line power matrix composed of the power of all lines at time t, is the line limit power matrix composed of all line power limit values, l∈N line , N line The numbers of all lines on all bus nodes.
[0025] The power balance of the generator is:
[0026]
[0027] in N is the time t PV The power matrix of the output of photovoltaic units, is the output power of the jth photovoltaic unit at time t, P t load It is the active power on the electricity consumption side.
[0028] The upper and lower power constraints of the unit at any time t are:
[0029]
[0030] in is the minimum / maximum output power matrix of thermal power units, is the minimum / maximum output power of the kth thermal power unit.
[0031] The climbing power constraint:
[0032]
[0033] in is the matrix composed of the downward / upward climbing power of the thermal power unit, is the downward / upward climbing power of the nth thermal power unit.
[0034] The grid optimization objectives described in step 2 are as follows:
[0035] The operation target of the power grid mainly considers the operation cost of thermal power units and the consumption rate of photovoltaic units. for:
[0036]
[0037] Among them, α i ,β i ,γ i is the operating cost coefficient of unit i when the quadratic model is used. The photovoltaic absorption rate can be expressed as:
[0038]
[0039] in represents the short-term predicted maximum output value of the j-th photovoltaic power station at time t, Indicates the current output value of the j-th PV unit at time t.
[0040] The CGAN scene generation method described in step 3 generates training samples as follows:
[0041] CGAN is an extended version of the generative adversarial network, which guides the process of generating data by introducing conditional information. In the power system dispatching scenario, CGAN can generate realistic scene data based on specific input conditions (such as meteorological data, load demand, power generation plan, etc.), which is used to simulate the output or source-load scenarios of new energy sources such as photovoltaic and wind power. CGAN consists of two main components: the generator and the discriminator. The former generates a virtual sample G(z|c) that meets the conditions based on the input conditional variable c and random noise z. The latter judges whether the input sample is real data by inputting real samples and generated samples, and outputs the probability of the true value. On the input side of CGAN, the random noise matrix d is merged with the label matrix l and used as the input data of the generator G(d|l), and finally the generated sample x′=G(d|l) that meets the label value is obtained. The discriminator D(x|l) not only needs to obtain the size p(x) of the sample that meets the real sample distribution, but also needs to judge whether the sample belongs to condition l. Therefore, during the training process, the sample distribution p output by the generator G(d|l) should be as close as possible to the sample distribution p(x). g (x′) continuously fits the distribution p(x) of samples in historical data; while the discriminator D(x|l) should try its best to improve the ability to distinguish true and false samples. In the renewable energy power system, the output of photovoltaic and wind power is affected by many factors such as weather and time. Its volatility and uncertainty bring great challenges to grid dispatching. Through CGAN scenario generation technology, rich simulation data can be provided for the grid dispatching agent to enhance the robustness of the dispatching model.
[0042] The photovoltaic output data is divided into multiple scene sets with different characteristics according to the month or solar term as the condition value. Assuming that the data set is divided into m scene sets, considering the randomness of photovoltaic output, each moment of a sample is matched to a noise value, then the definition of the noise sequence d is as follows: d = {d1, d2, d3, ···, d i}Where: i is the total number of moments in a sample. Similarly, the conditional values are divided according to the number of scene sets, and the conditional sequence l is defined as follows: l = {l1, l2, l3, ···, l m}, the two sequences are concatenated and the input matrix is first passed through the fully connected layer for feature extraction, and then a nonlinear function is added between each fully connected layer. The addition of nonlinear factors to the fully connected layer enables the neural network to more accurately approximate the conditional label value and the mapping relationship between the noise distribution and the output distribution of the new energy unit. Then, after layer-by-layer calculation, the input and output of each layer are obtained, and finally the output value of the fully connected layer is obtained.
[0043] By learning the time series characteristics and error distribution of wind and solar power data, a wind and solar scene that conforms to the actual situation is generated. Specifically, the model first uses the LSTM network to capture the time series characteristics of wind and solar power data, and then combines the kernel density estimation method to fit the wind power prediction error and generate random noise that conforms to the error distribution. Next, these noise data are used to train the CGAN model. At the same time, the day-ahead forecast value of wind and solar power is used as a conditional input to help the model learn the mapping relationship between the conditional noise distribution and the actual data. The entire design aims to more accurately simulate the randomness and uncertainty of wind and solar power, and provide an effective solution for the generation of wind and solar scenes.
[0044] Train a dispatching agent that can adapt to various source-load scenarios, use the CGAN scenario generation method to generate photovoltaic output scenarios, and select three main source-load scenario samples of sunny, cloudy and rainy days based on the maximum output, half-load probability and output period of photovoltaic data in that period. These scenario data can be used to verify the dispatching strategy of the power grid under different photovoltaic output conditions.
[0045] The A2C algorithm optimization described in step 3 first obtains the state from the environment and then inputs it into the agent's network. The actor network will output the mean and variance, and at the same time construct the corresponding Gaussian distribution and sample the action. Finally, the action is input into the environment for execution. After the execution is completed, the environment outputs the state information at the next moment. The agent's critic network will calculate the corresponding value estimate based on the current state and the state at the next moment, and further calculate the mean square error based on the real-time reward feedback from the environment, and use it to update the agent's critic network. Finally, the actor network is updated in combination with the output action information.
[0046] In step 4, transfer learning is used to improve the optimization of the scheduling agent within a specific day:
[0047] Using network-based transfer learning, this migration method refers to reusing the trained network in the source domain, including its network structure and connection parameters, and converting it into deep neural network parameters for initial training in the target domain. The specific formula is as follows:
[0048]
[0049] Where π source ,π target Represent the policy networks of the source task and the target task respectively, θ source represents the network parameters of the source task policy network, θ t ' arget Represents the initial network parameters of the target task, so the parameters of the source task policy network are directly transferred to the target task policy network, and then continue to learn. This greatly improves the training efficiency and brings better results.
[0050] Different from the prior art, the above technical solution has the following beneficial effects:
[0051] Aiming at the problem of intraday optimization scheduling of power grids with photovoltaic new energy access, this paper proposes an intraday intelligent scheduling optimization method with A2C algorithm framework and knowledge transfer method, and uses CGAN scenario generation technology to expand the output samples of photovoltaic units, providing a large number of source-load scenario sets for scheduling agents. Then, the knowledge transfer method is used to transfer the scheduling knowledge of agents that adapt to multiple source-load scenarios to agents for specific intraday scheduling. Through this optimization method, the power grid can operate in the most economical way while ensuring safety. BRIEF DESCRIPTION OF THE DRAWINGS
[0052] Figure 1 This is a diagram of the sequential decision-making process for daily dispatch of power grid;
[0053] Figure 2 Flowchart optimized for A2C algorithm;
[0054] Figure 3 A simplified schematic diagram of the CGAN scene generation method;
[0055] Figure 4 This is a schematic diagram of scheduling knowledge transfer;
[0056] Figure 5 A comparison chart of the off-load power and the operating cost of thermal power units with and without transfer learning;
[0057] Figure 6 Optimization curve of total reward with and without transfer learning. DETAILED DESCRIPTION
[0058] In order to explain the technical content, structural features, achieved objectives and effects of the technical solution in detail, the following is a detailed description in conjunction with specific embodiments and accompanying drawings.
[0059] The present invention discloses a method for optimizing the daily dispatch of power grids that takes into account the uncertainty of photovoltaic output under severe weather conditions. Considering that photovoltaic output is greatly affected by season and weather, there is a large deviation between its day-ahead prediction result and the true value, and a single deep learning algorithm is difficult to meet the dispatch problem of power grids after large-scale photovoltaic units are connected. Therefore, a CGAN-based scenario generation method is used to construct a photovoltaic output sample set, and then the A2C algorithm framework is used to optimize the dispatching agent. The dispatching optimization method proposed in the present invention ensures the safe and economical operation of the power grid under various severe source-load scenarios.
[0060] The present invention is specifically described by taking a provincial power grid as an example, and the present invention is further described in detail below in conjunction with specific embodiments and drawings.
[0061] See Figure 1 As shown in Figure 1, it is a schematic diagram of the sequential decision-making process for daily dispatch of power grid.
[0062] Step 1: The sequential decision-making process for determining the intraday dispatch of the power grid is as follows:
[0063] First, the state information is obtained from the environment and then input into the agent's network. The agent's actor network will output the mean and variance, then construct the corresponding Gaussian distribution and sample the action. Secondly, the action is input into the environment for execution. After the execution is completed, the environment outputs the state information of the next moment.
[0064] The power optimization dispatch divides the dispatch problem within a period of time into multiple dispatch problems of smaller time scales. The dispatch time is 24 hours, and a dispatch plan is made every 15 minutes, with a total of 96 dispatch plans made per day. First, at the tth moment, the observed information is constructed into a state matrix and recorded as S t Including the current time t, the observation information of the power grid operation status ob t , and the load forecast value and new energy unit output forecast value for the next time t+1 t , so the state matrix can be written as follows:
[0065] S t =[t,ob t ,fore t ],t∈(0,1,2,...,T)
[0066] Among them, the observation matrix ob tIt is composed of variables in the operation of the power grid, mainly including bus load, information on the current power generation of each unit, and information on each node of the current system. When the agent receives a state matrix S at the tth moment t After being taken as input, the agent strategy network π will output a unit action adjustment A based on the information in the current state matrix t , the adjustment action matrix A of all units t It can be written in the form of the following matrix:
[0067]
[0068] The adjustment amount of the unit needs to be mapped to the current unit output value. The mapping formula is as follows:
[0069]
[0070] Where P t ther represents the output value of all thermal power generating units in the system at time t, so in is the output value of the i-th thermal power unit at time t. They represent the lower and upper limits of the output of all thermal power units in the system at time t, respectively. in is the adjustable lower bound or upper bound of the i-th thermal power unit at time t.
[0071] Step 2: Determine the optimization objectives and constraints of the grid’s daily dispatch:
[0072] The most important thing for the safe operation of the power system is to ensure the safe operation of the power grid during the daily power optimization dispatch. The dispatch safety is mainly achieved by ensuring that the power transmission of the power grid line does not exceed the power limit of the line, and at the same time, it is also necessary to ensure the balance between the source and the load. The safe operation of the power grid also needs to solve the problem of active power distribution of the generator set. At the same time, on this basis, certain constraints need to be met, such as line flow power constraints, upper and lower limits of thermal power unit climbing power constraints, and upper and lower limits of unit output power constraints. As shown in the following formula:
[0073] The line power flow constraint is:
[0074]
[0075] in, is the line power matrix composed of the power of all lines at time t, is the line limit power matrix composed of all line power limit values, l∈N line , N line The numbers of all lines on all bus nodes.
[0076] The safe operation of the power grid requires frequency stability. To ensure frequency stability, the active power on the power generation side must be the same as the active power on the power consumption side. Therefore, at any time t, the system should satisfy the following power balance equation:
[0077]
[0078] in N is the time t PV The power matrix of the output of photovoltaic units, is the output power of the jth PV unit at time t.
[0079] The upper and lower power constraints of the thermal power unit are:
[0080]
[0081] in is the minimum / maximum output power matrix of thermal power units, is the minimum / maximum output power of the kth thermal power unit.
[0082] The thermal power unit climbing constraint:
[0083]
[0084] in is the matrix composed of the downward / upward climbing power of the thermal power unit, is the downward / upward climbing power of the nth thermal power unit.
[0085] The operation target of the power grid mainly considers the operation cost of thermal power units and the absorption rate of photovoltaic units. The calculation formula is:
[0086]
[0087] Among them, α i ,β i ,γ i is the operating cost coefficient of unit i when the quadratic model is adopted. The photovoltaic absorption rate formula is as follows:
[0088]
[0089] in represents the short-term predicted maximum output value of the j-th photovoltaic power station at time t, It represents the current output value of the jth photovoltaic unit at time t. If the short-term predicted output of photovoltaic is zero, the absorption rate is 100%. In the actual optimization process, some constraints need to be introduced into the objective function as penalty terms. The specific optimization objective function is given in the example.
[0090] Step 3: Use the CGAN scenario generation method to generate training samples, and use the A2C algorithm framework to optimize the scheduling agent that can adapt to a variety of source-load scenarios.
[0091] CGAN is an extended version of the generative adversarial network, which guides the process of generating data by introducing conditional information. In the power system dispatching scenario, CGAN can generate realistic scene data based on specific input conditions (such as meteorological data, load demand, power generation plan, etc.), which is used to simulate the output or source-load scenarios of new energy sources such as photovoltaic and wind power. CGAN consists of two main components: the generator and the discriminator. The former generates a virtual sample G(z|c) that meets the conditions based on the input conditional variable c and random noise z. The latter judges whether the input sample is real data by inputting real samples and generated samples, and outputs the probability of the true value. On the input side of CGAN, the random noise matrix d is merged with the label matrix l and used as the input data of the generator G(d|l), and finally the generated sample x′=G(d|l) that meets the label value is obtained. The discriminator D(x|l) not only needs to obtain the size p(x) of the sample that meets the real sample distribution, but also needs to judge whether the sample belongs to condition l. Therefore, during the training process, the sample distribution p output by the generator G(d|l) should be as close as possible to the sample distribution p(x). g (x′) continuously fits the distribution p(x) of samples in historical data; while the discriminator D(x|l) should try its best to improve the ability to distinguish true and false samples. In the renewable energy power system, the output of photovoltaic and wind power is affected by many factors such as weather and time. Its volatility and uncertainty bring great challenges to grid dispatching. Through CGAN scenario generation technology, rich simulation data can be provided for the grid dispatching agent to enhance the robustness of the dispatching model.
[0092] The photovoltaic output data is divided into multiple scene sets with different characteristics according to the month or solar term as the condition value. Assuming that the data set is divided into m scene sets, considering the randomness of photovoltaic output, each moment of a sample is matched to a noise value, then the definition of the noise sequence d is as follows: d = {d1, d2, d3, ···, d i}Where: i is the total number of moments in a sample. Similarly, the conditional values are divided according to the number of scene sets, and the conditional sequence l is defined as follows: l = {l1, l2, l3, ···, l m}, the two sequences are concatenated and the input matrix is first passed through the fully connected layer for feature extraction, and then a nonlinear function is added between each fully connected layer. The addition of nonlinear factors to the fully connected layer enables the neural network to more accurately approximate the conditional label value and the mapping relationship between the noise distribution and the output distribution of the new energy unit. Then, after layer-by-layer calculation, the input and output of each layer are obtained, and finally the output value of the fully connected layer is obtained.
[0093] By learning the time series characteristics and error distribution of wind and solar power data, a wind and solar scene that conforms to the actual situation is generated. Specifically, the model first uses the LSTM network to capture the time series characteristics of wind and solar power data, and then combines the kernel density estimation method to fit the wind power prediction error and generate random noise that conforms to the error distribution. Next, these noise data are used to train the CGAN model. At the same time, the day-ahead forecast value of wind and solar power is used as a conditional input to help the model learn the mapping relationship between the conditional noise distribution and the actual data. The entire design aims to more accurately simulate the randomness and uncertainty of wind and solar power, and provide an effective solution for the generation of wind and solar scenes.
[0094] Train a dispatching agent that can adapt to various source-load scenarios, use the CGAN scenario generation method to generate photovoltaic output scenarios, and select three main source-load scenario samples of sunny, cloudy and rainy days based on the maximum output, half-load probability and output period of photovoltaic data in that period. These scenario data can be used to verify the dispatching strategy of the power grid under different photovoltaic output conditions.
[0095] See Figure 2 As shown in FIG. 1 , the flowchart of A2C algorithm optimization includes the following steps:
[0096] Step 3.1: First, use the scenario generation method based on conditional generative adversarial networks to construct a set of photovoltaic output samples under various weather scenarios, and select the three main source-load scenario samples of sunny days, cloudy days and rainy days according to the characteristics of the maximum output, half-load probability and output period of the photovoltaic data in that period.
[0097] Step 3.2: Then, based on the generated sample set, the dominant actor-critic algorithm is used to optimize the dispatching agent that can adapt to various source-load scenarios. The intraday dispatch plan needs to be made every 15 minutes, so the timeliness requirement is relatively high. This section uses the data-driven deep reinforcement learning algorithm to first learn the intraday dispatching agent offline based on historical data and scenario generation technology, and then use the dispatching agent after offline learning to receive real-time information of the power grid and real-time forecast information of new energy and load as input for online decision-making. The intraday dispatching agent that adapts to various source-load scenarios is trained using the framework of the A2C algorithm. First, the new energy output and load forecast curves are input as part of the state matrix S into the general agent. In addition, the state matrix S also includes bus load and other node information as well as the output plan of the current unit. After receiving the input, the agent will output the adjustment action vector A of each unit and calculate the output value of the current unit through the action. Then the output value of the current unit will be imported into the power grid environment, and the safety index, economic index, clean index and load abandonment penalty of the system will be calculated after the power system flow calculation and constraint condition check, and then the above indicators will be weighted according to the importance to obtain the corresponding cost reward R. Finally, the power grid simulation environment will update the final executed unit output plan to the current unit output plan in the state.
[0098] In order to improve the learning efficiency of the agent, the parallel architecture of CPU multi-threading is used to accelerate its calculation. The specific implementation details are to divide the entire policy network into a main network and multiple parallel independent sub-threads. Each sub-thread interacts with its own environment independently by loading the parameters of the main network to obtain an independent sampling trajectory. At the same time, in the operation environment of the power grid, different source-load scenarios are set for each sub-grid environment, that is, the 1st to nth sub-environments are set with different samples of sunny day operation scenarios, the n+1th to 2nth sub-environments are set with different samples of cloudy day operation scenarios, and the 2n+1th to 3nth sub-environments are set with different samples of rainy day operation scenarios. By setting the scenario parameters of different sub-environments, the policy network can explore and learn in power grid operation scenarios with different climates, so that the dispatching agent can learn decision-making methods under different conditions and improve the ability of the agent to operate safely in various harsh power grid scenarios.
[0099] Step 4: Use transfer learning to improve the optimization of the scheduling agent for a specific day.
[0100] Step 4.1: Use network-based transfer learning to reuse the network trained in the source domain, including its network structure and connection parameters, and convert it into deep neural network parameters for initial training in the target domain. The purpose of the transfer is to use the source domain D s In the learning task T sThe experience of how to safely and effectively adjust the unit output for various climate source and load scenarios can help improve the target domain D t In the learning task T t The learning speed of the policy network and its ability to make safe scheduling decisions in harsh scenarios are as follows:
[0101]
[0102] Where π source ,π target Represent the policy networks of the source task and the target task respectively, θ source represents the network parameters of the source task policy network, θ t ' arget Represents the initial network parameters of the target task, so the parameters of the source task policy network are directly transferred to the policy network of the target task, and then continue learning.
[0103] The present invention takes into account the fact that traditional deep learning algorithms are difficult to fully meet the needs of grid dispatch optimization when dealing with the volatility and uncertainty of photovoltaic output in power systems with large-scale photovoltaic access, and proposes a grid daily dispatch optimization method that takes into account the uncertainty of photovoltaic output under severe weather conditions.
[0104] It should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that the process, method, article or terminal device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or terminal device. In the absence of further restrictions, the elements defined by the sentence "include..." or "comprise..." do not exclude the existence of other elements in the process, method, article or terminal device including the elements. In addition, in this article, "greater than", "less than", "exceed" and the like are understood to exclude the number itself; "above", "below", "within" and the like are understood to include the number itself.
[0105] Although the above embodiments have been described, once those skilled in the art know the basic creative concepts, they can make additional changes and modifications to these embodiments. Therefore, the above description is only an embodiment of the present invention and does not limit the patent protection scope of the present invention. Any equivalent structure or equivalent process transformation made by using the contents of the specification and drawings of the present invention, or directly or indirectly used in other related technical fields, are also included in the patent protection scope of the present invention.
Claims
1. A method for optimizing grid intraday dispatch considering the uncertainty of photovoltaic output under severe weather conditions, characterized in that: The following steps are involved: Step 1: Determine the sequential decision-making process of grid daily dispatch and establish a mathematical model; Step 2: Determine the optimization objectives and constraints of the daily dispatch of the power grid and establish its learning optimization model; Step 3: Use the CGAN scenario generation method to generate training samples, and use the A2C algorithm framework to optimize the scheduling agent that can adapt to various source-load scenarios; Step 4: Use transfer learning to obtain the intraday scheduling agent for specific source-load scenarios.
2. The method for optimizing grid daytime dispatch considering the uncertainty of photovoltaic output under severe weather conditions as claimed in claim 1, characterized in that: The sequential decision-making process of the intraday dispatch of the power grid in step 1 is specifically modeled as follows: The sequential decision model first needs to construct the observed power grid operation information into a state matrix at the tth moment, which is recorded as S t The state matrix mainly includes the current time t, the observation information ob of the power grid operation state t , and the load forecast value and new energy unit output forecast value for the next time t+1 t , so the state matrix can be written as follows: S t =[t,ob t ,fore t ],t∈(0,1,2,...,T) The time interval is 15 minutes, so T = 95, and the agent receives a state matrix S t After being taken as input, the agent strategy network π will output a unit action adjustment A based on the information in the current state matrix t ; Assume that the system contains N ther thermal power unit, then the action adjustment amount of the nth thermal power unit at the tth time is recorded as The adjustment action matrix A of all units t It can be written in the form of the following matrix: The adjustment amount of the unit needs to be mapped to the current unit output value. The mapping formula is as follows: Where P t ther represents the output value of all thermal power generating units in the system at time t, They represent the allowable lower and upper limits of the output of all thermal power units in the system at time t respectively; As mentioned above, the state matrix S at time t t Input into the strategy π and decide the adjustment amount A t After that, the power grid adjusts the corresponding thermal power units and executes them, and enters the next decision-making process. The decision-making process can be described by mathematical symbols: S t →π→A t →S t+1 →π→A t+1 …→S T 。 3. The method for optimizing grid daytime dispatch considering the uncertainty of photovoltaic output under severe weather conditions as claimed in claim 1, characterized in that: The constraints of the grid intraday dispatch model in step 2 include line flow power constraints, generation and consumption power balance, upper and lower limit constraints of thermal power unit ramp power, upper and lower limit constraints of unit output power, etc., which are as follows: The line flow power constraints: in, is the line power matrix composed of the power of all lines at time t, is the line limit power matrix composed of all line power limit values, l∈N line , N line The numbers of all lines on all bus nodes; The power balance of the generator is: in N is the time t PV The power matrix of the output of photovoltaic units, is the output power of the jth photovoltaic unit at time t, P t load It is the active power on the electricity consumption side; The upper and lower power constraints of the unit at any time t are: in is the minimum / maximum output power matrix of thermal power units, is the minimum / maximum output power of the kth thermal power unit; The climbing power constraint: in is the matrix composed of the downward / upward climbing power of the thermal power unit, is the downward / upward climbing power of the nth thermal power unit.
4. The method for optimizing grid daytime dispatch considering the uncertainty of photovoltaic output under severe weather conditions as claimed in claim 1, characterized in that: The power grid optimization objectives in step 2 are as follows: The operation target of the power grid mainly considers the operation cost of thermal power units and the consumption rate of photovoltaic units. for: Among them, α i ,β i ,γ i When the quadratic model is used, the operating cost coefficient of unit i and the photovoltaic absorption rate can be expressed as: in represents the short-term predicted maximum output value of the j-th photovoltaic power station at time t, Indicates the current output value of the j-th PV unit at time t.
5. The method for optimizing grid daytime dispatch considering the uncertainty of photovoltaic output under severe weather conditions as claimed in claim 1, characterized in that: The CGAN scene generation method in step 3 generates training samples as follows: CGAN consists of two main components: a generator and a discriminator. The generator generates virtual samples G(z|c) that meet the conditions according to the input conditional variables c and random noise z. The discriminator judges whether the input samples are real data by inputting real samples and generated samples, and outputs the probability of true value. In the new energy power system, the output of photovoltaic and wind power is affected by multiple factors such as weather and time. Its volatility and uncertainty bring great challenges to grid dispatching. Through CGAN scenario generation technology, rich simulation data can be provided for the grid dispatching intelligent body to enhance the robustness of the dispatching model. Train a scheduling agent that can adapt to various source-load scenarios, use the CGAN scenario generation method to generate photovoltaic output scenarios, and select the three main source-load scenario samples of sunny days, cloudy days and rainy days based on the maximum output, half-load probability and output period of photovoltaic data in that period.
6. The method for optimizing grid daytime dispatch considering the uncertainty of photovoltaic output under severe weather conditions as claimed in claim 1, characterized in that: In the step 3, the A2C algorithm is optimized by first obtaining the state from the environment and then inputting it into the network of the agent. The actor network will output the mean and variance, and at the same time construct the corresponding Gaussian distribution and sample the action; finally, the action is input into the environment for execution, and after the execution is completed, the environment outputs the state information at the next moment; the critic network of the agent will calculate the corresponding value estimate based on the current state and the state at the next moment, and further calculate the mean square error based on the real-time reward feedback from the environment, and use it to update the critic network of the agent, and finally update the actor network in combination with the output action information.
7. The method for optimizing grid daytime dispatch considering the uncertainty of photovoltaic output under severe weather conditions as claimed in claim 1, characterized in that: In step 4, transfer learning is used to improve the optimization of the scheduling agent within a specific day: Use network-based transfer learning. This migration method refers to reusing the trained network in the source domain, including its network structure and connection parameters, and converting it into deep neural network parameters for initial training in the target domain. The specific formula is as follows: Where π source ,π target Represent the policy networks of the source task and the target task respectively, θ source represents the network parameters of the source task policy network, θ t ' arget Represents the initial network parameters of the target task, so the parameters of the source task policy network are directly transferred to the policy network of the target task, and then continue learning.
Citation Information
Patent Citations
Deep reinforcement learning-based day-ahead-intra-day combined dispatching method for regional power grid
CN115441437A
Power distribution network optimization scheduling method considering intelligent soft switching and demand side response
CN116885688A
Short-term photovoltaic output prediction model construction method and prediction method
CN117454751A
Novel power system dispatching scene sample data generation method and model construction method
CN117933760A
Photovoltaic short-term output scene generation method based on improved CGAN
CN118154355A
Cited By
Method and system for evaluating photovoltaic bearing capacity of distributed photovoltaic transformer area
CN121395572A