Multi-scale optimization scheduling method and system for distribution network based on deep reinforcement learning
Through a multi-scale optimization scheduling method based on deep reinforcement learning, using probability box theory and multi-agent deep reinforcement learning algorithm, a grid optimization scheduling model with centralized training and decentralized execution architecture is constructed, which solves the problem of instability in power supply and demand caused by source load uncertainty in the distribution network, and achieves efficient new energy consumption and cost reduction.
Patent Information
- Application Number
- CN202510857711.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-25
- Publication Date
- 2025-08-22
- Estimated Expiration
- 2045-06-25
AI Technical Summary
The existing distribution network optimization scheduling methods are difficult to effectively deal with the high uncertainty on both sides of the source and load, resulting in problems such as imbalance in power supply and demand, voltage fluctuations and frequency deviations. Especially in power grids with high proportion of new energy access, traditional deterministic prediction methods are difficult to capture the randomness and load characteristics of distributed energy such as wind power and photovoltaics.
A multi-scale optimization scheduling method based on deep reinforcement learning is adopted, and the probability box theory is used to characterize the source load prediction error probability distribution, and combined with the multi-agent deep reinforcement learning algorithm, a power grid optimization scheduling model with centralized training and decentralized execution architecture is constructed, including a few days, day and real-time optimization scheduling model, and dynamic balance is used to achieve flexible resource optimization configuration.
It significantly improves the operating stability and economy of the distribution network under different time scales, enhances the adaptability to source load uncertainty, improves the level of new energy consumption and reduces electricity consumption costs.
Smart Images

Figure CN120373805B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of power systems, and in particular to a multi-scale optimization scheduling method and system for distribution networks based on deep reinforcement learning. Background Art
[0002] With the optimization of energy structure and the development of grid technology, distribution networks occupy a key position in modern power systems. Under the development trend of new energy as the main body, distributed energy such as wind energy and solar energy, as well as flexible resources such as energy storage systems, demand-side response, and electric vehicles are connected on a large scale in the distribution network. This has optimized the energy structure and enhanced the grid regulation capability to a certain extent. However, the operation of traditional distribution networks mainly adopts a one-way dispatching mode of source following load. This dispatching method based on deterministic prediction has become difficult to adapt to the complex operating environment brought about by the high proportion of new energy access.
[0003] The core problem currently facing the optimization and dispatching of distribution networks stems from the high uncertainty on both the source and load sides. On the power supply side, the output of distributed energy sources such as wind power and photovoltaics has significant randomness and volatility, and their power generation capacity is affected by weather conditions, environmental factors, etc. and shows strong nonlinear characteristics. Traditional deterministic prediction methods are difficult to accurately capture their changing patterns; on the load side, with the access of flexible resources such as demand-side response, electric vehicle charging, and distributed energy storage, the load characteristics have changed from passive acceptance to active participation, and the uncertainty of electricity consumption behavior has increased significantly. This uncertainty on both the source and load sides not only increases the complexity of grid operation and dispatch, but may also cause a series of problems such as imbalance in electricity supply and demand, voltage fluctuations, and frequency deviations. In severe cases, it may even lead to large-scale power outages, posing a huge challenge to the stable operation of the power grid.
[0004] In summary, existing technologies have many shortcomings in dealing with these problems and are unable to effectively meet the requirements for optimized operation of distribution networks under new power systems. Therefore, how to accurately characterize the uncertainties on both the source and load sides and achieve efficient optimized scheduling of distribution networks has become a key issue that needs to be urgently addressed in the current distribution network field. It is of great significance to ensure the safe and stable operation of the power grid and optimize resource allocation. Summary of the Invention
[0005] To solve the above technical problems, the present invention provides a multi-scale optimization scheduling method and system for distribution networks based on deep reinforcement learning.
[0006] In a first aspect, the present invention provides a multi-scale optimization scheduling method for a distribution network based on deep reinforcement learning, the method comprising the following steps:
[0007] Based on the forecast data and actual source load data of the source load dispatch day, the probability box theory is used to model the probability distribution of the source load forecast error on the dispatch day and obtain the net load uncertainty interval.
[0008] According to the response characteristics of distribution network equipment, the distribution network dispatch cycle is divided into multiple time scale levels to obtain a multi-time scale dispatch level.
[0009] Based on the multi-timescale scheduling hierarchy and distribution network partition information, a multi-agent deep reinforcement learning algorithm is used to construct a power grid optimization scheduling model based on a centralized training and decentralized execution architecture; the power grid optimization scheduling model includes a day-ahead centralized optimization scheduling model, an intraday distributed rolling optimization scheduling model, and a real-time optimization scheduling model;
[0010] Taking minimizing the total operating cost of the day-ahead as the optimization goal, according to the net load uncertainty interval and the pre-acquired self-regulating temperature coefficient, the day-ahead dispatch strategy is obtained by solving the power grid optimization dispatch model;
[0011] Based on the day-ahead dispatching strategy, multi-scale rolling optimization is performed using the real-time operating status of distribution network equipment and ultra-short-term forecast deviation data, and a multi-scale full-cycle collaborative dispatching strategy is obtained through iterative optimization.
[0012] In a further embodiment, the step of modeling the probability distribution of the source load forecast error on the scheduling day using the probability box theory based on the source load scheduling day forecast data and the actual source load data to obtain the net load uncertainty interval includes:
[0013] Calculate the deviation between the daily source load scheduling forecast data and the actual source load data to obtain the source load forecast error, and use the source load forecast error as the uncertainty component of the random variable in the probability box theory;
[0014] Performing a probability statistical analysis on the source-load prediction error to obtain a probability distribution characteristic parameter boundary of the source-load prediction error;
[0015] Constructing a cumulative probability distribution function of the source-load prediction error according to the probability distribution characteristic parameter boundary, and obtaining a cumulative probability distribution boundary of the cumulative probability distribution function;
[0016] Calculating the power prediction error boundary of each source load according to the cumulative probability distribution boundary, and obtaining the source load power uncertainty interval according to the power prediction error boundary of each source load and the source load scheduling daily forecast data;
[0017] The net load prediction error boundary is calculated according to the source load power uncertainty interval to generate a net load uncertainty interval.
[0018] In a further embodiment, the multi-time-scale scheduling level includes at least three time-scale scheduling levels, which are a day-ahead scale scheduling level, an intraday scale scheduling level, and a real-time scale scheduling level.
[0019] In a further embodiment, the step of constructing a power grid optimization scheduling model based on a centralized training and decentralized execution architecture using a multi-agent deep reinforcement learning algorithm based on the multi-timescale scheduling hierarchy and distribution network partition information includes:
[0020] The distribution network is divided into several autonomous regions according to its geographical structure and the distribution of its equipment. Regional agents are deployed for each autonomous region. Each regional agent has an independent actor network and critic network.
[0021] Define the state space and action space of regional agents at different time-scale scheduling levels, and establish an initial optimization scheduling model for each regional agent based on a centralized training and decentralized execution architecture;
[0022] During the centralized training phase, the actor network parameters and critic network parameters of each regional agent are randomly initialized;
[0023] During the training process, the state information, action information, and reward information generated by the interaction between the regional agent and the environment are stored in the replay buffer. Each time the actor network parameters and the critic network parameters are updated, batches of data are randomly sampled from the replay buffer to form a training dataset.
[0024] Based on the training dataset, updating the critic network parameters by minimizing the soft Bellman residual to obtain the optimal action value estimate;
[0025] The action entropy is calculated based on the action selection probability distribution output by the actor network, and the self-regulating temperature coefficient is calculated based on the uncertainty quantification parameter and action entropy obtained in advance;
[0026] According to the optimal action value estimate and the self-regulating temperature coefficient, the actor network parameters of the regional agent are updated by gradient back propagation to obtain the global coordination strategy parameters;
[0027] In the decentralized execution phase, the updated actor network parameters are used to generate partition-independent action instructions based on the real-time local observation status of the agents in each region;
[0028] The initial optimization scheduling model is iteratively optimized at multiple time scales according to the global coordination strategy parameters and independent action instructions to obtain a power grid optimization scheduling model.
[0029] In a further embodiment, the steps of defining the state space and action space of the regional agent at different time scale scheduling levels and establishing an initial optimization scheduling model for each regional agent based on a centralized training decentralized execution architecture include:
[0030] The day-ahead state space is constructed based on the active power output data of micro-turbines, photovoltaic power output forecast values, wind power output forecast values, and load forecast data of each autonomous region.
[0031] Define all the control actions taken by the regional agent in the day-ahead phase to form the day-ahead action space;
[0032] Taking minimization of the day-ahead total operating cost of the distribution network as the optimization objective, a day-ahead centralized optimization scheduling model is constructed according to the day-ahead state space and the day-ahead action space;
[0033] The intraday state space is constructed based on the real-time data of regional wind and solar power output, regional load demand, and regional energy storage charge state data within the autonomous region under the jurisdiction of each regional intelligent agent, and the gas turbine output adjustment amount and energy storage charge and discharge adjustment amount are defined as the intraday action space;
[0034] The day-ahead optimization scheduling strategy obtained by solving the day-ahead centralized optimization scheduling model is used as the intraday initial condition, and the minimization of the operating cost within the autonomous region is taken as the optimization goal. An intraday distributed rolling optimization scheduling model is constructed according to the intraday state space and the intraday action space.
[0035] A real-time state space is constructed based on the real-time operating status of the distribution network in the real-time phase, and the energy storage charging and discharging power is defined as the real-time action space;
[0036] The intraday scheduling strategy output by the intraday distributed rolling optimization scheduling model is used as the real-time initial condition, minimizing the power imbalance in the autonomous area is taken as the optimization goal, and a real-time optimization scheduling model is constructed according to the real-time state space and the real-time action space.
[0037] In a further embodiment, the action entropy is defined as the negative logarithmic mean of the action selection probability distribution.
[0038] In a further implementation scheme, the total day-ahead operating cost is the sum of the distribution network power generation cost, distribution network active power loss cost, energy storage system control cost, flexible load control cost and flexible load penalty cost during the distribution network control cycle.
[0039] In a further embodiment, the operating cost within the autonomous region is the sum of the gas turbine adjustment cost and the energy storage adjustment cost of the distribution network during the regulation cycle.
[0040] In a further embodiment, the power imbalance amount within the autonomous region is the product of the net power fluctuation value of the distribution network during the regulation cycle and the penalty weight.
[0041] In a second aspect, the present invention provides a distribution network multi-scale optimization scheduling system based on deep reinforcement learning, the system comprising:
[0042] The source-load analysis module is used to model the probability distribution of the source-load forecast error on the dispatch day based on the source-load dispatch day forecast data and actual source-load data, and obtain the net load uncertainty interval;
[0043] A scale division module is used to divide the distribution network dispatching cycle into multiple time scale levels according to the response characteristics of the distribution network equipment to obtain a multi-time scale dispatching level;
[0044] A model building module is used to build a power grid optimization scheduling model based on a centralized training and decentralized execution architecture using a multi-agent deep reinforcement learning algorithm based on the multi-timescale scheduling hierarchy and distribution network partition information; the power grid optimization scheduling model includes a day-ahead centralized optimization scheduling model, an intraday distributed rolling optimization scheduling model, and a real-time optimization scheduling model;
[0045] a model solving module, configured to obtain a day-ahead dispatching strategy by solving the power grid optimization dispatching model based on the net load uncertainty interval and the pre-acquired self-regulating temperature coefficient, with minimizing the day-ahead total operating cost as the optimization objective;
[0046] The optimization scheduling module is used to perform multi-scale rolling optimization based on the day-ahead scheduling strategy, using the real-time operating status of distribution network equipment and ultra-short-term forecast deviation data, and obtain a multi-scale full-cycle collaborative scheduling strategy through iterative optimization.
[0047] The present invention provides a distribution network multi-scale optimization scheduling method and system based on deep reinforcement learning. The method uses probability box theory to model the probability distribution of source-load prediction errors on the scheduling day according to source-load scheduling daily forecast data and actual source-load data, and obtains a net load uncertainty interval; divides the distribution network scheduling cycle into multi-time scale levels according to the response characteristics of distribution network equipment, and obtains a multi-time scale scheduling level; based on the multi-time scale scheduling level and distribution network partition information, a multi-agent deep reinforcement learning algorithm is used to construct a power grid optimization scheduling model based on a centralized training and decentralized execution architecture; the power grid optimization scheduling model includes a day-ahead centralized optimization scheduling model, an intra-day distributed rolling optimization scheduling model, and a real-time optimization scheduling model; with the minimization of the day-ahead total operating cost as the optimization goal, the day-ahead scheduling strategy is obtained by solving the power grid optimization scheduling model according to the net load uncertainty interval and the pre-acquired self-regulating temperature coefficient; based on the day-ahead scheduling strategy, multi-scale rolling optimization is performed using the real-time operating status of the distribution network equipment and ultra-short-term prediction deviation data, and a multi-scale full-cycle collaborative scheduling strategy is obtained through iterative optimization. Compared with the existing technology, this method realizes the flexible resource optimization configuration of the distribution network in day-ahead, intraday and real-time scheduling by comprehensively utilizing probability box theory modeling, multi-time scale hierarchical division and multi-agent deep reinforcement learning algorithm to construct a power grid optimization scheduling model, enhances the adaptability to source and load uncertainty, significantly improves the operation stability and economy of the distribution network at different time scales, and makes the scheduling strategy more in line with actual operation needs. BRIEF DESCRIPTION OF THE DRAWINGS
[0048] Figure 1 This is a flow chart of a multi-scale optimization scheduling method for a distribution network based on deep reinforcement learning provided by an embodiment of the present invention;
[0049] Figure 2 is a schematic diagram of flexibility capacity provided by an embodiment of the present invention;
[0050] Figure 3 Schematic diagram of a multi-timescale optimization scheduling framework provided by an embodiment of the present invention;
[0051] Figure 4 Schematic diagram of an optimization scheduling framework based on a multi-agent deep reinforcement learning algorithm provided by an embodiment of the present invention;
[0052] Figure 5 Schematic diagram of the multi-agent reinforcement learning training process provided by an embodiment of the present invention;
[0053] Figure 6 is a schematic diagram of an IEEE-33 node provided by an embodiment of the present invention;
[0054] Figure 7 This is a schematic diagram of 24-hour source-load forecast data provided by an embodiment of the present invention;
[0055] Figure 8 This is a schematic diagram of source-load prediction data and equipment operating parameters provided by an embodiment of the present invention;
[0056] Figure 9 This is a schematic diagram of the optimization results of the previous 24 hours provided by an embodiment of the present invention;
[0057] Figure 10 2 is a schematic diagram comparing load data before and after day-ahead optimization provided by an embodiment of the present invention;
[0058] Figure 11 Schematic diagram of micro gas turbines and energy storage output in various regions provided by an embodiment of the present invention;
[0059] Figure 12 This is a schematic diagram comparing the flexibility capacity and demand provided by an embodiment of the present invention;
[0060] Figure 13 1 is a schematic diagram comparing the output of a micro gas turbine provided by an embodiment of the present invention;
[0061] Figure 14 1 is a schematic diagram comparing energy storage outputs provided by an embodiment of the present invention;
[0062] Figure 15 This is a block diagram of a multi-scale optimization and scheduling system for distribution networks based on deep reinforcement learning provided by an embodiment of the present invention.
[0063] Explanation of reference numerals: 101, source-load analysis module; 102, scale division module; 103, model construction module; 104, model solution module; 105, optimization scheduling module. DETAILED DESCRIPTION
[0064] The following describes the embodiments of the present invention in detail with reference to the accompanying drawings. The embodiments are provided for illustrative purposes only and are not to be construed as limiting the present invention. The accompanying drawings are provided for reference and illustration only and do not constitute a limitation on the scope of protection of the present invention. Many changes may be made to the present invention without departing from the spirit and scope of the present invention.
[0065] The multi-scale distribution network optimization scheduling method provided in this embodiment uses the probability box theory to characterize the intervals of wind power, photovoltaic power, and load forecast errors, constructs upper and lower bound cumulative probability distribution functions, quantifies the range of forecast uncertainty, and divides the scheduling cycle into three levels: day-ahead (1 hour), intraday (15 minutes), and real-time (5 minutes). These levels address medium- and long-term forecast deviations, short-term fluctuations, and ultra-short-term power fluctuations, respectively. This method effectively addresses the uncertainty of source-load forecasts, thereby improving the absorption of renewable energy and reducing electricity costs. Furthermore, this embodiment combines an improved multi-agent deep reinforcement learning algorithm with the introduction of a self-regulating temperature coefficient for dynamic balance exploration and utilization. It also uses a centralized training and decentralized execution architecture to improve decision-making efficiency, solving the optimal multi-scale, full-cycle coordinated scheduling strategy for the power grid. This coordinated scheduling strategy includes micro gas turbines, energy storage systems, and flexible loads, effectively reducing the decision-making time of the power grid scheduling model. This ensures the safe operation of the distribution network while effectively addressing the uncertainty of source-load forecasts, improving the absorption of renewable energy, and reducing electricity costs. This provides a solution for optimizing distribution networks with a high proportion of renewable energy access. Figure 1 This is a flow chart of a multi-scale optimization scheduling method for a distribution network based on deep reinforcement learning provided by an embodiment of the present invention. The embodiment of the present invention provides a multi-scale optimization scheduling method for a distribution network based on deep reinforcement learning, such as Figure 1 As shown, the method includes the following steps:
[0066] S1. Based on the source load scheduling day forecast data and actual source load data, the probability box theory is used to model the probability distribution of the source load forecast error on the scheduling day to obtain the net load uncertainty interval.
[0067] In some embodiments, the step of modeling the probability distribution of the source load forecast error on the scheduling day using the probability box theory based on the source load scheduling day forecast data and the actual source load data to obtain the net load uncertainty interval includes:
[0068] Calculate the deviation between the daily source load scheduling forecast data and the actual source load data to obtain the source load forecast error, and use the source load forecast error as the uncertainty component of the random variable in the probability box theory;
[0069] Performing a probability statistical analysis on the source-load prediction error to obtain a probability distribution characteristic parameter boundary of the source-load prediction error;
[0070] Constructing a cumulative probability distribution function of the source-load prediction error according to the probability distribution characteristic parameter boundary, and obtaining a cumulative probability distribution boundary of the cumulative probability distribution function;
[0071] Calculating the power prediction error boundary of each source load according to the cumulative probability distribution boundary, and obtaining the source load power uncertainty interval according to the power prediction error boundary of each source load and the source load scheduling daily forecast data;
[0072] The net load prediction error boundary is calculated according to the source load power uncertainty interval to generate a net load uncertainty interval.
[0073] When dealing with the uncertainty problem of the distribution network, the prediction time scale has a significant impact on the accuracy of the source-load prediction. Specifically, the uncertainty errors presented by predictions at different time scales vary in size, especially in the source-load prediction. For example, the day-ahead prediction is usually based on a time scale of 1 hour, which is affected by multiple uncertainties such as long-term weather changes, equipment aging or potential failures, and fluctuations in user electricity demand, resulting in large prediction errors. In contrast, the intraday prediction uses a shorter time scale of 15 minutes, mainly targeting short-term weather changes, short-term equipment failures, and fluctuations in user electricity demand. The prediction accuracy is improved and the uncertainty range is correspondingly reduced. The real-time prediction uses a time scale of 5 minutes, which can achieve The best source load prediction accuracy is achieved, and it performs outstandingly in terms of accuracy and confidence. In summary, different prediction time scales have their own advantages in source load prediction, among which real-time prediction is particularly outstanding in terms of accuracy and prediction confidence interval. It should be noted that the probability box theory, as a mathematical tool for dealing with uncertainty, can provide accurate and relatively conservative estimates by modeling uncertainty as a range and assigning the degree of uncertainty within the range according to the probability distribution. In the process of extracting the source load uncertainty interval, the probability box theory can give a reasonable uncertainty range when there is insufficient data or limited expert experience, comprehensively describe the nature of uncertainty, effectively deal with nonlinear and complex systems, and provide a more reliable basis for decision-making.
[0074] For the interval characterization method, this embodiment first uses the prediction algorithm to obtain the source-load prediction values of wind power, photovoltaic output and load demand on the dispatch day. These prediction values constitute the random variables in the probability box (p-box) theory. The uncertainty of random variables is reflected in the source load prediction error. By analyzing historical data, this embodiment can determine the probability box parameters and convert the random variable The upper bound of the cumulative probability distribution function of and the lower bound As the interval boundary, it is assumed that the source load prediction error obeys the normal distribution, so its mean and standard deviation will fluctuate within a certain range. The confidence level is 0.9, and the corresponding upper and lower boundary quantiles are 95% and 5%, respectively. The probability distribution parameters of the prediction error and the corresponding cumulative probability distribution function boundaries can be obtained by calculation, and then the prediction interval information is obtained. The cumulative probability distribution function boundaries corresponding to the probability distribution parameters of the prediction error are:
[0075]
[0076] Where, is the lower bound of the cumulative probability distribution function of the photovoltaic output forecast error, which corresponds to the minimum probability distribution of the photovoltaic forecast error; is the upper bound of the cumulative probability distribution function of photovoltaic output forecast error, which corresponds to the maximum probability distribution of photovoltaic forecast error; is the lower bound of the cumulative probability distribution function of wind power output forecast error, which corresponds to the minimum probability distribution of wind power forecast error; is the upper bound of the cumulative probability distribution function of wind power output forecast error, which corresponds to the maximum probability distribution of wind power forecast error; is a random variable (prediction error value) used as the independent variable of the probability distribution function; It is the lower bound mixing parameter of the normal distribution when the mean of the prediction error takes the lower bound and the standard deviation takes the upper bound; is the lower boundary parameter of the normal distribution of the forecast error; It is the lower bound mixing parameter of the normal distribution when the mean of the prediction error takes the upper bound and the standard deviation takes the lower bound; is the upper boundary parameter of the normal distribution of the forecast error, and the mean and standard deviation take the upper boundary of the interval; is the lower boundary of the interval of the mean prediction error; is the upper boundary of the interval of the mean prediction error; is the upper boundary of the interval of forecast error standard deviation; is the lower boundary of the interval of forecast error standard deviation; t is the time variable.
[0077] Similarly, after obtaining the predicted values of photovoltaic power, wind power, and load, this embodiment can use the probability box theory to calculate their uncertainty intervals. Since the net load is the difference between wind power and photovoltaic power output and load demand, the net load prediction error and its uncertainty interval can be calculated using the photovoltaic, wind power, and load error values. The net load uncertainty interval is:
[0078]
[0079] Where, is the lower boundary of the interval of net load forecast value at time t, which is the minimum net load after considering the error; is the upper boundary of the interval of net load forecast value at time t, which is the maximum net load after considering the error; is the net load at time t, which is the difference between wind power, photovoltaic output and load demand; is the lower bound of the interval of net load forecast error at time t; is the upper bound of the net load forecast error at time t.
[0080] Next, this embodiment quantifies the flexibility demand of the power system based on the net load forecast interval. By calculating the net load forecast value and its uncertainty interval, the interval range of the uncertainty demand capacity in the power system can be further evaluated, thereby clarifying the range of the flexibility supply capacity. This embodiment defines the flexibility demand at time t as the change in net load in adjacent time periods. By analyzing the change in net load forecast values between the previous moment and the current moment and the corresponding forecast interval, this embodiment can calculate the flexibility demand capacity. The specific calculation formula for the net load forecast value change is:
[0081]
[0082] Where, is the net load change between adjacent time periods, i.e., the flexibility demand; is the net load power at time t; is the net load power at time (t-1).
[0083] The formula for the change in the net load forecast value shows that the change in the net load forecast value is the flexibility demand. Without considering the forecast uncertainty interval, if the relationship between the flexibility supply capacity and the flexibility demand satisfies , then the system flexibility is balanced, where Provides flexibility to the system by supplying capacity in time intervals The amount of change within.
[0084] Since the net load forecast value has directionality, the flexibility of the power system also has directionality. Figure 2 This is a schematic diagram of flexibility capacity provided by an embodiment of the present invention. The flexibility capacity mainly includes two scenarios: net load increase and net load decrease. Since uncertainty intervals are introduced into the net load forecast value, the flexibility demand capacity interval must be calculated at the same time as the flexibility demand. The calculation formula for the flexibility demand capacity interval is:
[0085]
[0086] Where, is the net load change in time interval The lower boundary within , which corresponds to the minimum requirement for net load change; is the upper boundary of the flexibility demand interval, which corresponds to the maximum demand for net load change; is the lower bound of the net load forecast at time t; is the upper bound of the net load forecast at time (t-1); is the upper bound of the net load forecast at time t; is the lower bound of the net load forecast at time (t-1).
[0087] Through the above process, the upper and lower bounds of the flexibility demand range can be obtained. At the same time, the system needs to meet the upward flexibility capacity. The mathematical expression of the flexibility supply capacity constraint is:
[0088]
[0089]
[0090] Where, Provide capacity for upward flexibility; Provides capacity for downward flexibility.
[0091] According to the flexibility supply capacity constraint, upward flexibility supply capacity means that the discharge capacity can be increased, while downward flexibility supply capacity means that the discharge and charge capacities can be reduced. Figure 2 It can be seen from the above that when the net load increases, , upward flexibility supply capacity When the net load decreases, the upward adjustment resources (such as energy storage discharge) need to be called first; when the net load decreases, that is, , downward flexibility spare capacity If the net load is large, then the downward resources (such as energy storage charging) will be called. When the net load does not change, that is, , the flexibility supply capacity is equal to the flexibility demand capacity , which is equal to the net load forecast error .
[0092] S2. Divide the distribution network dispatching cycle into multiple time scale levels according to the response characteristics of the distribution network equipment to obtain a multi-time scale dispatching level.
[0093] This embodiment adopts a multi-timescale optimization strategy to manage the distribution network to cope with the multi-scale fluctuations of source-load forecast uncertainty in the day-ahead, intra-day, and real-time stages. Specifically, this embodiment adopts an improved multi-agent deep reinforcement learning algorithm to centrally train agents at different time scales, thereby constructing an efficient distribution network optimization control strategy. In terms of specific operations, this embodiment constructs a training data set based on historical data, incorporates renewable energy and load data into the environmental model, and uses distribution network operation data as the state input of the agent. The action instructions generated by the agent are converted into control signals, and the corresponding reward value is calculated according to a preset reward function. Figure 3 1 is a schematic diagram of a multi-time-scale optimization scheduling framework provided by an embodiment of the present invention. In this embodiment, the multi-time-scale scheduling level includes at least three time-scale scheduling levels, namely, a day-ahead scheduling level, an intraday scheduling level, and a real-time scheduling level.
[0094] In the day-ahead phase, uncertainty primarily stems from biases in medium- and long-term weather forecasts and the randomness of load variations. This embodiment, based on historical forecast error data, utilizes the probability box theory to determine the possible ranges and probabilistic distributions of forecast errors for photovoltaic power generation, wind power, and load. Because day-ahead forecasts are updated infrequently (typically daily), their error ranges are relatively stable and can cover uncertainties over a longer timeframe. Therefore, day-ahead optimization can generate 24-hour power generation plans and backup capacity. However, because day-ahead forecasts cannot capture short-term fluctuations, the optimization results may be conservative and less economical. In this embodiment, the primary function of the day-ahead phase is to provide a foundational reference for intraday and real-time optimization, ensuring system reliability and economic efficiency over a longer timeframe. Therefore, the intelligent agents in the day-ahead phase operate on a one-hour timescale, employing a centralized optimization approach to achieve the economically optimal daily operation of the power system. This day-ahead phase considers the uncertainty ranges in power generation and load forecasts, generating a 24-hour scheduling strategy that includes distributed micro-gas turbines, energy storage system scheduling, and flexible load regulation, while ensuring a certain degree of robustness.
[0095] In the intraday phase, uncertainty primarily stems from short-term meteorological changes (such as cloud movement and wind speed fluctuations) and random load fluctuations. The role of the probability box theory in this phase is to generate more accurate and targeted error intervals and their probability distributions by combining short-term forecast data. Due to the high update frequency of intraday forecasts (typically every 15 minutes to one hour), the probability box theory can dynamically capture short-term fluctuations, providing a more accurate description of uncertainty for rolling optimization. Based on the error intervals described by the probability box theory, the system can dynamically adjust power generation plans and reserve capacity, optimizing the output curves of controllable resources (such as energy storage and micro-gas turbines), effectively addressing short-term uncertainty. Compared to the day-ahead phase, the intraday phase offers advantages in both forecast data update frequency and optimization strategy flexibility, enabling a better balance between economy and system reliability. In this intraday phase, the intelligent agent operates on a 15-minute timescale, employing a decision-making approach of offline centralized training and online decentralized execution. Day-ahead decisions are modified based on intraday power generation and load fluctuations, generating a one-hour scheduling strategy for micro-gas turbines and energy storage systems in each region, ensuring overall system robustness and economy.
[0096] In the real-time stage, the intelligent agent uses ultra-short-term precise forecasting data on a 5-minute time scale to obtain the output fluctuations of photovoltaic and wind power. Faced with real-time fluctuations, the intelligent agent quickly adjusts the distributed energy storage system to suppress power fluctuations in the power grid and ensure the stable operation of the system. The rapid response and precise regulation at this stage provide strong guarantees for the real-time stability of the power grid.
[0097] S3. Based on the multi-timescale scheduling hierarchy and distribution network partition information, a multi-agent deep reinforcement learning algorithm is used to construct a power grid optimization scheduling model based on a centralized training and decentralized execution architecture; the power grid optimization scheduling model includes a day-ahead centralized optimization scheduling model, an intraday distributed rolling optimization scheduling model, and a real-time optimization scheduling model.
[0098] S4. Taking minimization of the total operating cost on the day before as the optimization goal, the day before scheduling strategy is obtained by solving the power grid optimization scheduling model according to the net load uncertainty interval and the pre-acquired self-regulating temperature coefficient.
[0099] S5. Based on the day-ahead dispatching strategy, multi-scale rolling optimization is performed using the real-time operating status of distribution network equipment and ultra-short-term forecast deviation data, and a multi-scale full-cycle collaborative dispatching strategy is obtained through iterative optimization.
[0100] In some embodiments, the step of constructing a power grid optimization scheduling model based on a centralized training and decentralized execution architecture using a multi-agent deep reinforcement learning algorithm based on the multi-timescale scheduling hierarchy and distribution network partition information includes:
[0101] The distribution network is divided into several autonomous regions according to its geographical structure and the distribution of its equipment. Regional agents are deployed for each autonomous region. Each regional agent has an independent actor network and critic network.
[0102] Define the state space and action space of regional agents at different time-scale scheduling levels, and establish an initial optimization scheduling model for each regional agent based on a centralized training and decentralized execution architecture;
[0103] During the centralized training phase, the actor network parameters and critic network parameters of each regional agent are randomly initialized;
[0104] During the training process, the state information, action information, and reward information generated by the interaction between the regional agent and the environment are stored in the replay buffer. Each time the actor network parameters and the critic network parameters are updated, batches of data are randomly sampled from the replay buffer to form a training dataset.
[0105] Based on the training dataset, updating the critic network parameters by minimizing the soft Bellman residual to obtain the optimal action value estimate;
[0106] The action entropy is calculated based on the action selection probability distribution output by the actor network, and the self-regulating temperature coefficient is calculated based on the uncertainty quantification parameter and action entropy obtained in advance;
[0107] According to the optimal action value estimate and the self-regulating temperature coefficient, the actor network parameters of the regional agent are updated by gradient back propagation to obtain the global coordination strategy parameters;
[0108] In the decentralized execution phase, the updated actor network parameters are used to generate partition-independent action instructions based on the real-time local observation status of the agents in each region;
[0109] The initial optimization scheduling model is iteratively optimized at multiple time scales according to the global coordination strategy parameters and independent action instructions to obtain a power grid optimization scheduling model.
[0110] Specifically, the Soft Actor-Critic (SAC) algorithm optimizes the decision-making strategy by balancing expected rewards and information entropy during training to avoid falling into the problem of local optimal solutions. This design enables the intelligent agent to more comprehensively learn the source-load uncertainty space, thereby improving the system's global optimization capabilities and ability to handle source-load uncertainty problems. In addition, this embodiment further encourages the intelligent agent to explore by introducing the intelligent agent's action entropy into the value function, thereby enhancing the stability of the algorithm. The action entropy is defined as the negative logarithmic mean of the action selection probability distribution. In this embodiment, the expression of the action entropy is:
[0111]
[0112] Where, In state Next, the agent action strategy The action entropy is used to measure the agent’s state The uncertainty of action selection under , the larger the value, the stronger the exploration; is the expectation operator, which represents the action strategy of the agent Next, action In state The expected value of the sample at is the action strategy of the agent, and its parameters are A neural network that outputs the probability distribution of the agent's action choices; is the environmental state of the agent at time t, that is, the system state observed by the agent at time t (such as wind and solar power output, load, etc.); In state Next, the agent selects an action The logarithmic probability of action Log-likelihood under policy; To show that the agent follows the strategy at time t The selected action, for example, the control instruction output by the strategy at time t (such as energy storage charging and discharging power).
[0113] Action entropy specifically reflects the probability distribution characteristics of the agent's action strategy. Among them, the temperature coefficient of action entropy is used to regulate the degree of influence of action entropy on rewards. The size of the temperature coefficient has a direct impact on the model's ability to explore uncertainty, and is therefore related to the diversity of exploration behavior and the optimization efficiency during training. In order to enhance the degree of learning about uncertainty while reducing unnecessary exploration, this embodiment introduces a method of self-adjusting temperature coefficient, setting the self-adjusting temperature coefficient to a coefficient that dynamically changes with the size of environmental uncertainty. In each step of the optimization process, the calculation formula for updating the self-adjusting temperature coefficient takes into account the measurement coefficient of the size of the scheduling environment uncertainty (this coefficient is mainly related to the size of the predicted power and prediction error) and the uncertainty quantification parameter determined by the historical prediction error data and sensitivity to the environment. In this way, this embodiment can dynamically adjust the impact of action entropy on rewards according to the uncertainty in the environment, thereby achieving more efficient exploration and optimization. The update calculation formula of the self-adjusting temperature coefficient in each step of the optimization process is:
[0114]
[0115]
[0116] Where, is the self-regulating temperature coefficient after updating at time t; is the dynamic self-regulating temperature coefficient at time t, which is used to dynamically adjust the impact of action entropy on rewards; is the learning rate for self-regulating temperature coefficient update, which is used to control Update rate hyperparameters; is the trajectory expectation, which represents the strategy Next, the state-action pair ( ) Expected value under is the initial temperature coefficient of action entropy, which is used to characterize the influence of action entropy on reward; is the initial action entropy; It is an uncertainty quantification parameter, which is determined by the historical data of prediction error and sensitivity to the environment, and is used to dynamically adjust Sensitivity parameters; is the source-load prediction error at time t, the absolute error of wind-solar and load prediction at time t; is the predicted power at time t.
[0117] As a reinforcement learning algorithm within the actor-critic framework, the soft action-critic algorithm models action strategies through an actor network and outputs a probability distribution of actions. Simultaneously, the critic network evaluates the strategies generated by the actor network based on a state-value function and a soft Q-function. During training, the algorithm alternately updates the parameters of the actor and critic networks using gradient backpropagation, thereby gradually optimizing the strategies and approximating the optimal Q-value function. The state-value function in the soft action-critic algorithm is defined as follows:
[0118]
[0119] Where, For the state ( ), the Critic network parameters The state value function is used to estimate the state ( ) The expected reward of the agent under For the strategy Next, action expected value; For the state ( ), it takes action ( ) is an action-value function that is used to evaluate the expected reward of the action.
[0120] Parameters of the Critic Network It can be updated by minimizing the soft Bellman residual. The Critic network updates the parameters of the Critic network. Learn accurate Q values, guide the Actor network to optimize the strategy through divergence update and reparameterization algorithm, and minimize the soft Bellman residual to update the parameters of the Critic network The formula is:
[0121]
[0122] Where, To represent the loss function of the Critic network, used to update the Critic network parameters ; is the state-action pair in the experience replay buffer D ( )’s expected value of experience replay; In state Next, take action The action-value function of In state Next, take action Instant rewards received; is the reward discount factor, which is used to weigh the impact of immediate rewards and future rewards; For the environmental dynamics model Next, generate the next state ( )’s expected value of state transition; is the target critic network parameter The target state value function; D is the experience replay buffer, which stores the sample four-element group ( , , , ); are the parameters of the Critic network; is the target critic network parameter; is the immediate reward at time t; i is the agent index.
[0123] Parameters of the actor network The Kullback-Leibler (KL) divergence can be used for updating, and the update formula is:
[0124]
[0125] Where, The loss function of the Actor network is used to update the parameters of the Actor network. ; is the state in the experience replay buffer D expected value; is the Kullback-Leibler (KL) divergence, which is used to measure the distance between two probability distributions; In state The action value function under the state Q-value mapping of the next action; is a normalization factor used to ensure that the sum of the probability distribution is 1; is an exponential transformation of the action-value function.
[0126] This embodiment can learn the policy parameters by minimizing the equation In order to obtain a lower variance estimate, a reparameterization algorithm is used for the strategy output. The reparameterization algorithm is:
[0127]
[0128] Where, is the state in the experience replay buffer D and Gaussian noise expected value; To reparameterize the function, the noise and status Mapping to Action ; is Gaussian noise, which is usually sampled from a standard normal distribution; In state Next, take sampling action The action-value function of .
[0129] After each update of the critic network parameters, the target critic network needs to be soft-updated to avoid overestimation of the Q value and ensure the stability of the power grid dispatch strategy. The soft update formula is:
[0130]
[0131] Where, are the parameters of the target Critic network; is a hyperparameter for soft update, which is used to control the weight update ratio between the target network and the current network.
[0132] Figure 4 This is a schematic diagram of an optimization scheduling framework based on a multi-agent deep reinforcement learning algorithm provided by an embodiment of the present invention. Multi-agent training is extended using the Centralized Training with Decentralized Execution (CTDE) model. To ensure the stability of each agent during training, the Critic network of each agent shares global information. During the centralized training phase, the Critic network incorporates the observations and actions of other agents as additional information. During the decentralized execution phase, the Actor network only needs to make decisions based on the agent's own private observations, ensuring real-time decision-making. Figure 5 This is a schematic diagram of the multi-agent reinforcement learning training process provided by an embodiment of the present invention. During the training process, the actions, states, and rewards of each agent are stored in the replay buffer. The critic network obtains the observation and action information of other agents through the replay buffer to update the neural network. Therefore, the four-element group of the experience replay buffer is expanded to include the set of elements of all agents ( , , , ),in, is the joint state space, which is the set of observations of all agents at time t; is the joint action space, which is the set of actions of all agents at time t; is the joint reward signal, which is the reward set of all agents at time t; is the joint state at the next moment, which is executed The set of states to transfer to later.
[0133] Assume that there are n agents in the environment, and the Actor network of agent i is , each Actor network Parameters , then according to the CTDE framework and the soft action-critic algorithm, the Actor network parameters of agent i are The update formula is:
[0134]
[0135] Where, is the Actor network loss function of agent i, which is used to update the parameters of the Actor network ; To sample the local state of agent i in the experience replay buffer D and noise expected value; is the Actor network of agent i, parameterized as , used to generate action strategies; is the action sampling function of agent i, which is used to reparameterize the action from the action probability distribution; is the local observation of agent i at time step t; are the parameters of the Actor network of agent i; i is the index of the agent; To represent the Critic network of agent i in state and actions The action-value function under .
[0136] Each agent has an independent Actor network and Critic network. Agent i uses its own observation state Get independent action , the difference is that the Critic network samples the states of all agents from the experience replay pool D and actions , which enables the agent to adapt to the multi-agent collaborative environment, and its parameters The update formula is:
[0137]
[0138]
[0139] Where, is the loss function of the Critic network of agent i, which is used to update the parameters of the Critic network ; For agent i in state Take action Rewards obtained; For the environmental dynamics model Next, next state expected value; For the target Critic network in state The state value function under ; In state The state value function under ; For the strategy Next, action expected value; For the target Critic network in state and actions The action-value function under For the target Critic network in state Next Generate Action The probability distribution of is the action of agent i at time step (t+1); is the state of agent i at time step (t+1).
[0140] In the experience replay pool D, the algorithm uses random sampling to extract fixed-size sample batches. Perform approximate calculations. In the specific implementation, agent i will observe the value of the next moment Input to the target Actor network , and then randomly sample the action probability distribution based on the network output to determine the subsequent action strategy of the agent It should be noted that, when running online, each agent is based on local observations , generating actions through the Actor network , without accessing other intelligent agent information, ensuring the scalability of power grid scheduling.
[0141] In some embodiments, the steps of defining the state space and action space of the regional agent at different time scale scheduling levels and establishing an initial optimization scheduling model for each regional agent based on a centralized training decentralized execution architecture include:
[0142] The day-ahead state space is constructed based on the active power output data of micro-turbines, photovoltaic power output forecast values, wind power output forecast values, and load forecast data of each autonomous region.
[0143] Define all the control actions taken by the regional agent in the day-ahead phase to form the day-ahead action space;
[0144] Taking minimization of the day-ahead total operating cost of the distribution network as the optimization objective, a day-ahead centralized optimization scheduling model is constructed according to the day-ahead state space and the day-ahead action space;
[0145] The intraday state space is constructed based on the real-time data of regional wind and solar power output, regional load demand, and regional energy storage charge state data within the autonomous region under the jurisdiction of each regional intelligent agent, and the gas turbine output adjustment amount and energy storage charge and discharge adjustment amount are defined as the intraday action space;
[0146] The day-ahead optimization scheduling strategy obtained by solving the day-ahead centralized optimization scheduling model is used as the intraday initial condition, and the minimization of the operating cost within the autonomous region is taken as the optimization goal. An intraday distributed rolling optimization scheduling model is constructed according to the intraday state space and the intraday action space.
[0147] Construct a real-time state space based on the real-time operating status of the distribution network in the real-time phase, and define the energy storage charging and discharging power as the real-time action space;
[0148] The intraday scheduling strategy output by the intraday distributed rolling optimization scheduling model is used as the real-time initial condition, minimizing the power imbalance in the autonomous area is taken as the optimization goal, and a real-time optimization scheduling model is constructed according to the real-time state space and the real-time action space.
[0149] Specifically, when constructing the day-ahead optimization model, this embodiment selects the active power output of k micro-gas turbine units, the predicted photovoltaic power output, the predicted wind power output, the predicted load, and the power purchase price of the distribution network as state variables. Based on these state variables, a day-ahead state space is established. This day-ahead state space includes key factors that affect the day-ahead scheduling decision. At the same time, this embodiment sets the control objects for day-ahead regulation to mainly include energy storage systems, distributed generators, flexible loads, and photovoltaic and wind power inverters. For these control objects, this embodiment defines the day-ahead action space of the energy storage system, distributed generators, flexible loads, and photovoltaic and wind power inverters. In addition, in the soft action-critic algorithm, the policy network learns how to maximize the reward. To minimize the objective function, this embodiment sets the reward value as a negative objective function. This design enables the policy network to adjust toward the optimization goal during training, thereby gradually approaching the optimal solution. By setting a reward function, this embodiment can guide the intelligent agent to make decisions in the state space that are more conducive to achieving the optimization goal. The mathematical expression of the reward function is:
[0150]
[0151] Where, is the reward function; is the objective function of the total day-ahead operating cost of the distribution network.
[0152] In the day-ahead phase, this embodiment uses a centralized optimization method to optimize the distribution network. Each intelligent agent aims to optimize the economic efficiency of the 24-hour operation of the grid. On the basis of ensuring the safe operation of the grid, it regulates the micro gas turbine, distributed energy storage, and flexible load on an hourly basis to ensure the economic operation of the grid. The day-ahead economic optimization scheduling is based on the full absorption of wind and solar power generation. The day-ahead total operating cost objective function of the distribution network is:
[0153]
[0154]
[0155]
[0156]
[0157]
[0158]
[0159] Where T is the total number of time periods in the regulation cycle; is the power generation cost of the distribution network at time t; is the active power loss cost of the distribution network at time t; is the energy storage system regulation cost at time t; is the flexible load control cost at time t; is the penalty cost for violating constraint X at time t; X is the type of constraint (e.g., energy storage state of charge, line power, flexibility capacity, etc.); is a set of constraints, including energy storage SOC, line power, etc. is the unit electricity price; is the exchange power between the distribution network and the upper power grid at time t; is the electricity price of distributed generators; is the discharge power of the distributed generator at time t; is the loss cost coefficient; is the active network loss of the distribution network at time t; is the energy storage control cost coefficient; The charging efficiency of the energy storage system; is the charging power of the energy storage system at time t; is the discharge power of the energy storage system at time t; is the discharge efficiency of the energy storage system; is the total number of flexible loads; is the cost coefficient for flexible load regulation; is the original power of the i-th type flexible load at time t; is the regulated power of the i-th type of flexible load at time t; is the lower limit penalty coefficient of constraint condition X; is the minimum allowed value of the constraint condition X; is the actual value of the constraint X at time t; is the maximum allowed value of the constraint condition X; is the upper penalty coefficient of constraint X.
[0160] During the day-ahead optimization of the distribution network, constraint adjustment is key to ensuring system stability and efficient operation. The constraints for the day-ahead scheduling interval optimization problem include power flow constraints, voltage amplitude constraints, micro-turbine output constraints, energy storage charging and discharging power and capacity constraints, flexible load constraints, line transmission power constraints, and reserve supply capacity constraints. To simplify the expression, this embodiment omits the scheduling period t in the constraints. The mathematical expression for the power flow constraint is:
[0161]
[0162] Where, is the active power injection of node i; is the voltage amplitude of node i; is the voltage amplitude at node j; is the conductance between node i and node j; is the voltage phase angle difference between nodes i and j; is the susceptance between nodes i and j; is the reactive power injection of node i; j is the node index connected to node i; N is the set of nodes connected to node i.
[0163] The mathematical expression of voltage amplitude constraint is:
[0164]
[0165] Where, is the lower limit of the voltage amplitude at node j; is the voltage amplitude of node j at time t; is the upper limit of the voltage amplitude at node j.
[0166] The mathematical expression of the micro gas turbine output constraint is:
[0167]
[0168]
[0169] Where, is the lower limit of the active power output of the micro gas turbine; is the active power output of the micro gas turbine at time t; is the active power output of the micro gas turbine at time (t-1); is the change in active power output of the micro gas turbine at time t; is the upper limit of the active power output of the micro gas turbine; is the maximum downward ramp power of the microturbine; is the maximum upward ramp power of the microturbine; is the lower limit of reactive power output of micro gas turbine; is the reactive power output of the micro gas turbine at time t; It is the upper limit of reactive power output of micro gas turbine.
[0170] The mathematical expression of energy storage charging and discharging power and capacity constraints is:
[0171]
[0172]
[0173] Where, It is the negative upper limit of the charging and discharging power of the energy storage system; is the charge and discharge power of the energy storage system at time t; The upper limit of the charge and discharge power of the energy storage system (discharge is positive, charge is negative); The minimum state of charge for the energy storage system; is the state of charge of the energy storage system at time (t-1); is the scheduling time interval; It is the maximum state of charge of the energy storage system.
[0174] There are various forms of flexible loads on the user side. Among them, interruptible loads can be used as a virtual power source, and their control strategy can be considered equally with distributed power sources. Shiftable loads can be regarded as a special type of transferable load. Therefore, this paper takes transferable loads as an example to explore the control mode of flexible loads in the distribution network. The constraint condition is that the total load remains unchanged during the control cycle. The mathematical expression of the flexible load constraint is:
[0175]
[0176] In addition, the regulation capacity of the transferable load in each regulation period is subject to certain restrictions, that is, there are upper and lower limits on the power of the transferable load:
[0177]
[0178] Where, is the regulated power of the i-th type of flexible load at time t; is the original power of the i-th type flexible load at time t; is the lower power limit of the i-th type flexible load at time t; is the upper limit of the power of the i-th type flexible load at time t.
[0179] The mathematical expression of line transmission power constraint is:
[0180]
[0181] Where, is the square of the active transmission power of the line between node i and node j at time t; is the square of the reactive power transferred between nodes i and j at time t; is the square of the apparent power of the line between node i and node j at time t.
[0182] In the intraday phase, the scheduling strategy of the day-ahead phase is modified by distributed micro-turbines and energy storage. Therefore, in the day-ahead phase, the source of reserve capacity is the reserve capacity of micro-turbines and energy storage, and the reserve capacity constraints of micro-turbines and energy storage are:
[0183]
[0184]
[0185]
[0186]
[0187]
[0188]
[0189]
[0190]
[0191]
[0192] Where, Upward reserve capacity for energy storage; upward reserve capacity for gas turbines; is the change in active power output of the micro gas turbine; For the time interval The upward reserve capacity required by the system within the range; Provide downward reserve capacity for energy storage; is the change in active power of the energy storage system; Downward reserve capacity for gas turbines; is the maximum downward ramp power of the microturbine; is the actual active power output of the micro gas turbine; is the lower limit of the active power output of the micro gas turbine; For the time interval The downward reserve capacity required by the system within the range.
[0193] In the intraday stage, this embodiment divides the distribution network into several autonomous regions according to the geographical environment and regional absorption conditions. Each autonomous region sets up a regional intelligent agent. Each intelligent agent issues a dispatching instruction to the controllable equipment in the region through the collected ultra-short-term power generation forecast, load forecast data and equipment status information of the renewable energy in the region. The neural network of each regional intelligent agent adopts a centralized training and decentralized execution framework. The network parameters of each intelligent agent are obtained during the centralized training process, and then the trained intelligent agent is distributedly controlled. In the t-th scheduling period, the regional intelligent agent uses the strategy network to output the dispatching decision and obtain the regional reward according to the status of the region. Since the intraday optimization is adjusted on the basis of the day-ahead optimization, its control range is relatively small. In the day-ahead optimization process, A certain amount of control capacity is reserved within the domain, so that the region can fully cope with the challenges and fluctuations brought by the uncertainty of new energy. The consumption of new energy is adjusted to a certain extent with the help of day-ahead planning on the basis of the lowest regional operating cost. The purpose of intraday optimization is to correct the energy supply shortage and new energy waste caused by day-ahead forecast errors. Intraday optimization ensures the balance of power grid supply and demand and promotes the local consumption of new energy by adjusting micro gas turbines and energy storage systems. While ensuring real-time backup capacity, it also ensures the economic operation of the power grid. The objective function of the intraday distributed rolling optimization scheduling model mainly includes the adjustment cost of the micro gas turbine and the adjustment cost of the energy storage system to ensure the long-term economic efficiency of the power grid operation. The objective function of the intraday distributed rolling optimization scheduling model is specifically:
[0194]
[0195]
[0196]
[0197] Where, It is the objective function of the intraday distributed rolling optimization scheduling model; Adjusting costs for microturbines; Adjusting costs for energy storage systems; is the actual output of the gas turbine at time t; is the planned output deviation of the gas turbine at time t; Adjust cost factors for gas turbines; Adjusting cost factors for energy storage; is the actual charging and discharging power of the energy storage at time t; is the charge and discharge deviation of the energy storage plan at time t.
[0198] The intraday flexibility capacity constraint mainly serves as a reserve capacity constraint for the energy storage system, ensuring that the energy storage system can flexibly smooth out wind and solar fluctuations during real-time optimization.
[0199]
[0200] Where, Penalize costs for flexibility.
[0201] The constraints of the intraday distributed rolling optimization scheduling model include gas turbine constraints, energy storage constraints, and reserve capacity constraints. The specific gas turbine constraints are:
[0202]
[0203] Where, is the upper limit of the change in active power output of the micro gas turbine at time t; It is the lower limit of the change in active power output of the micro gas turbine at time t.
[0204] The energy storage constraints are as follows:
[0205]
[0206] Where, is the change in charging and discharging power of the energy storage system at time t.
[0207] The specific reserve capacity constraints are:
[0208]
[0209]
[0210] Intraday economic optimization is mainly based on regional internal regulation. The state variables of the agent include the active power output of the micro gas turbine group in the kth region, the photovoltaic output forecast value, the wind power output forecast value, the regional load forecast value, and the energy storage capacity in the region at the previous moment. Based on these state variables, the intraday state space of the agent can be constructed. The control objects of intraday regulation are mainly the energy storage system and distributed generators. The energy storage system and distributed generator control objects are used to construct the intraday action space. In the multi-agent system, the reward value of each agent is determined based on its performance in the region. The formula for calculating the intraday reward value of the agent is:
[0211]
[0212] Where, is the intraday reward value; is the economic weight coefficient; is the flexibility weight coefficient. and Balance the importance of different optimization objectives in the reward function and guide the agent to make decisions that are more conducive to achieving intraday economic optimization.
[0213] Real-time optimization is mainly used to smooth out power fluctuations caused by wind and solar distributed energy and loads, and to further supplement the source-load forecast error data. This embodiment uses high-precision ultra-short-term forecast data with a time scale of 5 minutes as input to perform real-time optimization and regulation of the distribution network. The objective function of the real-time optimization scheduling model is specifically:
[0214]
[0215]
[0216] Where, Penalize costs for real-time power fluctuations; is the grid penalty coefficient at time t; is the internal power imbalance of region k at time t after grid regulation, which mainly comes from the imbalance between demand power and generated power caused by fluctuations in power output and load demand; is the fluctuation of wind power output in region k; is the fluctuation of photovoltaic output in region k; is the load demand fluctuation in area k.
[0217] Real-time optimization is mainly based on regional internal regulation. In terms of real-time state space construction, this embodiment selects photovoltaic output, wind power output, regional load demand, and the energy storage capacity in the region at the previous moment as real-time state variables to establish a real-time state space. These real-time state variables fully reflect the real-time energy status and load demand in the region, providing basic information for real-time regulation. This embodiment sets the control objects of real-time regulation as energy storage systems and distributed generators, and determines the action space based on the control objects of real-time regulation. At the same time, in the real-time reward function, the multi-agent reward value is determined based on the performance within the region. The real-time reward value of the agent is specifically:
[0218]
[0219] Where, is the real-time reward value.
[0220] This embodiment uses a real-time reward function to guide the intelligent agent to make decisions that are conducive to the optimal allocation of regional energy in real-time regulation. Through iterative optimization, a multi-scale full-cycle collaborative scheduling strategy that meets safety constraints is obtained. Based on the multi-scale full-cycle collaborative scheduling strategy, full-time equipment control instructions (such as energy storage charging and discharging power and gas turbine output) are generated. The full-time equipment control instructions are issued in real time to the micro gas turbine, energy storage system, and flexible load controller for corresponding control and scheduling. To verify the effectiveness of the proposed multi-time scale control method for the distribution network, this embodiment uses an improved IEEE-33 node expansion system for simulation verification. The voltage reference value of the distribution network is 12.66kV. Figure 6 This is a schematic diagram of the IEEE-33 node provided by an embodiment of the present invention. In this embodiment, the test system is divided into Figure 7 In the three autonomous regions shown in Figure 1, the photovoltaic power generation system is distributed at nodes {17, 32}, the wind power is located at nodes {19, 30}, the micro gas turbine is distributed at nodes {10, 24, 28}, the electric energy storage is also distributed at nodes {10, 24, 28}, and the flexible load is located at node 33. The source-load prediction data and equipment operating parameters are shown in Figure 11. Figure 8 ,The performance of the deep reinforcement learning algorithm is closely related to ,network parameters.,In the simulation example, the day-ahead and intraday state observations are represented as 9-dimensional ,array vectors, the day-ahead action is represented as a 9-dimensional array vector, and the intraday action is represented as a 2-dimensional ,array vector.
[0221] The SAC algorithm is an efficient deep reinforcement learning method. It achieves exploration-exploitation trade-off optimization by simultaneously maximizing the expected reward and entropy value of the strategy, thereby significantly improving learning efficiency and performance. This embodiment introduces an improved adaptive temperature coefficient adjustment mechanism into the SAC framework, so that the temperature coefficient can be dynamically adjusted to better adapt to environmental uncertainty and specific task requirements, achieving faster convergence speed and higher stability. Then, this embodiment compares and analyzes the standard SAC algorithm and the improved algorithm under the same environmental conditions, and statistics the cumulative reward value of the agent during the training process. The agent undergoes 1,000 rounds of training. The improved SAC algorithm obtains significantly higher reward values within 100 rounds, while the standard SAC algorithm requires approximately 400 rounds of training to achieve a similar performance level. This shows that the improved SAC algorithm is superior to the standard algorithm in both convergence speed and learning efficiency.
[0222] Figures 9 to 14 The results of optimizing the environment by using the multi-time scale optimization scheduling method proposed in this embodiment are shown. In the day-ahead scheduling stage, Figure 9As shown in the figure, the source-load-storage coordinated operation is optimized based on the forecast data, and the optimization decision of this stage is realized through the centralized scheduling method. The system can obtain the optimal output plan of various flexible resources. In the day-ahead scheduling model, the source-load-storage coordinated optimization significantly improves the utilization rate of wind and solar resources. The active interaction of flexible resources effectively controls the overall operating cost of the system in the intraday stage. Figure 10 It can be seen that the optimized scheduling redistributed the adjustable load from the original 13:00-18:00 period to the 8:00-13:00 period, achieving the goals of peak shaving and valley filling and load balancing. Figure 11 The operating characteristics of equipment in each region during the day-ahead scheduling phase are demonstrated. During the 10:00-20:00 period, the output fluctuations of micro gas turbines and energy storage systems in each region remained at a low level. Specifically, due to the high proportion of distributed renewable energy in the region, the total output of micro gas turbines and energy storage systems was relatively low, which is related to the inherent uncertainty of renewable energy generation.
[0223] It should be noted that Figure 11 The data showed that both regions experienced a rapid increase in load between 4:00 and 6:00. To meet the high load demand during this period, the output of the micro gas turbines and energy storage systems in each region increased significantly. As the load uncertainty increased, the total system output gradually decreased to maintain flexible regulation capabilities. In addition, the charging behavior of the energy storage system showed significant electricity price response characteristics, especially before 5:00 and after 20:00.
[0224] Figure 12 The distribution of system flexibility supply and demand within a day is shown. Figure 12 It can be seen that the system flexibility demand is always higher than the supply level. After the demand curve reaches its peak near midday, the supply curve shows a similar trend of change. This phenomenon may be due to the Gaussian distribution characteristics of photovoltaic, wind power and adjustable load forecast data. Compared with the intervals on both sides, the forecast fluctuation range in the central interval is more significant. In addition, since photovoltaic power generation is concentrated in the period of 6:00 to 18:00, the power fluctuation in this period is more prominent.
[0225] The intraday optimization process is implemented based on the existing day-ahead scheduling strategy. This embodiment dynamically adjusts the operation plans of the micro gas turbine and the energy storage system to maximize the consumption of renewable energy while ensuring the balance of power grid supply and demand. Figure 13The results show the difference in output changes of distributed micro gas turbines in different regions during day-ahead and intraday scheduling, reflecting the impact of load fluctuations. During the 9:00-13:00 period, the actual load demand is significantly lower than the day-ahead forecast value, and the unit output is reduced accordingly; while in other periods, the load demand exceeds the forecast value, resulting in an increase in unit output during intraday scheduling.
[0226] The real-time optimization phase focuses on rapid resource re-optimization of the distributed energy storage system. This phase is based on the daily rolling scheduling plan to implement optimization and adapt to random supply and demand fluctuations by minimizing the adjustment range. Figure 14 The real-time regulation effect of the energy storage system is demonstrated. It can be seen that there is a significant deviation between the real-time scheduling results and the day-ahead plan. This is mainly due to the large difference between the day-ahead forecast and the intraday forecast data. Real-time optimization focuses on solving the problem of insufficient renewable energy consumption caused by supply and demand fluctuations by introducing energy storage strategy adjustments with higher frequency but smaller amplitude.
[0227] An embodiment of the present invention provides a multi-scale optimization and scheduling method for a distribution network based on deep reinforcement learning. The method uses probability box theory to model the probability distribution of the source-load prediction error on the scheduling day according to the source-load scheduling daily forecast data and the actual source-load data, and obtains the net load uncertainty interval; the distribution network scheduling cycle is divided into multiple time-scale levels according to the response characteristics of the distribution network equipment to obtain a multi-time-scale scheduling level; based on the multi-time-scale scheduling level and the distribution network partition information, a multi-agent deep reinforcement learning algorithm is used to construct a power grid optimization and scheduling model based on a centralized training and decentralized execution architecture; the power grid optimization and scheduling model includes a day-ahead centralized optimization and scheduling model, an intra-day distributed rolling optimization and scheduling model, and a real-time optimization and scheduling model; with the minimization of the day-ahead total operating cost as the optimization goal, the day-ahead scheduling strategy is obtained by solving the power grid optimization and scheduling model according to the net load uncertainty interval and the pre-acquired self-regulating temperature coefficient; based on the day-ahead scheduling strategy, multi-scale rolling optimization is performed using the real-time operating status of the distribution network equipment and ultra-short-term prediction deviation data, and a multi-scale full-cycle collaborative scheduling strategy is obtained through iterative optimization. Compared with the existing technology, this method realizes the flexible resource optimization configuration of the distribution network in day-ahead, intraday and real-time scheduling by comprehensively utilizing probability box theory modeling, multi-time scale hierarchical division and multi-agent deep reinforcement learning algorithm to construct a power grid optimization scheduling model, enhances the adaptability to source and load uncertainty, and significantly improves the operation stability and economy of the distribution network at different time scales, making the scheduling strategy more in line with actual operation needs and ensuring the safe, efficient and economical operation of the distribution network.
[0228] It should be noted that the size of the serial numbers of the above-mentioned processes does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiment of this application.
[0229] In one embodiment, Figure 15 As shown, an embodiment of the present invention provides a distribution network multi-scale optimization scheduling system based on deep reinforcement learning, the system comprising:
[0230] The source load analysis module 101 is used to model the probability distribution of the source load forecast error on the scheduling day based on the source load scheduling day forecast data and the actual source load data, using the probability box theory to obtain the net load uncertainty interval;
[0231] A scale division module 102 is used to divide the distribution network dispatching cycle into multiple time scale levels according to the response characteristics of the distribution network equipment to obtain a multi-time scale dispatching level;
[0232] A model building module 103 is configured to construct a power grid optimization scheduling model based on a centralized training and decentralized execution architecture using a multi-agent deep reinforcement learning algorithm based on the multi-timescale scheduling hierarchy and distribution network partition information; the power grid optimization scheduling model includes a day-ahead centralized optimization scheduling model, an intraday distributed rolling optimization scheduling model, and a real-time optimization scheduling model;
[0233] The model solving module 104 is configured to solve the power grid optimization scheduling model to obtain a day-ahead scheduling strategy based on the net load uncertainty interval and the pre-acquired self-regulating temperature coefficient, with minimizing the day-ahead total operating cost as the optimization objective;
[0234] The optimization scheduling module 105 is used to perform multi-scale rolling optimization based on the day-ahead scheduling strategy, using the real-time operating status of distribution network equipment and ultra-short-term forecast deviation data, and obtain a multi-scale full-cycle collaborative scheduling strategy through iterative optimization.
[0235] For the specific definition of a distribution network multi-scale optimization scheduling system based on deep reinforcement learning, please refer to the above-mentioned definition of a distribution network multi-scale optimization scheduling method based on deep reinforcement learning, which will not be repeated here. Those of ordinary skill in the art will appreciate that the various modules and steps described in conjunction with the embodiments disclosed in this application can be implemented in hardware, software, or a combination of both. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.
[0236] An embodiment of the present invention provides a distribution network multi-scale optimization and scheduling system based on deep reinforcement learning. The source-load analysis module of the system uses probability box theory to model the probability distribution of the source-load prediction error on the scheduling day based on the source-load scheduling daily forecast data and actual source-load data, and obtains the net load uncertainty interval; the scale division module divides the distribution network scheduling cycle into multiple time-scale hierarchies according to the response characteristics of the distribution network equipment to obtain a multi-time-scale scheduling hierarchy; the model construction module uses a multi-agent deep reinforcement learning algorithm to construct a power grid optimization and scheduling model based on a centralized training and decentralized execution architecture based on the multi-time-scale scheduling hierarchy and distribution network partition information; the power grid optimization and scheduling model includes a day-ahead centralized optimization and scheduling model, an intra-day distributed rolling optimization and scheduling model, and a real-time optimization and scheduling model; the model solution module takes minimizing the day-ahead total operating cost as the optimization goal, and obtains the day-ahead scheduling strategy through the power grid optimization and scheduling model according to the net load uncertainty interval and the pre-acquired self-regulating temperature coefficient; the optimization and scheduling module uses the real-time operating status of the distribution network equipment and the ultra-short-term prediction deviation data to perform multi-scale rolling optimization based on the day-ahead scheduling strategy, and obtains the multi-scale full-cycle collaborative scheduling strategy through iterative optimization. Compared with existing technologies, this system achieves flexible resource optimization configuration of distribution networks in day-ahead, intraday and real-time scheduling by comprehensively utilizing probability box theory modeling, multi-time scale hierarchical division and multi-agent deep reinforcement learning algorithm to construct a power grid optimization scheduling model, enhances the adaptability to source and load uncertainty, and significantly improves the operating stability and economy of the distribution network at different time scales, making the scheduling strategy more in line with actual operating needs and ensuring the safe, efficient and economical operation of the distribution network.
[0237] The above-described embodiments merely represent several preferred implementations of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the patent. It should be noted that a person skilled in the art could make several improvements and substitutions without departing from the technical principles of the present invention, and these improvements and substitutions should also be considered within the scope of protection of the present application. Therefore, the scope of protection of the present patent application shall be based on the scope of protection of the claims.
Claims
1. A multi-scale optimization scheduling method for distribution network based on deep reinforcement learning, characterized in that: The following steps are involved: Based on the forecast data and actual source load data of the source load dispatch day, the probability box theory is used to model the probability distribution of the source load forecast error on the dispatch day and obtain the net load uncertainty interval. According to the response characteristics of distribution network equipment, the distribution network dispatch cycle is divided into multiple time scale levels to obtain a multi-time scale dispatch level. Based on the multi-timescale scheduling hierarchy and distribution network partition information, a multi-agent deep reinforcement learning algorithm is used to construct a power grid optimization scheduling model based on a centralized training and decentralized execution architecture; the power grid optimization scheduling model includes a day-ahead centralized optimization scheduling model, an intraday distributed rolling optimization scheduling model, and a real-time optimization scheduling model; Taking minimizing the total operating cost of the day-ahead as the optimization goal, according to the net load uncertainty interval and the pre-acquired self-regulating temperature coefficient, the day-ahead dispatch strategy is obtained by solving the power grid optimization dispatch model; Based on the day-ahead dispatch strategy, multi-scale rolling optimization is performed using the real-time operating status of distribution network equipment and ultra-short-term forecast deviation data, and a multi-scale full-cycle collaborative dispatch strategy is obtained through iterative optimization; The steps of constructing a power grid optimization scheduling model based on a centralized training and decentralized execution architecture using a multi-agent deep reinforcement learning algorithm based on the multi-timescale scheduling hierarchy and distribution network partition information include: The distribution network is divided into multiple autonomous regions according to its geographical structure and the distribution of its equipment. A regional agent is deployed for each autonomous region. Each regional agent has an independent actor network and critic network. Define the state space and action space of regional agents at different time-scale scheduling levels, and establish an initial optimization scheduling model for each regional agent based on a centralized training and decentralized execution architecture; During the centralized training phase, the actor network parameters and critic network parameters of each regional agent are randomly initialized; During the training process, the state information, action information, and reward information generated by the interaction between the regional agent and the environment are stored in the replay buffer. Each time the actor network parameters and the critic network parameters are updated, batches of data are randomly sampled from the replay buffer to form a training dataset. Based on the training dataset, updating the critic network parameters by minimizing the soft Bellman residual to obtain the optimal action value estimate; The action entropy is calculated based on the action selection probability distribution output by the actor network, and the self-regulating temperature coefficient is calculated based on the uncertainty quantification parameter and action entropy obtained in advance; According to the optimal action value estimate and the self-regulating temperature coefficient, the actor network parameters of the regional agent are updated by gradient back propagation to obtain the global coordination strategy parameters; In the decentralized execution phase, the updated actor network parameters are used to generate partition-independent action instructions based on the real-time local observation status of the agents in each region; The initial optimization scheduling model is iteratively optimized at multiple time scales according to the global coordination strategy parameters and independent action instructions to obtain a power grid optimization scheduling model.
2. A distribution network multi-scale optimization scheduling method based on deep reinforcement learning according to claim 1, characterized in that: The step of modeling the probability distribution of the source load forecast error on the dispatching day using the probability box theory based on the source load dispatching day forecast data and the actual source load data to obtain the net load uncertainty interval includes: Calculate the deviation between the daily source load scheduling forecast data and the actual source load data to obtain the source load forecast error, and use the source load forecast error as the uncertainty component of the random variable in the probability box theory; Performing a probability statistical analysis on the source-load prediction error to obtain a probability distribution characteristic parameter boundary of the source-load prediction error; Constructing a cumulative probability distribution function of the source-load prediction error according to the probability distribution characteristic parameter boundary, and obtaining a cumulative probability distribution boundary of the cumulative probability distribution function; Calculating the power prediction error boundary of each source load according to the cumulative probability distribution boundary, and obtaining the source load power uncertainty interval according to the power prediction error boundary of each source load and the source load scheduling daily forecast data; The net load prediction error boundary is calculated according to the source load power uncertainty interval to generate a net load uncertainty interval.
3. The multi-scale optimization scheduling method for distribution network based on deep reinforcement learning according to claim 1, characterized in that: The multi-time-scale scheduling level includes at least three time-scale scheduling levels, which are a day-ahead scheduling level, an intraday scheduling level, and a real-time scheduling level.
4. The multi-scale optimization scheduling method for distribution network based on deep reinforcement learning according to claim 1, characterized in that: The steps of defining the state space and action space of the regional agent at different time scale scheduling levels and establishing an initial optimization scheduling model for each regional agent based on a centralized training decentralized execution architecture include: The day-ahead state space is constructed based on the active power output data of micro-turbines, photovoltaic power output forecast values, wind power output forecast values, and load forecast data of each autonomous region. Define all the control actions taken by the regional agent in the day-ahead phase to form the day-ahead action space; Taking minimization of the day-ahead total operating cost of the distribution network as the optimization objective, a day-ahead centralized optimization scheduling model is constructed according to the day-ahead state space and the day-ahead action space; The intraday state space is constructed based on the real-time data of regional wind and solar power output, regional load demand, and regional energy storage charge state data within the autonomous region under the jurisdiction of each regional intelligent agent, and the gas turbine output adjustment amount and energy storage charge and discharge adjustment amount are defined as the intraday action space; The day-ahead optimization scheduling strategy obtained by solving the day-ahead centralized optimization scheduling model is used as the intraday initial condition, and the minimization of the operating cost within the autonomous region is taken as the optimization goal. An intraday distributed rolling optimization scheduling model is constructed according to the intraday state space and the intraday action space. Construct a real-time state space based on the real-time operating status of the distribution network in the real-time phase, and define the energy storage charging and discharging power as the real-time action space; The intraday scheduling strategy output by the intraday distributed rolling optimization scheduling model is used as the real-time initial condition, minimizing the power imbalance in the autonomous area is taken as the optimization goal, and a real-time optimization scheduling model is constructed according to the real-time state space and the real-time action space.
5. The method for multi-scale optimization and scheduling of distribution networks based on deep reinforcement learning according to claim 1, characterized in that: The action entropy is defined as the negative logarithmic mean of the action selection probability distribution.
6. The method for multi-scale optimization and scheduling of distribution networks based on deep reinforcement learning according to claim 4, characterized in that: The total day-ahead operating cost is the sum of the distribution network's power generation cost, distribution network active power loss cost, energy storage system regulation cost, flexible load regulation cost, and flexible load penalty cost during the distribution network's regulation cycle.
7. The method for multi-scale optimization and scheduling of distribution networks based on deep reinforcement learning according to claim 4, characterized in that: The operating cost within the autonomous region is the sum of the gas turbine adjustment cost and the energy storage adjustment cost of the distribution network during the regulation cycle.
8. The method for multi-scale optimization and scheduling of distribution networks based on deep reinforcement learning according to claim 4, characterized in that: The power imbalance amount within the autonomous region is the product of the net power fluctuation value of the distribution network during the regulation cycle and the penalty weight.
9. A multi-scale optimization and dispatching system for distribution networks based on deep reinforcement learning, characterized in that: The system comprises: The source-load analysis module is used to model the probability distribution of the source-load forecast error on the dispatch day based on the source-load dispatch day forecast data and actual source-load data, and obtain the net load uncertainty interval; A scale division module is used to divide the distribution network dispatching cycle into multiple time scale levels according to the response characteristics of the distribution network equipment to obtain a multi-time scale dispatching level; A model building module is used to build a power grid optimization scheduling model based on a centralized training and decentralized execution architecture using a multi-agent deep reinforcement learning algorithm based on the multi-timescale scheduling hierarchy and distribution network partition information; the power grid optimization scheduling model includes a day-ahead centralized optimization scheduling model, an intraday distributed rolling optimization scheduling model, and a real-time optimization scheduling model; a model solving module, configured to obtain a day-ahead dispatching strategy by solving the power grid optimization dispatching model based on the net load uncertainty interval and the pre-acquired self-regulating temperature coefficient, with minimizing the day-ahead total operating cost as the optimization objective; An optimization scheduling module is used to perform multi-scale rolling optimization based on the day-ahead scheduling strategy, using the real-time operating status of distribution network equipment and ultra-short-term forecast deviation data, and obtain a multi-scale full-cycle collaborative scheduling strategy through iterative optimization; Wherein, the model building module is specifically used to: The distribution network is divided into multiple autonomous regions according to its geographical structure and the distribution of its equipment. A regional agent is deployed for each autonomous region. Each regional agent has an independent actor network and critic network. Define the state space and action space of regional agents at different time-scale scheduling levels, and establish an initial optimization scheduling model for each regional agent based on a centralized training and decentralized execution architecture; During the centralized training phase, the actor network parameters and critic network parameters of each regional agent are randomly initialized; During the training process, the state information, action information, and reward information generated by the interaction between the regional agent and the environment are stored in the replay buffer. Each time the actor network parameters and the critic network parameters are updated, batches of data are randomly sampled from the replay buffer to form a training dataset. Based on the training dataset, updating the critic network parameters by minimizing the soft Bellman residual to obtain the optimal action value estimate; The action entropy is calculated based on the action selection probability distribution output by the actor network, and the self-regulating temperature coefficient is calculated based on the uncertainty quantification parameter and action entropy obtained in advance; According to the optimal action value estimate and the self-regulating temperature coefficient, the actor network parameters of the regional agent are updated by gradient back propagation to obtain the global coordination strategy parameters; In the decentralized execution phase, the updated actor network parameters are used to generate partition-independent action instructions based on the real-time local observation status of the agents in each region; The initial optimization scheduling model is iteratively optimized at multiple time scales according to the global coordination strategy parameters and independent action instructions to obtain a power grid optimization scheduling model.
Citation Information
Patent Citations
Optimized scheduling method and terminal based on probability box and conditional value-at-risk
CN115632438A
Power distribution network load prediction and electric quantity balance optimization method and system based on big data
CN118889419A