An active power distribution network deep reinforcement learning real-time scheduling method and system
By employing an improved deep deterministic strategy gradient algorithm and event-driven mechanism, the scheduling problem of uncertain new energy sources and loads in active distribution networks was solved, achieving stable, economical, and low-carbon operation of active distribution networks and improving the accuracy and efficiency of scheduling strategies.
Patent Information
- Application Number
- CN202210999422.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-08-19
- Publication Date
- 2025-11-07
- Estimated Expiration
- 2042-08-19
AI Technical Summary
Existing technologies struggle to effectively handle the complex uncertainties of new energy sources and loads in active power distribution networks. Traditional deep reinforcement learning methods lead to discretization errors in the state-action space and cannot quickly and accurately compensate for power deviations, affecting the system's economic efficiency and low-carbon attributes.
An improved deep deterministic policy gradient algorithm is used to construct a multi-objective optimization scheduling model. Combined with a deep reinforcement learning framework, the model is trained on continuous states and action spaces using the improved deep deterministic policy gradient algorithm. By combining an event-driven mechanism and schedulability evaluation, a real-time scheduling strategy is realized.
It enables precise scheduling of active distribution networks under multiple uncertainties, improves system stability and economy, ensures reliable power supply and low-carbon operation, shortens calculation time, and improves the efficiency and accuracy of scheduling strategy implementation.
Smart Images

Figure CN115207977B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the technical field of energy internet, and particularly relates to a kind of active power distribution network deep reinforcement learning real-time scheduling method and system. BACKGROUND
[0002] The power distribution network is an intermediate link that receives electric energy from the power transmission system or power plant and then distributes it to power users, and plays an important role in supplying power to users. The construction of a regional flexible, strong and reliable high-elasticity power distribution network can effectively improve the consumption capacity of green and clean new energy, and provides a strong cornerstone for continuously optimizing the business environment and improving the "power availability" index.
[0003] In the face of the diversification and large-scale access of distributed power sources, the concept of active intelligent power distribution network is proposed. The active intelligent power distribution network is based on the main functional framework of the intelligent power distribution network, uses advanced information communication technology and automation technology, and is a new type of intelligent power distribution system that faces the diversified and high-penetration distributed energy access. Its basic characteristics are to cope with the high-penetration access of distributed energy and to ensure the reliability of power supply and the safety of the system. The addition of distributed energy changes the power distribution network from a passive network to an active network, and changes the single role of electric energy distribution to a new type of electric power exchange system of electric energy production, storage, transmission and distribution. The structure and operation control mode of the power distribution network will change greatly, and the installation location and operation mode of the distributed power source will have a great impact on the power flow, voltage, network loss and safe and stable operation of the power distribution network. In order to adapt to the development trend of large-scale grid connection of distributed power sources and the continuous increase of the proportion of direct current load, the development of AC / DC hybrid power distribution network has become an important direction of current power distribution network research. The AC / DC power distribution network can fully exert the advantages and potential power supply capacity of the DC link under the premise of retaining part of the AC system, and has the advantages of high distributed energy accommodation capacity, flexible and diverse network formation and control mode of the power distribution network, etc.
[0004] Reinforcement learning is a model-free method that does not rely on the distribution knowledge of uncertainty, so it does not need to predict or model the source and load in advance as traditional methods do. In recent years, the deep reinforcement learning model combined with deep neural networks and reinforcement learning has better adaptive learning ability and optimization decision-making ability for non-convex and nonlinear problems, and has been gradually applied in the optimization and scheduling of power systems. The dynamic scheduling problem of the active power distribution network is a stochastic sequential decision problem, and reinforcement learning, as an important machine learning method, focuses on how an agent takes action in an environment to obtain the maximum cumulative reward, which is essentially consistent with the design goal of dynamic economic dispatching of the active power distribution network, i.e. focusing on how the active power distribution network makes scheduling decisions to obtain the optimal operation cost of the system at a certain scheduling stage. Therefore, the research on real-time scheduling strategy of the active power distribution network under deep reinforcement learning has a prominent practical significance. SUMMARY
[0005] The technical problem to be solved by the present application is to provide a kind of active power distribution network deep reinforcement learning real-time scheduling method and system for the above-mentioned deficiencies in the prior art, provide real-time scheduling plan arrangement under multiple uncertainties for active power distribution network, to solve the complex uncertainty modeling of new energy and load in active power distribution network, the error caused by the discretization of state-action space of traditional deep reinforcement learning method, and the technical problems that power deviation cannot be quickly and accurately compensated in active power distribution network considering the economy and low carbon property of equipment.
[0006] The present application adopts the following technical solutions:
[0007] An active power distribution network deep reinforcement learning real-time scheduling method comprises the following steps:
[0008] S1, establish the equipment operation model of photovoltaic unit, combined heat and power unit, air source heat pump, electric gasification equipment and electric refrigeration equipment in active power distribution network, and construct multi-objective optimization scheduling model considering economy and low carbon property;
[0009] S2, the multi-objective optimization scheduling model obtained in step S1 is expressed as the framework of reinforcement learning, and the improved deep deterministic policy gradient algorithm is used to train and online schedule the multi-objective optimization scheduling model under continuous state and action space, to obtain the benchmark result of intraday scheduling;
[0010] S3, according to the intraday scheduling benchmark result obtained in step S2, determine the real-time scheduling strategy based on event-driven, which comprehensively considers the imbalance degree of system supply and demand side and the adjustable capacity of response subject;
[0011] S4, realize the real-time scheduling of active power distribution network based on deep reinforcement learning under multiple uncertainties according to the real-time scheduling strategy determined in step S3.
[0012] Specifically, in step S1, the equipment operation model comprises:
[0013] Combined heat and power unit operation model:
[0014]
[0015]
[0016]
[0017] wherein, is the electric power output of the CHP unit in the t period, is the gas consumption rate of the CHP unit in the t period, is the power generation efficiency of the CHP unit, LHV of natural gas, is the output heat power of the CHP unit at the t th time interval; is the loss coefficient, and are the upper and lower limits of the electric power of the CHP unit, respectively;
[0018] Air source heat pump operation model:
[0019]
[0020]
[0021] wherein, is the load power of the heat pump at the t th time interval, is the heat power of the heat pump at the t th time interval, is the energy efficiency ratio of the heat pump, and are the upper and lower limits of the electric power of the heat pump, respectively;
[0022] Electricity-to-gas equipment operation model:
[0023]
[0024]
[0025] wherein, and are the gas output and the consumed electric power of the electricity-to-gas equipment at the t th time interval, respectively; η P2G is the energy conversion efficiency; and are the upper and lower limits of the consumed electric power of the electricity-to-gas equipment, respectively.
[0026] Electricity-to-refrigeration equipment operation model:
[0027]
[0028]
[0029] wherein, and are the cold output and the consumed electric power of the ice storage air conditioner at the t th time interval, respectively; η ISAC represents the energy conversion efficiency of the ISAC; and are the upper and lower limits of the consumed electric power of the ice storage air conditioner;
[0030] Photovoltaic unit operation model:
[0031]
[0032] wherein, forecasted output of photovoltaic; forecasted output of photovoltaic.
[0033] Specifically, in step S1, the multi-objective optimization scheduling model considering economy and low carbon is:
[0034] minf=C purchase +C carbon
[0035]
[0036]
[0037] Wherein, f is the total cost of the system, C purchase is the operation cost of the system, mainly including the purchase cost of electricity from the upper grid and the purchase cost of gas from the natural gas source, C carbon is the carbon emission cost of the system, including the carbon emission cost of fossil fuel combustion and the carbon emission cost of purchased electricity, T is the total operation time of the system, is the purchase power of electricity at t, is the purchase price of electricity from the upper grid at t, is the purchase power of gas at t, is the purchase price of gas from the natural gas source, is the consumption power of natural gas at t, C coef is the carbon dioxide emission coefficient including multiple coefficients such as low calorific value, unit calorific value carbon content and carbon oxidation rate, λ carbon is the carbon dioxide emission cost, is the carbon dioxide emission factor of purchasing electricity from the upper grid.
[0038] Further, the constraint condition is:
[0039]
[0040] Wherein, is the purchase power of electricity at t, is the planned output of photovoltaic, is the output electric power of the CHP unit in the t period, is the electric power consumed by the ice storage air conditioner at t, is the load power of the heat pump in the t period, is the electric power consumed by the electric-gas conversion equipment at t.
[0041] Specifically, in step S2, the improved deep deterministic policy gradient algorithm is specifically:
[0042] At each time period t, the output of all devices is used as the action space in the active power distribution network; the electric load demand, photovoltaic power generation, park electricity purchase price, park natural gas purchase price and carbon dioxide emission price are used as the environmental state in the active power distribution network, and the goal of the low-carbon economic dispatch of the active power distribution network is to find the optimal strategy π * The parameters are updated using small batch gradient descent to maximize the action value function, and the data obtained by Monte Carlo sampling is used as the environmental state of the active power distribution network to train the parameters of the deep deterministic policy gradient algorithm network offline. When the offline training process is completed, the deep deterministic policy gradient algorithm parameters under the optimal strategy are obtained to solve the low-carbon economic dispatch problem of the active power distribution network in actual operation.
[0043] Specifically, in step S3, based on event-driven, specifically:
[0044] At each dispatch time, the state variables in the active power distribution network are monitored in real time, and when an uncertain disturbance event occurs, the active power distribution network formulates a real-time power distribution plan according to the current collected information and issues control instructions to each device; if the deviation between the daily reference dispatch plan and the real-time device operation plan is less than the set deviation penalty value, no event trigger signal is generated, and the device executes the control instruction according to the daily dispatch result.
[0045] Specifically, in step S3, the imbalance degree of the system supply and demand sides and the response subject adjustable capacity are comprehensively considered, and specifically:
[0046] A unified evaluation standard is set to evaluate the adjustable capacity of each device, and the compensation power corresponding to the adjustable capacity is allocated according to the adjustable capacity. The adjustable capacity quantitatively represents the adjustment capacity of different devices when they are connected to the system to participate in the ultra-short-term regulation and operation of the system, including the reliability factor, positive / negative adjustment capacity, operation and maintenance loss factor and carbon emission factor indexes of the device.
[0047] Specifically, in step S3, the real-time scheduling strategy is specifically:
[0048] When the system is running in real time, the monitored load, photovoltaic and various device state information are read in at each dispatch time, no event trigger signal is generated at present, the control instruction is issued according to the daily 15min dispatch reference value, there is an event trigger signal at present, the real-time power distribution strategy is used, and the dispatch correction is performed according to the preset algorithm.
[0049] Specifically, in step S4, the real-time scheduling based on deep reinforcement learning is specifically:
[0050] In the dynamic control phase, a low-carbon economic dispatch model is established, and dynamic optimization with a scale of 15min is performed in the optimization time domain;
[0051] In the real-time power allocation stage, the time scale is shortened to 5 minutes in the dynamic optimization framework, and real-time power allocation operations are implemented on each response subject according to the comprehensive evaluation system, so that the real-time control of the optimal operation of the active power distribution network is realized.
[0052] In a second aspect, the embodiment of the present application provides an active power distribution network deep reinforcement learning real-time scheduling system, comprising:
[0053] The construction module establishes the equipment operation model of the photovoltaic unit, the combined heat and power unit, the air source heat pump, the electricity-to-gas equipment and the electric refrigeration equipment in the active power distribution network, and constructs a multi-objective optimization scheduling model considering economy and low carbon;
[0054] The improved module expresses the multi-objective optimization scheduling model obtained by the construction module into the framework of reinforcement learning, trains and schedules the multi-objective optimization scheduling model under the continuous state and action space by using the improved deep deterministic policy gradient algorithm, and obtains the benchmark result of the intraday scheduling;
[0055] The strategy module determines the real-time scheduling strategy based on event driving which comprehensively considers the imbalance degree of the system supply and demand sides and the adjustable capacity of the response subject according to the intraday scheduling benchmark result obtained by the improved module;
[0056] The scheduling module realizes the real-time scheduling of the active power distribution network based on deep reinforcement learning under multiple uncertainties according to the real-time scheduling strategy determined by the strategy module.
[0057] Compared with the prior art, the present application has at least the following beneficial effects:
[0058] The active power distribution network deep reinforcement learning real-time scheduling method adopts the improved deep deterministic policy gradient algorithm applied in the real-time scheduling of the active power distribution network, adapts to the uncertain changes of the photovoltaic and load, avoids modeling the complex uncertainty, defines the high-dimensional and continuous state and action space appropriately, avoids the error caused by discretization, and optimizes and improves the model generalization and convergence mechanism. Due to the instability of the reinforcement learning result and the deviation of the prediction, and the requirement of the power system for real-time safety and stability control, a heuristic real-time power allocation strategy based on the adjustable capacity is established on the basis of the scheduling result of the deep deterministic policy gradient algorithm, so as to improve the efficiency and accuracy of the implementation of the scheduling strategy and ensure the continuous and stable operation of the active power distribution network. The implemented case study shows that the intraday-real-time multi-time scale active power distribution network low-carbon economic scheduling method proposed herein can achieve a performance close to that of the optimization method with perfect prediction information under the condition of limited historical data information, and the calculation time is greatly shortened.
[0059] Further, by constructing a multi-objective optimization scheduling model of the active power distribution network, the multi-time scale optimization scheduling of the equipment can be more accurately realized.
[0060] Further, the multi-objective optimization scheduling of the active power distribution network comprehensively considers the economy and low carbon of the operation of the active power distribution network, is conducive to responding to the national energy-saving and emission-reducing policy, and can help the active power distribution network to achieve the purpose of saving cost.
[0061] Further, by constructing the active power balance constraint condition of the active power distribution network, the safe and stable operation of the active power distribution network can be guaranteed, and the active power distribution network is a good access unit for the large power grid.
[0062] Further, by designing the traditional deep deterministic policy gradient algorithm and improving the convergence mechanism, the algorithm converges to the optimal value, the learning strategy of the stochastic gradient reduces the correlation between the training data of the algorithm, and the learning effect of the algorithm is improved.
[0063] Further, the design of the event-driven mechanism enables the response subject in the active power distribution network to quickly and accurately respond to the energy compensation demand, and effectively realizes the safe and stable operation of the active power distribution network.
[0064] Further, the real-time control strategy comprehensively considers the imbalance degree of the supply and demand sides of the active power distribution network and the dispatchable capacity of the response subject, can guarantee the reasonable allocation of the equipment output, and avoid that one device excessively undertakes the system regulation task.
[0065] Further, the designed real-time scheduling strategy meets the reliable supply of the electrical load in the active power distribution network, fully utilizes the new energy output, and comprehensively considers the economy of the operation of the active power distribution network.
[0066] Further, the improved deep deterministic policy gradient algorithm adapts to the uncertain changes of the photovoltaic and load, avoids modeling the complex uncertainty, and the improvement of the model generalization and convergence mechanism makes the algorithm more conducive to convergence. The real-time control strategy comprehensively considers the imbalance degree of the supply and demand sides of the system and the dispatchable capacity of the response subject, the design of the event-driven mechanism enables the response subject to quickly and accurately respond to the energy compensation demand, and effectively realizes the safe and stable operation of the system.
[0067] It can be understood that the beneficial effects of the above-mentioned second aspect can be referred to the related description in the above-mentioned first aspect, which will not be repeated here.
[0068] In summary, the present application avoids modeling complex uncertainty, and the improved model generalization and convergence mechanism makes the algorithm more conducive to convergence, effectively realizes the safe and stable operation of the system, improves the credibility, explainability and robustness of the multi-objective optimization scheduling model of the active power distribution network, ensures the efficiency and precision of the implementation of the scheduling strategy, and guarantees the continuous and stable operation of the active power distribution network.
[0069] The technical solutions of the present application will be further described in detail below with the help of the accompanying drawings and embodiments. BRIEF DESCRIPTION OF DRAWINGS
[0070] Figure 1 The flowchart of the method of the present application is shown in Figure 1.
[0071] Figure 2 The reward curve graph during the training process is shown in Figure 2.
[0072] Figure 3 The intraday scheduling result graph of the active power distribution network based on deep reinforcement learning is shown in Figure 3.
[0073] Figure 4 The real-time generation scheduling result graph of the active power distribution network is shown in Figure 4.
[0074] Figure 5 The real-time load scheduling result graph of the active power distribution network is shown in Figure 5. DETAILED DESCRIPTION
[0075] The technical solutions in the embodiments of the present application will be described clearly and completely below with the help of the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without any creative effort belong to the scope of protection of the present application.
[0076] In the description of the present application, it should be understood that the terms “include” and “contain” indicate the existence of described features, whole, steps, operations, elements and / or components, but do not exclude the existence or addition of one or more other features, whole, steps, operations, elements, components and / or sets thereof.
[0077] It should also be understood that the terms used in the present application specification are only for the purpose of describing specific embodiments and are not intended to limit the present application. As used in the present application specification and the appended claims, unless otherwise clearly indicated by the context, the singular forms “a”, “an” and “the” are intended to include the plural forms.
[0078] It should be further understood that the term "and / or" as used herein refers to and encompasses any and all possible combinations of one or more of the associated listed items, and that the term "or" as used herein refers to and encompasses any one of the associated items and that the term "at least one of" as used herein refers to and encompasses any one of the associated items or combination of one or more of the associated items.
[0079] It should be understood that, although the terms first, second, third, etc. can be used herein to describe various ranges, etc., these ranges should not be limited by these terms. These terms are only used to distinguish one range from another. For example, a first range could be termed a second range without departing from the scope of the embodiments.
[0080] The word "if" as used herein means "when" or "upon" or "in response to a determination" or "in response to a detection," depending on the context. Similarly, the phrase "if it is determined" or "if [a stated condition or event] is detected" can mean "when it is determined" or "in response to a determination" or "when [a stated condition or event] is detected" or "in response to a detection [of a stated condition or event]," depending on the context.
[0081] Various structural diagrams according to the disclosed embodiments of the present application are shown in the accompanying drawings. These diagrams are not drawn to scale, in which certain details are shown in a somewhat exaggerated manner for purposes of clarity and understanding, and certain details can be omitted. The shapes and relative sizes of the various regions, layers, and their relative positions shown in the drawings are merely exemplary, and in actuality can be deviated due to manufacturing tolerances or technical limitations, and regions / layers with different shapes, sizes, and relative positions can be additionally designed by those skilled in the art according to actual needs.
[0082] The application provides a kind of active power distribution network deep reinforcement learning real-time scheduling method, improved deep deterministic policy gradient algorithm is applied to active power distribution network real-time scheduling, the uncertainty of photovoltaic and load is adapted, the modeling of complex uncertainty is avoided, the algorithm used is suitable for the definition of high-dimensional, continuous state, action space, avoid the error caused by discretization, at the same time, model generalization and convergence mechanism are also optimized and improved. Due to the instability of reinforcement learning results and the deviation of prediction, and the requirement of real-time safety and stability control of power system, therefore, on the basis of the scheduling result of DDPG (Deep Deterministic Policy Gradient, deep deterministic policy gradient algorithm), in order to improve the efficiency and accuracy of scheduling strategy implementation, a heuristic real-time power distribution strategy based on schedulable capacity is established to ensure the continuous and stable operation of active power distribution network.
[0083] Please refer to Figure 1 The application provides a kind of active power distribution network deep reinforcement learning real-time scheduling method, including the following steps:
[0084] S1, the equipment operation model of photovoltaic unit, combined heat and power unit, air source heat pump, electric gasification equipment and electric refrigeration equipment in active power distribution network is established, and a multi-objective optimization scheduling model and constraint condition considering economy and low carbon are constructed;
[0085] The equipment model in active power distribution network is:
[0086] 1) combined heat and power unit
[0087] The combined heat and power unit taking micro gas turbine as core equipment supplies electric energy and heat energy for park, realizes energy cascade utilization, and has good social and economic benefits. Its physical model is as follows:
[0088]
[0089]
[0090]
[0091] Among them, is the electric power output of CHP unit in t period, is the gas consumption rate of CHP unit in t period, is the power generation efficiency of CHP unit, is the low heat value of natural gas, is the heat power output of CHP unit in t period; is the loss coefficient, and are the upper and lower limits of electric power of CHP unit respectively.
[0092] 2) Air source heat pump
[0093] Air source heat pump uses air as heat source, and converts low-grade heat energy in outdoor air into high-grade heat energy under the drive of electricity. Its physical model is shown below:
[0094]
[0095]
[0096] wherein, is the load power of the heat pump in the t period, is the heat power of the heat pump in the t period, is the energy efficiency ratio of the heat pump, and are the upper and lower limits of the electric power of the heat pump, respectively.
[0097] 3) Electricity-to-gas equipment
[0098] Electricity-to-gas equipment produces hydrogen by electrolyzing water, and further reacts with carbon dioxide to produce methane, which can absorb a certain amount of carbon dioxide. The calculation formula is as follows:
[0099]
[0100]
[0101] wherein, and are the gas output and the consumed electric power of the electricity-to-gas equipment at t time, respectively; η P2G is the energy conversion efficiency; and are the upper and lower limits of the electric power consumed by the electricity-to-gas equipment, respectively.
[0102] 4) Electric refrigeration equipment
[0103] Ice storage air conditioning uses electricity to produce refrigeration through a compressor. In this paper, ice storage air conditioning is studied as a balancing unit of the cooling system, and the calculation formula is as follows:
[0104]
[0105]
[0106] wherein, and are the cooling output and the consumed electric power of the ice storage air conditioning at t time, respectively; η ISAC represents the energy conversion efficiency of ISAC; and are the upper and lower limits of the electric power consumed by the ice storage air conditioning.
[0107] 5) Photovoltaic plant
[0108] The photovoltaic plant output is predicted, the predicted value being the maximum value of the output, the actual output being between 0 and the predicted value:
[0109]
[0110] wherein, is the planned photovoltaic output; is the predicted photovoltaic output.
[0111] In order to realize the economic operation of the active power distribution network system and meet the policy background of low-carbon dispatching, a multi-objective optimization dispatching model considering economic and low-carbon is constructed.
[0112] minf = C purchase + C carbon
[0113]
[0114]
[0115] wherein, f is the total cost of the system, C purchase is the system operation cost, mainly including the purchase power cost from the upper grid and the purchase gas cost from the natural gas source, C carbon is the system carbon emission cost, including the carbon emission cost of fossil fuel combustion and the carbon emission cost of purchased power, T is the total operation time of the system, is the purchase power at time t, is the purchase power price from the upper grid at time t, is the purchase gas power at time t, is the purchase gas price from the natural gas source, is the consumption power of natural gas at time t, C coef is the carbon dioxide emission coefficient containing multiple coefficients such as low calorific value, unit calorific value carbon content and carbon oxidation rate, λ carbon is the carbon dioxide emission cost, is the carbon dioxide emission factor of the purchase power from the upper grid.
[0116] The constraint conditions are:
[0117]
[0118] S2, the multi-objective optimization scheduling model of the active power distribution network obtained in step S1 is expressed in the framework of reinforcement learning, and the improved deep deterministic policy gradient algorithm is used to train and schedule the multi-objective optimization scheduling model under the continuous state and action space, so as to obtain the benchmark result of the intraday scheduling for subsequent real-time scheduling;
[0119] The decision refers to selecting the DDPG algorithm based on the actor-critic framework for function approximation, so as to make it suitable for the low-carbon economic scheduling problem of the active power distribution network with continuous state and action space. The optimal policy function is estimated by the deep neural network, which can not only avoid the curse of dimensionality, but also save the information of the entire action domain.
[0120] (1) Action space design
[0121] In the design of the reinforcement learning algorithm based on deep learning, the research is on the short-term scheduling of 15min within a day, so in each period t, the action space of the active power distribution network is represented by the output of all devices.
[0122]
[0123] (2) State space design
[0124] The environmental state in the active power distribution network includes the electric load demand, the photovoltaic power generation, the park electricity purchase price, the park natural gas purchase price and the carbon dioxide emission price, and is represented as:
[0125]
[0126] (3) Reward function design
[0127] The objective function of the model is designed in the form of the immediate reward in the reinforcement learning, and part of the constraint conditions is added in the form of the penalty function. Since too many penalty functions will affect the convergence and stability of the algorithm, part of the constraint conditions of the model is considered when the state and action space are defined.
[0128] The reward obtained by the agent at period t is represented as:
[0129]
[0130] Among them, 1 / 1000 is the corresponding scaling of the cost term. The reward function scaling function is introduced here to reduce the influence of environmental uncertainty on the learning process and improve the convergence speed.
[0131] When the state s t of the active power distribution network is determined, the pros and cons of the low-carbon economic scheduling action a t of the system can be determined by the action-value function Qπ (s, a) to evaluate, i.e.
[0132]
[0133] where the discount factor γ∈[0, 1] represents the importance of the reward at future time to the current reward. For the low-carbon economic dispatch problem studied, the decision at the current period will have an important impact on the future, so γ should take a larger value.
[0134] The goal of low-carbon economic dispatch of active power distribution network is to find the optimal strategy π * to maximize the action-value function, i.e.
[0135]
[0136] (4) Algorithm mechanism improvement
[0137] 1) Mini-Batch Gradient Descent (MBGD)
[0138] In the learning process, due to the sequential interaction between the agent and the environment, there is a correlation between samples, which does not meet the assumption of independent and identically distributed samples of deep reinforcement learning. Therefore, the experience replay mechanism is used in the DDPG algorithm. By storing the training experience of the agent D = (s, a, r, s') at each period, a replay memory sequence is formed. During training, a small batch of experience samples is randomly extracted from D each time, and the network parameters are updated based on the gradient rule. The experience replay mechanism breaks the correlation between data by randomly sampling historical data, and the repeated use of experience also increases the data usage efficiency.
[0139] Compared with batch gradient descent, using mini-batch gradient descent to update parameters is faster and more conducive to robust convergence, avoiding convergence to local optimum; compared with stochastic gradient descent, using mini-batch gradient descent has higher computational efficiency, which can help quickly train the model.
[0140] 2) Uncertainty processing-data augmentation
[0141] Unlike the optimization-based scheduling algorithm, the training time of reinforcement learning is long, and if the trained model can only be used for specific scheduling scenarios, the model is not universal and cannot be applied to actual systems. In order to make the trained model be able to process multiple uncertainties and make scheduling decisions on the output of the device, the sources of uncertainty including the output of the electrical load and the output of the photovoltaic are considered. However, due to the imperfection of sensors and other devices or the short online running time of the system, part of the historical data in the system cannot be completely and accurately obtained. Therefore, Monte Carlo sampling is used to realize the acquisition of the historical data of the active distribution network, so as to be applied to the training of the reinforcement learning parameter model.
[0142] 3) Random exploration
[0143] In the algorithm design, by adding random noise v t to the action a t , the exploration ability of the DDPG algorithm in the interaction of the integrated energy system can be improved to learn more optimized dynamic scheduling strategies.
[0144] a t = π(s t | θ π ) + v t
[0145] In this paper, Ornstein-Uhlenbeck (OU) noise is used, which is a random variable based on the OU process and is commonly used to simulate time-dependent noise sets.
[0146] At the same time, in the later stage of algorithm parameter update, since the training parameters of the neural network have learned the optimal strategy, in order to ensure the convergence of the results, a decay factor is added to the noise to attenuate the randomness of exploration, so as to stabilize the optimal strategy and no longer continue to explore.
[0147] 4) Model convergence effect improvement
[0148] In the conventional DDPG algorithm, the learning rate of the actor network and the critic network is always kept unchanged, but in fact, this fixed learning rate parameter convergence effect is not the best. In the later stage of network parameter learning, due to the high learning rate, it is very likely to fall into a local optimal value. Therefore, referring to the method of random exploration, this paper also superimposes a decay factor on the learning rate of the network parameters, which is used for the continuous decay of the network parameter learning rate in the later stage. This method is more conducive to algorithm convergence.
[0149] At the same time, the ratio of exploration and learning has a great influence on the convergence result. In the early stage of the algorithm, due to insufficient exploration of the environment, high-frequency learning is likely to make the algorithm fall into local optimization. Therefore, by increasing the ratio of exploration and learning, the environment is fully explored before high-frequency learning is performed.
[0150] 5) DDPG solving algorithm
[0151] (1) Value network training
[0152] For the value network, the model parameters of the value network are trained by minimizing the loss function L(θ Q ):
[0153] L(θ Q ) = E(y t -Q(s t ,a t |θ Q ) 2 )
[0154] Where y t is the Q value of the target network, and E(·) is the expectation function.
[0155] y t =r t +γQ'(s t+1 ,π'(s t+1 θ π′ )θ Q′ )
[0156] At time t, the multi-energy equipment in the integrated energy system performs a scheduling action a t , and then enters the next state s t+1 , that is, the predicted load power of electricity, cold, heat, gas, photovoltaic power generation, and electricity and gas purchase and carbon price at the next time period.
[0157] The gradient of L(θ Q ) with respect to θ Q is:
[0158]
[0159] Where, is a function representing gradient calculation.
[0160] In the above formula, y t -Q(s t ,a t |θ Q ) is the timing differential error (TD-error). According to the gradient rule, the network is updated, and the update formula is:
[0161]
[0162] where μ Q is the learning rate of the value network.
[0163] (2) Strategy network training
[0164] For the strategy network, it provides gradient information as the direction of action improvement. In order to update the model parameters of the strategy network, the sampled policy gradient is used:
[0165]
[0166] According to the deterministic policy gradient update formula, the training parameters θ π of the strategy network are updated:
[0167] where μ π is the learning rate of the strategy network.
[0168] Finally, the training parameters θ Q′ and θ π′ of the target network are updated:
[0169] θ Q′ ← τθ Q +(1-τ)θ Q′ θ π′ ← τθ π +(1-τ)θ π′ where τ is the soft update coefficient, τ << 1.
[0170] (3) Online scheduling
[0171] The data obtained by random sampling is used as the environment state of the active power distribution network, and is used to train the parameters of the DDPG algorithm network offline. When the offline training process is completed, the DDPG algorithm parameters under the optimal strategy obtained by training will be fixed, and used to solve the low-carbon economic scheduling problem in actual operation of the park.
[0172] Specifically, when the scheduling task arrives, at each scheduling period t, the system will select the scheduling action a t according to the current running state s t using the trained strategy network and execute it, and then enter the next environment state s t+1 and obtain the reward r t at the same time; in this way, the dynamic scheduling action reference value for the whole day can be obtained for subsequent real-time power allocation.
[0173] S3, based on the reference results of the intraday scheduling obtained in step S2 and the idea of model predictive control, an event-driven real-time scheduling strategy is proposed, which comprehensively considers the imbalance degree of the supply and demand sides of the system and the schedulable capacity of the response subject;
[0174] The event-driven is based on:
[0175] The traditional periodic driving method divides the entire scheduling process into multiple operation periods, and schedules and controls the equipment before each scheduling period. In order to simplify the system operation mode and ensure that the system has certain emergency response capability when dynamic events occur, an event-driven real-time scheduling control strategy is established. At each scheduling time, the state variables in the system are monitored in real time. Once some uncertain disturbance events occur in the system, the system will develop a real-time power distribution plan according to the information collected at the moment, and issue control instructions to each device. If the system state changes little, no event trigger signal will be generated, and the device will execute the control instruction according to the result of the intraday scheduling.
[0176] Comprehensively considering the imbalance degree of the supply and demand sides of the system and the schedulable capacity of the response subject is:
[0177] The focus of real-time scheduling research is to transfer the power deviation compared with the 15-minute scheduling plan to the devices in the system for adjustment. Since there are multiple devices, it is necessary to reasonably allocate the output of each device to avoid excessive adjustment by a certain device. In order to allocate different power compensation tasks to different devices, a unified evaluation standard needs to be set to evaluate the adjustable capacity of each device under the consideration of various factors, and according to the adjustable capacity, the compensation power corresponding to the adjustable capacity is allocated. Therefore, the concept of schedulable capacity is proposed to quantitatively represent the adjustment capacity of different devices when they are connected to the system to participate in the ultra-short-term regulation and operation of the system. This evaluation index comprehensively analyzes the reliability factor, positive / negative adjustment capacity, operation and maintenance loss factor, and carbon emission factor of the device.
[0178] 1) Reliability factor: used to represent the completion of the device participating in system adjustment within a certain period of time.
[0179]
[0180] Where N is the total number of times the device participates in scheduling within a certain period of time; Pi(t) is the power value allocated to device i in real-time scheduling at time period t; Pi(t) is the actual output power value of device i at time period t. The higher the credit degree, the smaller the difference between the actual output power and the planned scheduling output power, which can be considered to have the smallest impact on the microgrid.
[0181] 2) Positive / negative adjustment capacity
[0182] From the above formula, the greater the real-time power of the device, the stronger its reverse regulation ability; the smaller the real-time power of the device, the stronger its forward regulation ability.
[0183] 3) Operation and maintenance loss factor
[0184] Wherein, is the unit operation and maintenance cost of the ith device. The higher the scheduling operation and maintenance cost of the device, the weaker its scheduling ability, and vice versa.
[0185] 4) Carbon emission factor
[0186] If the device is scheduled to produce carbon dioxide, the weaker its scheduling ability.
[0187] Since the attribute values of each index have different dimensions and value ranges, it is not convenient to analyze the numerical values uniformly, so it is necessary to normalize the attribute values of each index.
[0188]
[0189] The real-time scheduling strategy is specifically:
[0190] Real-time scheduling should give priority to meeting the reliable supply of electrical load in the system, and then fully utilize new energy output and comprehensively consider the economy of system operation. When the system is running in real time, the monitored load, photovoltaic and various device state information are read at each scheduling time, and it is judged whether there is a trigger signal. If not, the control instruction is issued according to the 15-minute scheduling reference value within the day, and if there is an event trigger signal, the real-time power distribution strategy is adopted, and the scheduling correction is executed according to the preset algorithm. First, the system evaluates the schedulable ability of all devices according to the preset index, and determines the scheduling priority of each device according to the evaluation result; further, combined with the power compensation demand of the system within each scheduling period, the power distribution criteria of the device are formulated, and the proposed power distribution criteria can ensure the reasonable distribution of device output and avoid that a device excessively undertakes the system regulation task.
[0191] S4, according to the above intra-day-real-time multi-time scale scheduling framework, realize the real-time scheduling of active distribution network based on deep reinforcement learning under multiple uncertainties.
[0192] The real-time scheduling based on deep reinforcement learning is specifically:
[0193] In the dynamic control stage, in order to effectively cope with the uncertain environment and improve the control precision of the system, a low-carbon economic dispatching model is established based on a model prediction control method, dynamic optimization with a scale of 15 minutes is carried out in an optimization time domain, and a power reference value is provided for a real-time power distribution stage with a shorter sampling period; in the real-time power distribution stage, the time scale is further shortened to 5 minutes in the dynamic optimization framework, and real-time power distribution operations are carried out on each response subject according to the above-mentioned comprehensive evaluation system, so that the real-time control of the optimal operation of the active power distribution network is realized.
[0194] In another embodiment of the present application, an active power distribution network deep reinforcement learning real-time scheduling system is provided, which can be used to realize the above-mentioned active power distribution network deep reinforcement learning real-time scheduling method.
[0195] Wherein,.
[0196] In order to make the purpose, technical scheme and advantages of the embodiments of the present application clearer, the technical scheme in the embodiments of the present application will be described clearly and completely below in combination with the drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, rather than all the embodiments. The components of the embodiments of the present application described and shown in the drawings herein can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present application provided in the drawings is not intended to limit the scope of the claimed present application, but only represents selected embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor fall within the scope of the present application.
[0197] Compared with the traditional real-time scheduling method, the improved DDPG algorithm adapts to the uncertain changes of photovoltaic and load, avoids modeling complex uncertainty, and the improved model generalization and convergence mechanism makes the algorithm more conducive to convergence. The real-time control strategy comprehensively considers the imbalance degree of the supply and demand sides of the system and the schedulable ability of the response subject, and the design of the event-driven mechanism enables the response subject to quickly and accurately respond to the energy compensation demand, effectively realizing the safe and stable operation of the system. The implemented case study shows that the proposed intraday-real-time multi-time scale active power distribution network low-carbon economic dispatching method can achieve a performance close to that of the optimization method with perfect prediction information under the condition of limited historical data information, and the calculation time is greatly shortened.
[0198] The method of the application is applied in an actual active power distribution network. In order to highlight the compatibility and scalability of the proposed method, a case study is designed. Three typical scenarios are designed respectively to verify the effectiveness of the proposed method. Scenario one is the real-time scheduling of the active power distribution network using the model predictive control method, scenario two is the real-time scheduling of the active power distribution network using the traditional deep reinforcement learning method, and scenario three is the proposed intraday-real-time multi-time scale active power distribution network real-time scheduling.
[0199] Please refer to Figure 2 , the reward function curve in the algorithm training process converges after 2000 rounds. It can be observed that, due to the unfamiliarity of the agent with the environment at the beginning, the reward value obtained after decision-making is small. With the deepening of the training process, the agent takes actions that are more beneficial to obtaining rewards based on the experience obtained from interacting with the environment. Finally, the reward value converges to the maximum value, indicating that the agent has learned the optimal scheduling strategy to minimize the total system operation cost. Due to the continuous change of load data and photovoltaic data and the random exploration of actions, the total reward of the scheduling period has a certain degree of oscillation, which is reasonable.
[0200] Please refer to Figure 3 , the intraday scheduling result diagram of the equipment in the active power distribution network during the test period. From the diagram, it can be seen that during the period from 8:00 to 14:00, the electricity price has reached the maximum value, so the natural gas production of the electric-gas equipment is very low. In order to reduce the electricity purchase cost, the system mainly purchases gas from the external gas network during this period to achieve system power balance. From the above results, it can be seen that the improved DDPG algorithm can ensure the economic operation of the system and reduce the carbon emission cost of the system.
[0201] Please refer to Figure 4 and Figure 5 , the real-time generation and load power scheduling results of the test period can be seen. It can be seen that the rolling scheduling based on deep reinforcement learning alone has the problem of unstable results due to the large scheduling time scale, and cannot respond to the fluctuations of new energy and load in the power system. The system real-time energy supply deviation is large. The multi-time scale scheduling strategy proposed in this paper can balance the economic efficiency and operation reliability of short-term and ultra-short-term time scales, suppress the random fluctuations of new energy and load, ensure the real-time source-load balance of the system, and meet the requirements of safe and stable operation of the active power distribution network.
[0202] In summary, the active power distribution network deep reinforcement learning real-time scheduling method and system, the improved DDPG algorithm is adaptive to the uncertainty change of photovoltaic and load, avoids modeling the complex uncertainty, and the improvement of the model generalization and convergence mechanism makes the algorithm more conducive to convergence. The real-time control strategy comprehensively considers the imbalance degree of the system supply and demand sides and the schedulable ability of the response subject, the design of the event-driven mechanism enables the response subject to quickly and accurately respond to the energy compensation demand, and effectively realizes the safe and stable operation of the system. Finally, the learning mode of model data fusion breaks the limitations of human experience and small sample problems of a single physical model or data-driven model, effectively improves the credibility, explainability and robustness of the model, and the event-driven real-time power distribution strategy can ensure the efficiency and accuracy of the implementation of the scheduling strategy, and ensures the continuous and stable operation of the active power distribution network.
[0203] Those skilled in the art will understand that embodiments of the application can be provided as methods, systems, or computer program products. Therefore, the application can take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware aspects. Moreover, the application can take the form of a computer program product implemented on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROMs, optical storage devices, etc.) containing computer-usable program code.
[0204] The application is described with reference to flowcharts and / or block diagrams according to the methods, devices (systems), and computer program products of the embodiments of the application. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and the combination of flows and / or blocks in the flowcharts and / or block diagrams can be implemented by computer program instructions. These computer program instructions can be provided to a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to produce a machine, so that the instructions executed by the computer or other programmable data processing devices produce a device that implements the functions specified in the flowcharts and / or block diagrams. Figure 1 one or more flows and / or blocks Figure 1 an apparatus that carries out the functions specified in one or more flows and / or blocks.
[0205] These computer program instructions can also be stored in a computer-readable memory that can guide the computer or other programmable data processing devices to work in a specific way, so that the instructions stored in the computer-readable memory produce a product including instruction devices, which implement the functions specified in the flowcharts and / or block diagrams. Figure 1 one or more flows and / or blocks Figure 1 an apparatus that carries out the functions specified in one or more flows and / or blocks.
[0206] These computer program instructions can also be loaded into a computer or other programmable data processing devices, so that a series of operational steps are performed on the computer or other programmable data processing devices to generate a computer implemented process, so that the instructions executed on the computer or other programmable data processing devices provide a process for implementing the flowchart Figure 1 one flowchart or multiple flowcharts and / or blocks Figure 1 one flowchart or multiple flowcharts and / or blocks
[0207] The above is only to illustrate the technical idea of the present application, and cannot limit the protection scope of the present application. Any modification made according to the technical idea of the present application on the basis of the technical scheme falls within the protection scope of the claims of the present application.
Claims
1. An active power distribution network deep reinforcement learning real-time scheduling method, characterized in that, Comprise the following steps: S1, the establishment of active power distribution network photovoltaic unit, combined heat and power unit, air source heat pump, electric gas conversion equipment and electric refrigeration equipment equipment operation model, build considering economic and low carbon multi-objective optimization scheduling model; S2, the multi-objective optimization scheduling model obtained in step S1 is expressed as a framework of reinforcement learning, and an improved deep deterministic policy gradient algorithm is used to train and online schedule the multi-objective optimization scheduling model under continuous state and action space, to obtain the benchmark result of intraday scheduling, and the improved deep deterministic policy gradient algorithm is specifically: In each time period , the output of all devices is used as the action space in the active power distribution network; the demand for electrical load, photovoltaic power generation, park electricity purchase price, park natural gas purchase price and carbon dioxide emission price are used as the environmental state in the active power distribution network, and the goal of low-carbon economic dispatch of the active power distribution network is to find the optimal strategy The maximum action value function is used to update the parameters using small batch gradient descent, and the data obtained by Monte Carlo sampling is used as the environmental state of the active power distribution network to train the parameters of the deep deterministic policy gradient algorithm network offline. When the offline training process is completed, the parameters of the deep deterministic policy gradient algorithm under the optimal strategy are obtained to solve the low-carbon economic dispatch problem of the active power distribution network in actual operation; S3, according to the intraday scheduling benchmark result obtained in step S2, determine the real-time scheduling strategy based on event-driven, which comprehensively considers the imbalance degree of system supply and demand sides and the adjustable capacity of response subject, and the real-time scheduling strategy is specifically: When the system is running in real time, read in the monitored load, photovoltaic and various equipment state information at each scheduling time, there is no event trigger signal at present, according to the 15min scheduling benchmark value, issue control instruction, there is event trigger signal at present, adopt real-time power distribution strategy, and execute scheduling correction according to preset algorithm; S4, realize the real-time scheduling of active power distribution network based on deep reinforcement learning under multiple uncertainties according to the real-time scheduling strategy determined in step S3. 2.The active power distribution grid deep reinforcement learning real-time scheduling method according to claim 1, characterized in that, In step S1, the equipment operation model comprises: Combined heat and power unit operation model: wherein, PCHP is the electrical power output of the CHP unit, PCHP is the electrical power output of the CHP unit, PCHP is the electrical power output of the CHP unit, PCHP is the electrical power output of the CHP unit, PCHP is the electrical power output of the CHP unit, PCHP is the electrical power output of the CHP unit, PCHP is the electrical power output of the CHP unit, PCHP is the electrical power output of the CHP unit; PCHP is the electrical power output of the CHP unit, PCHP is the electrical power output of the CHP unit, PCHP is the electrical power output of the CHP unit, Air source heat pump operation model: wherein, is the heat pump's first period load power, is the heat pump's first period heat power, is the heat pump's coefficient of performance, and are the heat pump's electric power upper and lower limits, respectively. Electric gas conversion equipment operation model: wherein, and respectively are the gas output and the consumed electric power of the electric-gas conversion plant at the instant of time; is the energy conversion efficiency; is the energy conversion efficiency; and respectively are the upper and lower limits of the consumed electric power of the electric-gas conversion plant; Electric refrigeration equipment operation model: in, and Ice storage air conditioners The constant output power and power consumption; This indicates the energy conversion efficiency of ISAC; and The upper and lower limits of the power consumption of ice storage air conditioners; Photovoltaic unit operation model: wherein, the planned power output of the photovoltaic; the predicted power output of the photovoltaic. 3.The active power distribution grid deep reinforcement learning real-time scheduling method according to claim 1, characterized in that, In step S1, the multi-objective optimization scheduling model considering economic and low carbon is: wherein, is the total cost of the system, is the operation cost of the system, mainly including the cost of purchasing electricity from the upper grid and the cost of purchasing gas from the natural gas source, is the carbon emission cost of the system, including the carbon emission cost of fossil fuel combustion and the carbon emission cost of purchased electricity, is the total operation time of the system, is the purchasing power of electricity at the moment, is the purchasing unit price of electricity from the upper grid at the moment, is the purchasing power of gas at the moment, is the purchasing unit price of gas from the natural gas source, is the consumption power of natural gas at the moment, is the carbon dioxide emission coefficient including multiple coefficients such as low calorific value, unit calorific value carbon content and carbon oxidation rate, is the carbon dioxide emission cost, is the carbon dioxide emission factor of purchasing electricity from the upper grid.
4. The active power distribution grid deep reinforcement learning real-time scheduling method according to claim 3, characterized in that, The constraint condition is: wherein, is the electricity purchase power at the moment, is the planned output of the photovoltaic, is the electricity output of the CHP unit in the time period, is the electricity consumption of the ice storage air conditioner in the time period, is the heat pump in the time period, is the electricity consumption of the electricity-to-gas device in the time period. 5.The active power distribution grid deep reinforcement learning real-time scheduling method according to claim 1, wherein, In step S3, based on event-driven, specifically: At each scheduling time, the state variables in the active power distribution network are monitored in real time, when the uncertain disturbance event occurs, the active power distribution network formulates real-time power distribution plan according to the current collected information, and issues control instruction to each equipment; If the deviation between the intraday benchmark scheduling plan of the active power distribution network and the real-time equipment operation plan is less than the set deviation penalty value, no event trigger signal is generated, and the equipment executes the control instruction according to the intraday scheduling result. 6.The active power distribution grid deep reinforcement learning real-time scheduling method according to claim 1, wherein, In step S3, the comprehensive consideration of the imbalance degree of system supply and demand sides and the adjustable capacity of response subject is specifically: A unified evaluation standard is set to evaluate the adjustable capacity of each equipment, and the corresponding compensation power is allocated according to the adjustable capacity. The adjustable capacity quantitatively represents the adjustment capacity of different equipment when participating in the ultra-short-term regulation and operation of the system, including the reliability factor, positive / negative adjustment capacity, operation and maintenance loss factor and carbon emission factor indexes of the equipment. 7.The active power distribution grid deep reinforcement learning real-time scheduling method according to claim 1, wherein, In step S4, the real-time scheduling based on deep reinforcement learning is specifically: In the dynamic control stage, the low-carbon economic scheduling model is established, and the dynamic optimization with a scale of 15min is carried out in the optimization time domain; In the real-time power distribution stage, the time scale is shortened to 5min in the dynamic optimization framework, and the real-time power distribution operation is carried out on each response subject according to the comprehensive evaluation system, so as to realize the real-time control of the optimal operation of the active power distribution network.
8. An active power distribution grid deep reinforcement learning real-time scheduling system, characterized in that, Comprise: The construction module establishes a device operation model of photovoltaic units, combined heat and power units, air source heat pumps, electric-to-gas equipment and electric refrigeration equipment in the active power distribution network, and constructs a multi-objective optimization scheduling model considering economy and low carbon; The improvement module expresses the multi-objective optimization scheduling model obtained by the construction module into a framework of reinforcement learning, trains and schedules the multi-objective optimization scheduling model under continuous state and action space by using an improved deep deterministic policy gradient algorithm, and obtains a benchmark result of intraday scheduling, and the improved deep deterministic policy gradient algorithm is specifically as follows: In each time period , the output of all devices is used as the action space in the active power distribution network; the demand for electrical load, photovoltaic power generation, park electricity purchase price, park natural gas purchase price and carbon dioxide emission price are used as the environmental state in the active power distribution network, and the goal of low-carbon economic dispatch of the active power distribution network is to find the optimal strategy The maximum action value function is used to update the parameters using small batch gradient descent, and the data obtained by Monte Carlo sampling is used as the environmental state of the active power distribution network to train the parameters of the deep deterministic policy gradient algorithm network offline. When the offline training process is completed, the parameters of the deep deterministic policy gradient algorithm under the optimal strategy are obtained to solve the low-carbon economic dispatch problem of the active power distribution network in actual operation. The strategy module determines an event-driven real-time scheduling strategy considering the imbalance degree of the system supply and demand sides and the adjustable capacity of the response subject according to the benchmark result of intraday scheduling obtained by the improvement module, and the real-time scheduling strategy is specifically as follows: During real-time operation of the system, at each scheduling time, the monitored load, photovoltaic and state information of various devices are read in, no event trigger signal is currently triggered, a control instruction is issued according to the 15-minute scheduling benchmark value, there is an event trigger signal currently, a real-time power distribution strategy is adopted, and scheduling correction is performed according to a preset algorithm; The scheduling module realizes real-time scheduling of the active power distribution network based on deep reinforcement learning under multiple uncertainties according to the real-time scheduling strategy determined by the strategy module.
Citation Information
Patent Citations
Dispatching optimization method for distributed energy participating in peak regulation of distribution network based on reinforcement learning
CN110365057A
Multi-park energy scheduling method and system based on deep reinforcement learning
CN114091879A