Automatic control method and system for distribution network switches based on reinforcement learning theory
Through the reinforcement learning-based distribution network switch automatic control method, the Distflow power flow optimization and approximate dynamic programming algorithm are used to optimize the distribution network switching strategy, solve the fluctuation and failure problems of distributed power generation units, and improve the power supply reliability and economy.
Patent Information
- Application Number
- CN202310099843.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-02-06
- Publication Date
- 2025-10-03
- Estimated Expiration
- 2043-02-06
AI Technical Summary
The traditional distribution network has a weak structure, low degree of automation, and distributed power generation units have frequent fluctuations and failures, which affect the reliability and economy of power supply.
Based on reinforcement learning theory, a Distflow power flow optimization constraint model and an approximate dynamic programming algorithm are established. Through the Markov decision process MDP model, the distribution network switch automatic control strategy is optimized in real time, taking into account the output fluctuations of distributed generation units and the network topology structure, providing efficient and stable switch control.
It improves the power supply reliability and economy of the distribution system, optimizes the working status of distributed power generation units, solves fault and fluctuation problems, and improves investment benefits.
Smart Images

Figure CN116260143B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of power grid dispatching, and in particular to a distribution network switch automatic control method and system based on reinforcement learning theory. Background Art
[0002] The distribution network undertakes the arduous task of receiving and distributing electricity in the power system. It directly faces the electricity users and is closely connected with the daily production and life of the general public. The traditional distribution network has a "tree" structure. This "radiating" structure has many weaknesses: serious problems such as large damage caused by faults, poor mutual supply capacity, and low degree of automation. Although the manufacturing cost has been greatly reduced, its reliability is not high. In recent years, more and more distributed generation systems (DG) such as wind power generation and photovoltaic power generation have been connected to the distribution network. Such optimization strategies can, to a certain extent, meet the development requirements of the clean, environmentally friendly, low-cost, efficient and reliable power industry.
[0003] However, the integration of these renewable energy sources has also had numerous adverse impacts on distribution networks. The most significant issue is that renewable energy generation modes like wind and solar power are subject to random fluctuations, uncertainty, and instability, influenced by environmental factors. Furthermore, due to natural and man-made disasters, nodes and branches in the distribution network are prone to failure. These unfavorable factors complicate and variate the operation of the distribution system, directly impacting its safety. Summary of the Invention
[0004] The present invention aims to address, at least to a certain extent, one of the technical problems in the related art. To this end, a first object of the present invention is to provide a method for automatically controlling distribution network switches based on reinforcement learning theory. This method can address operational fluctuations, failures, and malfunctions of distributed generation units in the distribution network, thereby improving the power supply reliability and investment returns of the distribution system.
[0005] The second object of the present invention is to provide a distribution network switch automatic control system based on reinforcement learning theory.
[0006] To achieve the above object, the present invention is implemented through the following technical solutions:
[0007] A distribution network switch automatic control method based on reinforcement learning theory, comprising:
[0008] Step S1: Establish a Distflow power flow optimization constraint model and determine a reinforcement learning algorithm for approximate dynamic programming;
[0009] Step S2: Obtain historical output data of photovoltaic and wind power distributed generation units, and establish an output power fluctuation and conversion model of the distributed generation units based on the historical output data;
[0010] Step S3: Determine the topology of the distribution network, obtain active power output data of distributed generation units, status information of controllable section switches and tie switches of the distribution network, and calculation results of the Distflow power flow optimization constraint model, and establish a Markov decision process (MDP) model of the distribution network based on the output power fluctuation and conversion model of the distributed generation units, the distribution network topology, and the information obtained in step S3;
[0011] Step S4: using the reinforcement learning algorithm of approximate dynamic programming to solve the Markov decision process MDP model of the distribution network, and outputting the optimal strategy for automatic control of distribution network switches in real time.
[0012] Optionally, step S1 includes:
[0013] Step S11: establishing the Distflow power flow optimization constraint model according to the distribution network topology constraints and the basic theory of distribution network power flow calculation;
[0014] Step S12: Determine the reinforcement learning algorithm of approximate dynamic programming according to reinforcement learning theory and Bellman optimal equation.
[0015] Optionally, in step S2, before establishing the output power fluctuation and conversion model of the distributed power generation unit, the method also includes: analyzing and quantifying the output power fluctuation and uncertainty conditions of each distributed power generation unit, so as to use the quantified value as the input of the Markov decision process MDP model of the distribution network.
[0016] Optionally, the output power fluctuation and conversion model of the distributed power generation unit can simulate the fluctuation and uncertainty of the output power of each distributed power generation unit.
[0017] Optionally, step S3 includes:
[0018] Step S31: Modeling the output power fluctuation status quantization value of each distributed power generation unit and the corresponding distribution network topology as state parameters of the distribution network Markov decision process MDP model;
[0019] Step S32: Modeling the action states of the switches connected to the branches of the power distribution network as action combination parameters of the Markov decision process (MDP) model of the power distribution network;
[0020] Step S33: Selecting the load shedding and line operation cost modeling in the Distflow power flow optimization constraint model considering distributed generation failure as the reward function reference index of the distribution network Markov decision process MDP model;
[0021] Step S34: defining the state transition probability of the Markov decision process (MDP) model of the power distribution network, so as to take into account the uncertainty caused by the probability of change of the power output level of each distributed generation unit.
[0022] Optionally, in step S4, before outputting the optimal strategy for automatic control of distribution network switches in real time, the method further includes: performing offline iterative training on the Markov decision process MDP model of the distribution network.
[0023] Optionally, the step of performing offline iterative training on the distribution network Markov decision process MDP model includes: inputting active power output data of distributed generation units, outputting automatic control strategies for distribution network switches, and feeding back evaluation indicators of real-time operation of the distribution network.
[0024] Optionally, in step S4, after outputting the optimal strategy for automatic control of the distribution network switches in real time, the method further comprises: displaying the decision result in text and the real-time distribution network topology on a human-computer interaction interface.
[0025] Optionally, the action state of the switch includes: the switch changing the switch state at the current moment and maintaining the switch state at the current moment.
[0026] To achieve the above-mentioned object, the second aspect of the present invention provides a distribution network switch automatic control system based on reinforcement learning theory, comprising:
[0027] Establish a module for establishing the Distflow power flow optimization constraint model;
[0028] A determination module for determining a reinforcement learning algorithm that approximates dynamic programming;
[0029] An acquisition module is used to acquire historical output data of photovoltaic and wind power distributed generation units, so that the establishment module can establish an output power fluctuation and conversion model of the distributed generation units based on the historical output data;
[0030] The establishment module is further configured to establish a Markov decision process (MDP) model of the distribution network based on the distribution network topology determined by the determination module, the active power output data of the distributed generation units acquired by the acquisition module, the controllable section switches and tie switch status information of the distribution network, and the calculation result information of the Distflow power flow optimization constraint model, as well as the output power fluctuation and conversion model of the distributed generation units;
[0031] The calculation module is used to solve the Markov decision process MDP model of the distribution network according to the reinforcement learning algorithm of approximate dynamic programming, and output the optimal strategy for automatic control of distribution network switches in real time.
[0032] The present invention has at least the following technical effects:
[0033] 1. The present invention establishes a distributed power generation unit output power fluctuation and conversion model through the historical output data of distributed power generation units such as photovoltaic and wind power, and obtains a Markov decision process (MDP) model through the distributed power generation unit output power fluctuation and conversion model. Since the distributed power generation unit output power fluctuation and conversion model fully considers the working efficiency fluctuation of new energy power generation modes such as wind power generation and solar power generation due to environmental influences, the Markov decision process (MDP) model obtained by this model can focus on the power output fluctuation of each distributed power generation unit and provide an efficient and stable distribution network switch automatic control optimization strategy.
[0034] 2. The present invention establishes a Distflow power flow optimization constraint model, which can solve the problems of distribution network optimization reconstruction and fault reconstruction. Since the distribution network reconstruction can achieve the purpose of optimizing power flow distribution and improving power supply reliability and economy by selecting the power supply path of the user, the Markov decision process MDP model established by the Distflow power flow optimization constraint model can solve the distribution system fault problem and improve the reliability of the distribution system. Based on the MDP model, the optimal strategy within each time interval can be dynamically formulated according to the real-time state and state transition probability.
[0035] 3. The present invention adopts the approximate dynamic programming (ADP) algorithm to solve the above-mentioned MDP model, which can solve the "dimensionality curse" problem.
[0036] 4. The present invention aims to optimize and solve the problems of operating fluctuations, failures and malfunctions of distributed power generation units in the distribution network, improve the power supply reliability of the distribution system, and enhance the investment benefits of the distribution system.
[0037] Additional aspects and advantages of the present invention will be set forth in part in the description which follows and, in part, will be obvious from the description which follows, or may be learned through practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] Figure 1 This is a flow chart of a distribution network switch automatic control method based on reinforcement learning theory according to an embodiment of the present invention.
[0039] Figure 2 This is an overall framework diagram of the system corresponding to the distribution network switch automatic control method based on approximate dynamic programming according to an embodiment of the present invention.
[0040] Figure 3 This is a flow chart of an approximate dynamic programming algorithm according to an embodiment of the present invention.
[0041] Figure 4 This is a flowchart of an offline calculation process according to an embodiment of the present invention.
[0042] Figure 5 This is a flowchart of the online calculation process of an embodiment of the present invention.
[0043] Figure 6 This is an overall flow chart of the distribution network switch automatic control method based on reinforcement learning theory in an embodiment of the present invention.
[0044] Figure 7 This is a structural block diagram of a distribution network switch automatic control system based on reinforcement learning theory in an embodiment of the present invention. DETAILED DESCRIPTION
[0045] The present embodiment is described in detail below. Examples of the embodiment are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to explain the present invention, but are not to be construed as limiting the present invention.
[0046] The following describes the distribution network switch automatic control method and system based on reinforcement learning theory of this embodiment with reference to the accompanying drawings.
[0047] Figure 1 This is a flow chart of a distribution network switch automatic control method based on reinforcement learning theory provided by one embodiment of the present invention. Figure 1 As shown, the method includes:
[0048] Step S1: Establish a Distflow power flow optimization constraint model and determine a reinforcement learning algorithm for approximate dynamic programming.
[0049] The step S1 comprises:
[0050] Step S11: establishing a Distflow power flow optimization constraint model based on the distribution network topology constraints and the basic theory of distribution network power flow calculation.
[0051] Specifically, the network node structure of a distribution network is primarily tree-like, radial-like, and supplemented by ring-like structures. This is widely used in light- and medium-density load areas. For both economic and management reasons, the radial-like distribution network structure is widely adopted in power grids.
[0052] Assuming that a distribution network has n nodes and m connecting lines, the basic judgment equation of the distribution network "tree" structure is:
[0053] m=n-1 (1)
[0054] The above equation describes the basic requirements of a "tree" structure, but a spanning tree also needs to meet connectivity. A "tree" structure with connectivity is called a spanning tree. In a spanning tree, all nodes except the root node (substation node) must have only one parent node. This requirement can be achieved using the following equation:
[0055] β ij +βj i =α l , l=1,2,...,m (2)
[0056] ∑ j∈N(i) β ij =1,i=1,2,...,n (3)
[0057] β 0j =0, j∈N(0) (4)
[0058] β ij ∈{0,1} (5)
[0059] 0≤α l ≤1 (6)
[0060] By introducing two binary variables β ij and β ji To correspond to each connection line in the "tree" structure, and β ij =1 means node j is the parent node of node i, otherwise it is β ij =0. ∑ is a summation function. N(i) represents the set of all nodes connected to node i. In addition, the connection state (connected or disconnected) of any two nodes in the network is determined by the variable α. l or α ij This ensures that the distribution network corresponds to the spanning tree connecting the main substation. The above equations respectively indicate that line l actually exists in the spanning tree; that each node except the root node has exactly one parent node; and that the substation node (the root node) has no parent node. These five equations guarantee the connectivity of the "tree" structure, making it a spanning tree that simulates the structure of the distribution network.
[0061] Automatic control methods for distribution network switches must satisfy the operational constraints of the distribution network within each decision time t, including network topology constraints, power balance constraints, power flow constraints, voltage limit constraints, and line capacity constraints. The theoretical basis for satisfying these constraints is the DistFlow distribution network power flow calculation model. Power flow calculations in power systems utilize parameters such as the physical structure of each node in the network, voltage phasors, active and reactive power distribution, and line losses as operating conditions to determine the operating status of the entire power system.
[0062] According to Kirchhoff's voltage law, the law of conservation of energy and Ohm's law, we can get the following formula:
[0063] S1=S0-S loss1 -S L1 (7)
[0064]
[0065] V1∠θ=V0-z1I0 (9)
[0066] Among them, S i , i = 1, 2, ..., n represents the injection power of node i. For example, the injection power of node 0 is S0 = P0 + jQ0, that is, the injection power S0 is equal to the complex sum of the injected active power P0 and the reactive power Q0. loss1 represents the energy loss on the line from node 0 to node 1, S L1 = represents the load on node 1. The injected power on node 1 is equal to the injected power on node 0 minus the line loss power and the load demand power. z1 = r1 + jx1 represents the line impedance value connecting node 0 to node 1, r1 and x1 are line impedance variables. According to Ohm's law, the relationship between impedance and energy loss power can be expressed as V1∠θ represents the voltage phasor at node 1, and its relationship with the voltage phasor at node 0 can be expressed as V1∠θ=V0-z1I0, where I0 is the line current and θ is the phase angle.
[0067] According to the power calculation formula, the following formula is obtained:
[0068]
[0069] in, is the conjugate complex number of S0, then z1=r1+jx1, Substituting the above formula into V1∠θ=V0-z1I0, we can get the following formula:
[0070] V1∠θ=V0-(r1+jx1)(P0-jQ0) / V0 (11)
[0071] Taking the modulus of the phasor on both sides of the above equation, we get the following equation:
[0072]
[0073] After simplification, we can get the following formula:
[0074]
[0075] Based on the above analysis and using the above formula, we can extrapolate to i to get the recursive equation, and then get the following DistFlow equation:
[0076]
[0077]
[0078]
[0079] In order to conform to the actual situation of power operation, the above formula introduces the lowercase letters is the total power load consumption of node j. j and q j are the injected active power and injected reactive power of node j respectively, and They represent the active power and reactive power of the electric power load of node j respectively; and They represent the active power and reactive power of the generator or DG output power supply connected to node j, v j is the voltage at node j, r j and x j is the line impedance variable. It should be noted that this only exists when there is a generator or DG power generation unit at node j. and These two physical quantities.
[0080] Now we have the nonlinear Distflow equation. To better apply the above formula to practical research work, we propose the following two assumptions:
[0081] Assumption 1: The value of the nonlinear term in the power distribution model is very small and can be considered as zero.
[0082] Assumption 2: It is believed that V j ≈V0, then we can get the following formula:
[0083]
[0084] Based on the above two assumptions, the nonlinear Distflow equation obtained above can be transformed into a linear system of equations:
[0085]
[0086]
[0087]
[0088] Thus, a linearized Distflow power flow calculation model has been obtained. Its automatic control of the switches in each branch of the distribution network is essentially the planning and design of the distribution network topology. The above analysis and research are all based on the linearized Distflow power flow calculation model.
[0089] Based on the Distflow power flow calculation model, the following power flow optimization constraint model is considered:
[0090] min∑p sd,j (twenty one)
[0091]
[0092]
[0093]
[0094]
[0095]
[0096]
[0097]
[0098]
[0099] Among them, the new physical quantity p is introduced sd,i and q sd,i Represents the active load shedding and reactive load shedding at node i respectively. Load shedding is also called load reduction, i.e., reducing the load. Usually, when encountering line faults or natural disasters, in order to maintain the power balance and stability of the power system, the behavior of disconnecting part of the load from the power grid is called load shedding. In addition, P ik and Q ik denote the active power and reactive power flowing from node i to any node k, and They represent the active power and reactive power of the generator or DG output power supply of node i, P d,i and q d,i Represent the active and reactive load demands at node i, represents any node, and Respectively represent P ji and Q ji The maximum value of and Represents V i Minimum and maximum values; use the power factor of the load tanβ to express p sd,i and q sd,i Using the constraint relationship of the above power flow optimization constraint model, for a distribution network with n nodes, assuming that the voltage V1 of the root node, node 1, and the load demand P of each node are known, d,i , distribution network topology and each branch impedance r ji +x ji In the case of each node voltage V i 、V j , active power P injected into the load ji and Q ji As well as active load shedding and reactive load shedding p sd,i and q sd,i Optimize and solve physical quantities.
[0100] The on / off status of some lines in the distribution network can be controlled, so the following Distflow power flow optimization constraint model can be obtained:
[0101] min∑p sd,j (30)
[0102]
[0103]
[0104]
[0105]
[0106]
[0107]
[0108]
[0109]
[0110] The variable μ is introduced in the above formula ji Indicates the on / off status of the line. If μ ji =1 means the circuit is closed, otherwise μ ji =0.
[0111] Network reconfiguration is the essence of automatic switch control strategies. It involves changing the combined state of section switches and tie switches during normal network operation, thereby selecting the power supply path for each user. This optimizes power flow distribution and improves power supply reliability and economic efficiency. This model can be used to address issues such as distribution network optimization and fault reconfiguration.
[0112] Step S12: Determine a reinforcement learning algorithm for approximate dynamic programming based on reinforcement learning theory and Bellman's optimal equation.
[0113] Reinforcement learning is a class of problems within the field of machine learning that aims to enable intelligent agents to take optimal actions to maximize their rewards. It is used to find the best action to take in a given environment. Reinforcement learning differs from supervised learning in that, in supervised learning, the training data is labeled, so the model itself is trained using the correct answers. In reinforcement learning, while there are no labels, the agent gradually learns the optimal action or path through repeated experience and mistakes.
[0114] Dynamic programming (DP) originated in the fields of engineering and finance, where it tends to focus on continuous state and decision-making control problems. In contrast, in the field of artificial intelligence, DP primarily deals with discrete states and decisions. It is a model-based algorithm in reinforcement learning. In this reinforcement learning algorithm, the environment and model are known to the agent. Given a complete Markov decision process model, it can learn the optimal policy, thus being classified as a model-based method.
[0115] There are many high-dimensional problems in DP, such as the one studied in this paper, that are typically studied using tools from programming. Much of this work focuses on deterministic problems using tools from linear, nonlinear, or integer programming.
[0116] Like other reinforcement learning algorithms, the core idea of DP is still to find the optimal decision based on the state-value function. However, the state-value function in the DP algorithm is based on the Bellman optimality equation, based on which we can obtain:
[0117]
[0118] Among them, the v with an asterisk subscript * (S t ) represents the state-value function that satisfies value maximization, R t is the immediate reward at time t, γ is the discount factor, γ∈[0,1], the degree to which γ is close to 1 indicates the importance of future benefits; conversely, the degree to which γ is close to 0 indicates the importance of current benefits, G t+1is the sum of all rewards up to time t+1, S t is the state set, s is the state at time t, A t is the action set, a is the action at time t, To find the expected value, we can use the concept of expectation in probability theory to rewrite it:
[0119] v * (S t )=max a ∑ s′,r p(s′,r|s,a)[r+γG t+1 ] (40)
[0120] Among them, max is the maximum value function, s′ is the state at the next moment, r is the reward value of the environment feedback, and p(s′, r|s, a) is the probability of the environment feedback reward r and the transition to the next state s′ under the premise of state s and action a.
[0121] It should be noted that G t+1 With v * (S t+1 ) are interchangeable, and we get:
[0122] v * (S t )=max a ∑ s′,r p(s′,r|s,a)[r+γv * (S t+1 )] (41)
[0123] v * (S t )=max a {R t +∑ s′,r p(s′,r|s,a)*[γv * (S t+1 )]} (42)
[0124] The equation after substitution transforms the Bellman equation into a recursive update equation that approximates the ideal value function, which is the prototype of the DP algorithm. Based on the above formula, all dynamic programs can be written in a recursive way. The recursive process converts the state value v at a certain time t into a specific state value v. t (S t ) and the state value v entered at the next moment t′ t+1 (S t+1 ) are connected.
[0125] However, the DP algorithm has three "curses" in terms of dimension, which makes the solution process complicated, mainly manifested in the following three aspects: the state space S is too large, and the value function v cannot be calculated within an acceptable time. * (S t ); the decision space A is too large, and there are too many permutations and combinations of actions, which often increase exponentially, making it impossible to quickly find the optimal action; the result space is too large, and it is impossible to calculate the expected value of future rewards.
[0126] Approximate dynamic programming is based on an algorithmic strategy that progresses gradually over time, so it is also called forward dynamic programming. The DP algorithm needs to solve After converting into the form of probability distribution, we get ∑ s′,r p(s′,r|s,a)[r+γv * (S t+1 )]. By observing the formula, we find that too many states in the MDP (Markov Decision Process) will lead to slow or unsolvable problems. We propose an approximate dynamic programming algorithm to solve this problem, which converts the state-value function v * (S t ) and introduces a new concept:
[0127] The post-decision state refers to the state of the system after a decision is made and before any new information arrives. t After S t and S t+1 The state between
[0128] After proposing this concept, the above formula can be rewritten as:
[0129]
[0130] Instead, we solve for the expected value of future rewards.
[0131] Step S2: Obtain historical output data of photovoltaic and wind power distributed generation units, and establish an output power fluctuation and conversion model of the distributed generation units based on the historical output data.
[0132] In step S2, before establishing the output power fluctuation and conversion model of the distributed power generation units, the method further includes: analyzing and quantifying the output power fluctuation and uncertainty of each distributed power generation unit, so as to use the quantified value as one of the state inputs of the Markov decision process (MDP) model of the distribution network.
[0133] According to relevant historical data, the active power output of wind farms and photovoltaic farms on a long-term scale is highly random, which meets the requirements of exploring the volatility of new energy power generation units such as wind power generation and photovoltaic power generation with the natural environment.
[0134] The volatility and randomness of wind farm active power output are affected by a variety of objective factors, such as the geographical and climatic conditions of the wind farm area and the spatial distribution and arrangement of wind turbines. Using wind power generation active power output data with a sampling period of 15 minutes, the average daily wind power output is calculated as follows:
[0135]
[0136] Among them, P 日均出力 and W 日 Where represents the average daily wind power output and daily power generation, respectively, and P(t) represents the active power output at time t. The above formula is used to analyze the wind farm's average daily output. The active power output of a wind farm is affected by seasonal weather conditions and fluctuates significantly over the course of a year. Therefore, to qualitatively identify wind farm DG active power output fluctuation data, we selected output data from a typical wind farm output day (daily power generation = annual power generation / 365 days). Peak power generation on a typical output day occurs from 3:00 PM to 10:00 AM and from 10:00 PM to 10:00 AM, with power generation occurring approximately non-existent during the remaining periods.
[0137] The volatility and randomness of PV power generation's active output are primarily caused by the plant's sunshine hours, altitude, and natural disasters (drought, rainstorms, frost). Similarly, using actual PV power generation active output data from a typical output day with a sampling period of T = 15 minutes, weather conditions directly lead to significant fluctuations in PV power station output; output in clear weather is significantly better than in rainy weather. PV DG output is intermittent, with extended periods of power outages occurring at night.
[0138] In this embodiment, after the distributed power generation unit output power fluctuation and conversion model is established, the distributed power generation unit output power fluctuation and conversion model can be used to simulate the volatility and uncertainty of the output power of each distributed power generation unit.
[0139] Specifically, the quantified value of the output power fluctuation of each distributed generation unit (DG) and the corresponding distribution network topology are regarded as the state of the MDP, while the power output power of the DG has high uncertainty.
[0140] DG has k output levels in each time period. The larger the k value, the finer the discrete quantization of the DG output power analog quantity and the smaller the error. The probability of an output level k at time t to another output level k′ at time t+1 is expressed as ∏kk′ This transition probability can be represented by Monte Carlo simulation of the historical data of wind power generation DG. For example, the number of historical occurrences of output level k at time t is m times, and the number of transitions from output level k to output level k′ at time t+1 is n times, so ∏ kk′ Equal to n / m.
[0141] Step S3: Determine the topology of the distribution network, obtain the active output data of the distributed generation units, the status information of the controllable section switches and tie switches of the distribution network, and the calculation results of the Distflow power flow optimization constraint model, and establish a distribution network Markov decision process MDP model based on the output power fluctuation and conversion model of the distributed generation units, the distribution network topology and the information obtained in step S3.
[0142] Among them, the overall system framework diagram corresponding to the distribution network switch automatic control method based on approximate dynamic programming is as follows: Figure 2 shown. Figure 2 It consists of two parts: reinforcement learning algorithm and distributed distribution network. At each moment of the operation of the distributed distribution network, external conditions such as weather, natural disasters and network topology are used as state parameters S t That is, the external environmental state is input to the MDP process, which is specifically manifested in the external conditions such as weather and natural disasters affecting the output of photovoltaic power generation and wind power generation, and then the switching strategy A is output. t According to the pre-set parameters, such as network line operation cost, load shedding, etc., the current reward value is fed back, that is, the immediate reward R t The reinforcement learning (ADP) algorithm utilizes post-decision states and a forward dynamics algorithm. It iterative calculations are performed based on the above data, the real-time operating status of the distribution network, and accumulated rewards. The algorithm's results are fed back to the agent. Each complete time cycle is considered a training session. After hundreds of training sessions, each decision will produce a converged calculated value. These values are stored in a table and compared to find the optimal strategy at each moment.
[0143] The step S3 of establishing the Markov decision process MDP model of the power distribution network includes:
[0144] Step S31: Modeling the output power fluctuation status quantized value of each distributed power generation unit and the corresponding distribution network topology as state parameters of the distribution network Markov decision process MDP model.
[0145] State set S i,t :The switches on each line in the distribution network topology are defined in turn as the network topology:
[0146] Ξ t=[swt1,swt2,swt3,...,swt n ] (45)
[0147] Among them, swt n Indicates the current state of switch n. Each switch has two states: open and closed, represented by binary numbers 0 and 1 respectively. Therefore, the current distribution network topology Ξ can be represented by an n-bit binary number. t In the Markov decision model of the distribution network, the network topology and the power output level of the distributed generation units at time t are defined as the state set:
[0148] S i,t =[Ξ t |k 1,t , k 2,t , k 3,t ,...,k dg,t ,...,k DG,t ] (46)
[0149] Among them, t represents the topology of the distribution network at time t, k dg,t Indicates the DG power output level at time t.
[0150] Step S32: Modeling the action states of the switches connecting the branches of the power distribution network as action combination parameters of the Markov decision process (MDP) model of the power distribution network.
[0151] The action state of the switch includes the switch changing the switch state at the current moment and maintaining the switch state at the current moment.
[0152] The set of actions in state A t (S i,t ): The switches connecting each node [swt1, swt2, swt3, ..., swt n At each moment t, there are only two actions, namely, changing the switch state at moment t-1 (converting from closed state to open state, from open state to closed state) or maintaining the switch state at moment t-1, using the action a in a single switch state swt1 = a1 or a2. Each switch in the tree network has the above two actions at any time. t (S i,t ) can be expressed as:
[0153] A t (S i,t )=[a swt1 , a swt2 ,...,a swti ] (47)
[0154] Among them, a swti represents the action in the i-th switch state, S i,t Represents a state collection.
[0155] Step S33: Selecting load shedding in the Distflow power flow optimization constraint model considering distributed generation failures and line operation cost modeling as reward function reference indicators of the distribution network Markov decision process MDP model.
[0156] The instantaneous profit function R(S i,t , a t ): When a specific action a is applied at time t t ∈A t (S i,t ) the network status will change from s t ∈S i,t Change to s t+1 ∈S i,t , s t and s t+1 Represent the states at time t and t+1 respectively. The system observes an immediate reward value R(S i,t , a t ), which is measured based on the network flow calculation of the entire distribution system.
[0157] Specifically, the load shedding in the distribution network flow calculation model Distflow, which considers distributed generation failures, can be selected as the reward function R(S i,t , a t ) is a reference indicator. R(S i,t , a t ) represents the cost of load shedding and line operation cost in the distribution system, and its value is always negative. To find R(S i,t , a t ) is the largest, requiring the lowest load shedding cost. The magnitude of load shedding reflects the distribution system's input-output power balance and system stability. The smaller its value, the more it meets the distribution network's operational requirements and economic budget.
[0158] Step S34: defining the state transition probability of the Markov decision process (MDP) model of the power distribution network, so as to take into account the uncertainty caused by the probability of change of the power output level of each distributed generation unit.
[0159] In this embodiment, the uncertainty caused by the probability of change of the DG power output level can be considered to define the state transition probability of the Markov decision process MDP model in which the state is converted to the state under the action of the action.
[0160] State transition probability P(S j,t+1 |S i,t , a t ): Considering the probability of change of DG power output level from time t to t+1, ∏ kk′ The uncertainty brought about by state S i,t In action a t Under the action of j,t+1 The state transition probability P(S j,t+1 |S i,t , a t ) can be expressed as the following formula:
[0161] P(S j,t+1 |S i,t , a t )=∏ kk′ × P(Ξ t+1 |Ξ t , a t ) (48)
[0162] Where P(Ξ t+1 |Ξ t , a t ) represents the distribution network topology at time t. t In action a t Converted to Ξ t+1 The probability of . Due to Ξ t is changed to Ξ due to network reconstruction t+1 The topology conversion of the power distribution network is certain under a specific reconstruction operation, so the probability is "100%" or "0%". kk′ is the probability of an output level k at time t to another output level k′ at time t+1 when there is only one DG; if there are n DGs, the probability of the output level transition of each DG should be multiplied, that is, the above formula should be written as n ∏ kk′ The multiplication form of .
[0163] Step S4: Use the reinforcement learning algorithm of approximate dynamic programming to solve the Markov decision process MDP model of the distribution network and output the optimal strategy for automatic control of distribution network switches in real time.
[0164] In the step S4, before outputting the optimal strategy for automatic control of distribution network switches in real time, the method further comprises: performing offline iterative training on the Markov decision process MDP model of the distribution network.
[0165] The steps of offline iterative training of the Markov decision process (MDP) model of the distribution network include: inputting the active power output data of the distributed generation units, outputting the automatic control strategy of the distribution network switches, and feeding back the evaluation indicators of the real-time operation of the distribution network.
[0166] According to the distribution network Markov model and Bellman optimal equation established in step S3, and taking into account the uncertainty and volatility of DG output, the dynamic characteristics of the entire system are represented by the four-parameter probability distribution P(S j,t+1 |S i,t , a t ). The state-value function of recursive optimization can be expressed as follows:
[0167]
[0168] Where γ is the discount factor, γ∈[0,1], the degree to which γ is close to 1 indicates the importance of future returns; on the contrary, the degree to which γ is close to 0 indicates the importance of current returns, r(S i,t , a t ) is the state S at the current moment i,t and action a t The immediate reward value of the environment feedback, v t+1 (S j,t+1 ) is the state S at the next moment j,t+1 The state-value function value. Let γ = 1 and r(S i,t , a t ) put out the brackets, we get the following formula:
[0169]
[0170]
[0171] in, Indicates the current state S i,t and action a t Under the premise, the next moment state S j,t+1 The expected value of the state-value function, the immediate benefit function R(S i,t , a t ) consists of two parts: the cost of shedding loads in the distribution system and the cost of line operation, and its value is always negative. Therefore, R(S i,t , a t ) can be expressed as:
[0172] R(S i,t , a t )=∑ b∈B (c b *p sd,b )+∑ l∈L (c1*μ b,b′ ) (52)
[0173] Among them, B is the node set, c b is the load shedding cost coefficient, p sd,bis the active load shedding at node b, c1 is the line operation cost coefficient, μ b,b′ is the line operating cost from node b to node b′. After introducing the concept of the post-decision state, S i,t , a t Rewritten into the form of state variables after decision making Will Defined as This achieves the formal transformation of the mathematical expectation solution process and obtains the following formula:
[0174]
[0175]
[0176]
[0177]
[0178] in, Indicates the state after decision The state-value function value of Indicates the state S at the previous moment j,t-1 and action a t-1 Under the premise, the current state S i,t The expected value of the state-value function, R t+1 (S j,t+1 , a t+1 ) represents the state S at the next moment j,t+1 and action a t+1 The immediate reward value of the environment feedback, Indicates the state after decision The state-value function value of the above formula converts the multi-cycle and large-scale MDP-based stochastic model into a single-cycle deterministic model for each state in each decision cycle and can be solved by iteration
[0179] Introducing the variable n to express the nth iteration, we get the following formula:
[0180]
[0181] in, Indicates the state S in the nth iteration i,t The state-value function value of Indicates the state after decision in the nth iteration The state-value function value of the decision-making state variable is combined with the state-value function recursive form, and then the forward dynamic algorithm is used to update the state of the decision-making state variable.
[0182]
[0183] in, Indicates the state after decision in the nth iteration The state-value function value of Represents the state after decision in the n-1th iteration The state-value function value of Indicates the state S in the nth iteration j,t The state-value function value, the coefficient α is a smoothing parameter less than 1, relying on this iterative update formula and a sufficient number of iterations, the state-value of each state after the decision is A corresponding convergence value can be obtained. And the value of different switching actions in each Markov state can be obtained. Finally, the optimal action and decision can be found by comparing the sizes of these values.
[0184] The specific pseudo code of the ADP algorithm is as follows:
[0185] Step 1 Initialization.
[0186] Step 1a sets the network topology state and DG fluctuation output state.
[0187] Step 1b Initialization and State-value function table.
[0188] Step 1c sets the step size α∈(0, 1] and the number of iterations N.
[0189] Step 2 Do when t=1, 2, ..., T:
[0190] Step 2a uses Gurobi mathematical programming optimizer to solve the mixed integer linear programming R(S i,t , a t ), and solve for Among them, let a t To solve the maximum value problem.
[0191]
[0192] Step 2b uses the forward dynamic algorithm to update
[0193]
[0194] Step 2c From state S i,t and action a t To get the decision status Determine whether t is less than T, and if it is not less than T, perform the following steps;
[0195] Step 2d: Based on the DG fluctuation and network topology, the decision-making state Enter the next moment Markov state S j,t+1 .
[0196] Step 3 sets n=n+1. Determine whether n is less than or equal to the number of iterations N. If so, repeat step 2; otherwise, end the loop and proceed to step 4.
[0197] Step 4 returns the convergence value And its corresponding Markov state S i,t .
[0198] The flow chart of the approximate dynamic programming algorithm is as follows Figure 3 shown.
[0199] The learning and training process of the MDP model for automatic control of distribution network switches under DG output power fluctuation is an iterative process of offline calculation. Figure 4 The purpose of offline calculation is to obtain the convergence value of the state value after the decision is made. Approximate dynamic programming algorithm can achieve this goal. The input information of offline calculation includes the distribution system network topology and the uncertainty and volatility of DG output power. This information is fed into the ADP algorithm and returned as It is a A multi-cycle process that iterates repeatedly and eventually converges.
[0200] In step S4, after outputting the optimal strategy for automatic control of the distribution network switches in real time, the method further comprises: displaying the decision result textually and the real-time distribution network topology structure on the human-computer interaction interface.
[0201] After the training and learning of the MDP model for automatic control of distribution network switches under DG output power fluctuations, it is necessary to automatically predict the optimal switching strategy. The prediction stage and decision stage are an online calculation process. The online calculation flow chart is as follows: Figure 5 shown.
[0202] The role of online computing is to realize the real-time Markov state S observed by the agent at each moment. j,t The optimal switching strategy of the distribution network is obtained. Figure 5 As shown, the post-decision state value output in offline calculation and the Markov state S observed in each time period i,t As the input of the single-period deterministic model, the objective function is still R(S i,t , a t ), whose constraints are the distribution network structure and flow constraints proposed in step S1, and then return and the best action A t (S i,t ). In this case, the online calculation process can obtain the real-time Markov state S at each decision moment. i,t Maximum value strategy.
[0203] Figure 6 This is the overall flow chart of the automatic control method of distribution network switches based on reinforcement learning theory of the present invention. Figure 6 As shown, the proposed method consists of four main parts: derivation of relevant theoretical knowledge, establishment of an MDP model for automatic control of distribution network switches under DG fluctuations, solution of the MDP model and analysis of the results, and completion of a human-computer interaction app. The theoretical derivation section derives the Bellman equation based on reinforcement learning and Markov decision models, and introduces the basic principles of dynamic programming and approximate dynamic programming. The MDP model takes into account the fluctuations of renewable energy generation units such as wind and photovoltaic power generation due to the natural environment, and establishes a Markov decision model for automatic control of distribution network switches. First, a model for the transition fluctuations between different output levels of distributed generation units is established. Then, based on the MDP model concept from reinforcement learning theory, an MDP model for the distribution network is established. Finally, the corresponding program is written based on the structural constraints of the distributed distribution network and the Distflow power flow calculation constraints. The solution of the MDP model addresses the problems and difficulties encountered by traditional reinforcement learning algorithms in solving MDP models. An ADP algorithm for automatic control of distribution network switches is proposed to address these issues. The core method is the introduction of post-decision states and a forward dynamic algorithm. The goal is to transform multi-period and large-scale MDP-based stochastic models into single-period deterministic models for each state within each decision cycle. Finally, a human-computer interaction app (Application) will be developed to enable users to interact with the distribution network switch automatic control software system and directly view the text and image results of the optimal decision.
[0204] Figure 7 FIG is a structural block diagram of a distribution network switch automatic control system based on reinforcement learning theory according to an embodiment of the present invention. Figure 7 As shown, the distribution network switch automatic control system 100 based on reinforcement learning theory includes an establishment module 10, a determination module 20, an acquisition module 30 and a calculation module 40, wherein the establishment module 10 is connected to the determination module 20, the acquisition module 30 and the calculation module 40 respectively, and the calculation module 40 is also connected to the determination module 20.
[0205] Among them, the establishment module 10 is used to establish the Distflow flow optimization constraint model. The determination module 20 is used to determine the reinforcement learning algorithm of approximate dynamic programming. The acquisition module 30 is used to obtain the historical output data of photovoltaic and wind power distributed generation units, so that the establishment module 10 can establish the output power fluctuation and conversion model of the distributed generation unit based on the historical output data. The establishment module 10 is also used to establish the distribution network Markov decision process MDP model based on the distribution network topology determined by the determination module 20, the active output data of the distributed generation unit obtained by the acquisition module 30, the controllable segmented switches and the interconnection switch status information of the distribution network and the calculation result information of the Distflow flow optimization constraint model, as well as the output power fluctuation and conversion model of the distributed generation unit. The calculation module 40 is used to solve the distribution network Markov decision process MDP model according to the reinforcement learning algorithm of approximate dynamic programming, and output the optimal strategy for automatic control of the distribution network switches in real time.
[0206] It should be noted that the specific implementation of the distribution network switch automatic control system based on reinforcement learning theory in the embodiment of the present invention can refer to the specific implementation of the distribution network switch automatic control method based on reinforcement learning theory mentioned above. To avoid redundancy, it will not be repeated here.
[0207] In summary, the present invention proposes a method and system for automatic control of distribution network switches based on reinforcement learning theory. The present invention can fully consider the working efficiency fluctuations of renewable energy sources such as wind power generation and solar power generation, and establish a Markov decision process (MDP) model for automatic control of distribution network switches. The model combines the general distribution network configuration model with the power supply reliability rate, focuses on the power output fluctuations of each distributed power generation unit, and can provide an efficient and stable distribution network switch automatic control optimization strategy. The present invention uses the MDP process to achieve the main purposes of improving economy, reliability, reducing losses, and balancing loads, and completes the network reconstruction of the distribution network containing distributed power generation units. Based on this model, the optimal strategy within each time interval is dynamically formulated according to the real-time state and state transition probability. The present invention uses an approximate dynamic programming (ADP) algorithm to solve the above-mentioned MDP model, which can solve the "dimensionality curse" problem. The present invention aims to optimize and solve the working fluctuations, failures, and faults of distributed power generation units in the distribution network, improve the power supply reliability of the distribution system, and improve the investment efficiency of the distribution system.
[0208] It should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply the existence of any such actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or device comprising the element.
[0209] Although the present invention has been described in detail through the above preferred embodiments, it should be understood that the above description is not intended to limit the present invention. After reading the above description, various modifications and substitutions of the present invention will become apparent to those skilled in the art. Therefore, the scope of protection of the present invention should be defined by the appended claims.
Claims
1. A distribution network switch automatic control method based on reinforcement learning theory, characterized in that: include: Step S1: Establish a Distflow power flow optimization constraint model and determine a reinforcement learning algorithm for approximate dynamic programming; Step S2: Obtain historical output data of photovoltaic and wind power distributed generation units, and establish an output power fluctuation and conversion model of the distributed generation units based on the historical output data; Step S3: Determine the topology of the distribution network, obtain active power output data of distributed generation units, status information of controllable section switches and tie switches of the distribution network, and calculation results of the Distflow power flow optimization constraint model, and establish a Markov decision process (MDP) model of the distribution network based on the output power fluctuation and conversion model of the distributed generation units, the distribution network topology, and the information obtained in step S3; wherein step S3 includes: Step S31: Modeling the output power fluctuation status quantization value of each distributed power generation unit and the corresponding distribution network topology as state parameters of the distribution network Markov decision process MDP model; Step S32: Modeling the action states of the switches connected to the branches of the power distribution network as action combination parameters of the Markov decision process (MDP) model of the power distribution network; Step S33: Selecting the load shedding and line operation cost modeling in the Distflow power flow optimization constraint model considering distributed generation failure as the reward function reference index of the distribution network Markov decision process MDP model; Step S34: defining the state transition probability of the Markov decision process (MDP) model of the power distribution network, so as to take into account the uncertainty caused by the probability of change of the power output level of each distributed generation unit; Step S4: using the reinforcement learning algorithm of approximate dynamic programming to solve the Markov decision process MDP model of the distribution network, and outputting the optimal strategy for automatic control of distribution network switches in real time.
2. The automatic control method for distribution network switches based on reinforcement learning theory according to claim 1, characterized in that: The step S1 comprises: Step S11: establishing the Distflow power flow optimization constraint model according to the distribution network topology constraints and the basic theory of distribution network power flow calculation; Step S12: Determine the reinforcement learning algorithm of approximate dynamic programming according to reinforcement learning theory and Bellman optimal equation.
3. The automatic control method for distribution network switches based on reinforcement learning theory according to claim 1, characterized in that: In step S2, before establishing the output power fluctuation and conversion model of the distributed power generation unit, the method further includes: analyzing and quantifying the output power fluctuation and uncertainty of each distributed power generation unit, so as to use the quantified value as the input of the Markov decision process MDP model of the distribution network.
4. The automatic control method for distribution network switches based on reinforcement learning theory according to claim 3, characterized in that: The output power fluctuation and conversion model of the distributed power generation unit can simulate the fluctuation and uncertainty of the output power of each distributed power generation unit.
5. The automatic control method for distribution network switches based on reinforcement learning theory according to claim 1, characterized in that: In the step S4, before outputting the optimal strategy for automatic control of the distribution network switches in real time, the method further includes: performing offline iterative training on the Markov decision process MDP model of the distribution network.
6. The automatic control method for distribution network switches based on reinforcement learning theory according to claim 5, characterized in that: The steps of performing offline iterative training on the distribution network Markov decision process MDP model include: inputting active power output data of distributed generation units, outputting automatic control strategies for distribution network switches, and feeding back evaluation indicators of real-time operation of the distribution network.
7. The automatic control method for distribution network switches based on reinforcement learning theory according to claim 1, characterized in that: In the step S4, after the optimal strategy for automatic control of the distribution network switches is output in real time, the method further comprises: displaying the decision result in text and the real-time distribution network topology on a human-computer interaction interface.
8. The automatic control method for distribution network switches based on reinforcement learning theory according to claim 1, characterized in that: The action state of the switch includes: the switch changing the switch state at the current moment and maintaining the switch state at the current moment.
9. A distribution network switch automatic control system based on reinforcement learning theory, characterized in that: include: Establish a module for establishing the Distflow power flow optimization constraint model; A determination module for determining a reinforcement learning algorithm that approximates dynamic programming; An acquisition module is used to acquire historical output data of photovoltaic and wind power distributed generation units, so that the establishment module can establish an output power fluctuation and conversion model of the distributed generation units based on the historical output data; The establishment module is also used to establish a distribution network Markov decision process (MDP) model based on the distribution network topology determined by the determination module, the active power output data of the distributed generation units obtained by the acquisition module, the controllable section switches and the interconnection switch status information of the distribution network, and the calculation result information of the Distflow power flow optimization constraint model, as well as the output power fluctuation and conversion model of the distributed generation units, including: modeling the output power fluctuation status quantization value of each distributed generation unit and the corresponding distribution network topology as state parameters of the distribution network Markov decision process (MDP) model; modeling the action state of the switch connecting each branch line of the distribution network as the action combination parameter of the distribution network Markov decision process (MDP) model; selecting the load shedding and line operation cost modeling in the Distflow power flow optimization constraint model considering distributed generation faults as the reward function reference index of the distribution network Markov decision process (MDP) model; defining the state transition probability of the distribution network Markov decision process (MDP) model to consider the uncertainty brought about by the probability of change in the power output level of each distributed generation unit; The calculation module is used to solve the Markov decision process MDP model of the distribution network according to the reinforcement learning algorithm of approximate dynamic programming, and output the optimal strategy for automatic control of distribution network switches in real time.
Citation Information
Patent Citations
D3QN-based active power distribution network multi-target reactive power control method
CN113937829A
Distributed energy storage voltage regulation method adaptive to topology dynamic change of power distribution network
CN115276067A