A method for controlling reactive power and voltage of a power grid
By building an interactive training framework and policy network based on Markov decision-making in the distribution network, the problem of low efficiency and poor convergence of self-learning reactive voltage optimization of distribution network under multiple discrete equipment access is solved, and efficient and reliable grid reactive voltage control is achieved, which improves voltage quality and reduces grid loss.
Patent Information
- Application Number
- CN202310094269.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-02-08
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2043-02-08
AI Technical Summary
The existing distribution network self-learning reactive voltage optimization method is difficult to achieve efficient and reliable online learning under the condition of multiple discrete devices, resulting in low learning efficiency and poor convergence.
A reactive voltage control method for the power grid is proposed. By constructing an interactive training framework based on Markov decision-making, combining the policy network and evaluation network, outputting the decoupled policy function network and the value function of the affine structure, the generation of intelligent control strategies and the issuance of action variables is realized.
This method can improve optimization strategies through online learning and realize efficient and reliable multi-discrete equipment control, improve the voltage quality of the distribution network, and reduce operating network losses under the condition of incomplete models.
Smart Images

Figure CN116316659B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of power system operation and control, and in particular, to a method for controlling reactive power and voltage of a power grid. Background Art
[0002] The continuously increasing DG penetration rate in the distribution network poses severe challenges to the operation of the distribution network. Among them, voltage over-limit in high-penetration areas has greatly affected the normal grid connection operation of end-users and DGs. It is becoming increasingly important to adopt refined regulation and optimization methods to eliminate voltage over-limit and maintain the efficient operation of the distribution network. In addition, problems such as reverse power flow, deterioration of power quality, and device disconnection in the distribution network with high DG penetration rate also urgently need to be solved by exploring the regulation ability of flexible resources.
[0003] A series of data-driven model-free optimization methods represented by deep reinforcement learning have been widely applied in the field of power system operation and control, and have also become an important technical route for reactive power and voltage optimization in distribution networks with incomplete models. Based on the deep reinforcement learning algorithm, a number of self-learning reactive power optimizations for distribution networks have emerged in the industry, which can continuously improve the optimization strategy through continuous online interaction, so as to improve the optimization goal under the condition of incomplete models. However, the types of controllable resources in the distribution network are complex and their characteristics are different, which bring many problems to data-driven methods and reinforcement learning methods. Although DG has a certain reactive power regulation ability, the current main reactive power and voltage regulation means on site are still large-scale voltage regulators and compensators. These devices are all step-regulated, that is, their controlled variables are discrete values, simply referred to as discrete devices. Therefore, the reactive power and voltage optimization of the distribution network often needs to give the setting values of multiple discrete devices. In the optimization problem, this will change the nature of the problem into a mixed integer programming problem; in reinforcement learning, the access of multiple discrete devices means that the action space is a combined discrete space; in these cases, there are problems of combinatorial explosion. Therefore, the existing self-learning reactive power and voltage optimization methods for distribution networks are difficult to apply under the condition of accessing multiple discrete devices, and often have problems of low learning efficiency and poor convergence.
[0004] In view of this, this specification proposes a method for controlling reactive power and voltage of a power grid, which can not only improve the optimization strategy through online learning under the condition of incomplete models, but also achieve efficient and reliable online learning under the condition of accessing multiple discrete devices, and further support the continuous online operation of the self-learning reactive power optimization method for distribution networks, greatly improving the voltage quality of the distribution network and reducing the operating network loss. Summary of the Invention
[0005] The object of the present invention is to provide a method for controlling reactive power and voltage of a power grid, including: constructing a control optimization model based on the reactive power and voltage control objectives of the distribution network; the control optimization model is used to adjust a plurality of discrete devices under constraint conditions so that the distribution network reaches the optimization objective; constructing an interactive training framework based on Markov decision-making based on the control optimization model; the interactive training framework is used to determine an intelligent control strategy based on the configuration of the distribution network; inputting the current state information into an action generation model, and the model outputs an action variable based on the intelligent control strategy; and sending the action variable to the plurality of discrete devices.
[0006] Further, the optimization objective O T has the following expression:
[0007]
[0008] where T represents the number of control cycles of the plurality of discrete devices; t represents the selected moment; T O represents the tap position of the on-load tap changer OLTC; T B represents the tap position of the capacitor station; C O represents the OLTC adjustment cost, C B represents the capacitor station adjustment cost; C P represents the network loss cost coefficient; represents the tap change amount of the OLTC adjustment at time t; represents the tap change amount of the capacitor station adjustment at time represents the network loss of the active power at time t;
[0009] The expression of the constraint conditions includes:
[0010] Voltage constraint:
[0011]
[0012] where represents the set of all nodes of the distribution network; V i t represents the voltage amplitude of node i at time t; V and represent the lower limit and the upper limit of the node voltage respectively;
[0013] Regulation constraint:
[0014]
[0015]
[0016] where Represents the tap setting value of the $i$-th OLTC device at time $t$; Represents the number of taps of the $i$-th OLTC device; $n$ OLTC Represents the total number of OLTC devices; Is the tap setting value of the $j$-th capacitor bank at time $t$; Represents the number of taps of the $j$-th capacitor bank; $n$ CB Represents the total number of capacitor banks.
[0017] Furthermore, the intelligent control strategy includes state information $s$, feedback variable $r$, and action variable $a$;
[0018] The expression of the state information $s$ is:
[0019] $s=(P,Q,V,T$ O ,$T$ B ) t
[0020] where $P$ and $Q$ respectively represent the vectors composed of the active power injection and reactive power injection of each node in the distribution network; $V$ represents the vector composed of the voltages of each node in the distribution network; $T$ O represents the vector composed of the taps of each OLTC device; $T$ B represents the vector composed of the taps of each capacitor bank; $t$ represents the moment when the state information is obtained;
[0021] The expression of the feedback variable $r$ is:
[0022] $r = -C$ O $T$ O,loss $(s,s') - C$ B $T$ B,loss $(s,s') - C$ P $P$ loss $(s') - C$ V $V$ loss $(s')$
[0023] where $s$ and $s'$ respectively represent the current state information at time $t$ and the current state information at time $t + 1$; $T$ O,loss $(s,s')$ represents the tap change amount of the OLTC adjustment from state $s$ to state $s'$; $T$ B,loss $(s,s')$ represents the tap change amount of the capacitor bank adjustment from state $s$ to state $s'$; $P$ loss $(s')$ represents the network loss at time $t + 1$; $C$ V represents the voltage violation rate cost coefficient; $V$ loss $(s')$ represents the voltage violation amount at time $t + 1$;
[0024] The expression of the action variable $a$ is:
[0025] $a=(T$a O , T a B ) t
[0026] where T a O represents the vector composed of the positions of each OLTC device after performing action a; T a B represents the vector composed of the positions of each capacitor station after performing action a.
[0027] Furthermore, it further includes: constructing the action generation model; initializing the action generation model based on the actual situation of the distribution network; and updating the action generation model.
[0028] Furthermore, the constructing of the action generation model includes:
[0029] constructing a policy network; the input of the policy network is the state information s, and the output is the probability of taking action a when the state information is s, i.e., a ∼ π(·|s); the constructing of the policy network includes:
[0030] constructing a first feature processing module; the first feature processing module is composed of a neural network with three hidden layers; the input of the first layer is the state information, and the inputs of the second and third layers are the outputs of the previous layer; the output size of each layer is 512, each layer contains 512 neurons, and the activation function is selected as the ReLU function;
[0031] constructing a first multi-head output module; the first multi-head output module includes multiple first output modules, and the multiple first output modules respectively correspond to the multiple discrete devices; for each first output module, it is composed of a neural network with two hidden layers, the input of the first layer is the output of the feature processing module, the output size is 512, it contains 512 neurons, and the activation function is selected as ReLU; the input of the second layer is the output of the first layer, the output size is the total number of the multiple discrete devices, and the activation function is selected as the Softmax function;
[0032] constructing a value network; the input of the value network is the state information s, and the output is the value Q(s, a) of taking action a when the state information is s; the constructing of the value network includes:
[0033] constructing a second feature processing module; the second feature processing module is composed of a neural network with three hidden layers; the input of the first layer is the state information, and the inputs of the second and third layers are the outputs of the previous layer; the output size of each layer is 512, each layer contains 512 neurons, and the activation function is selected as the ReLU function;
[0034] Construct an affine coefficient module; the affine coefficient module is composed of a neural network including two hidden layers. The input of the first layer is the output of the second feature processing module. The output size of the first layer is 512, including 512 neurons, and the activation function selects the ReLU function. The input of the second layer is the output of the first layer. The output size of the second layer is N + 1, and the activation function selects the linear function; where N represents the total number of discrete devices.
[0035] Construct a second multi-head output module; the second multi-head output module includes a plurality of second output modules, and the plurality of second output modules respectively correspond to the plurality of discrete devices; for each second output module, it is composed of a neural network including two hidden layers. The input of the first layer is the output of the second feature processing module, the output size is 512, including 512 neurons, and the activation function selects the ReLU function. The input of the second layer is the output of the first layer, and the output size is the total number of the plurality of discrete devices; the activation function selects the linear function.
[0036] Further, the initialization of the action generation model includes:
[0037] Randomly initialize the network parameters θ of the policy network and the network parameters φ of the evaluation network;
[0038] Determine the discount factor γ, the maximum entropy parameter α of the action generation model, the learning step ρφ of the evaluation network, the learning step ρθ of the policy network, and the maximum learning period T of the action generation model max ;
[0039] Initialize the discrete time variable t, the control period Δt, and the model update period k;
[0040] Initialize the experience library D.
[0041] Further, the update of the action generation model includes:
[0042] Judge whether the current time is in the update period of the action generation model;
[0043] If not, input the state information of the next moment into the action generation model, and the model outputs the action variable of the next moment;
[0044] If so, extract the experience group from the experience library;
[0045] Input the experience group into the action model, and construct a loss function based on the output of the model and the data of the experience group;
[0046] Based on the loss function, update the network parameters θ of the policy network and the network parameters φ of the evaluation network;
[0047] Determine whether the action generation model meets a preset condition;
[0048] If so, stop updating the action generation model.
[0049] Furthermore, update the loss function of the policy network The expression is:
[0050]
[0051] where s represents state information; D B represents the extracted experience group; represents the expectation of V(s); V(s) represents the affine result of the value function;
[0052] The expression for updating the network parameter θ of the policy network is:
[0053]
[0054] where ρθ represents the learning step size of the policy network, represents the gradient with respect to θ;
[0055] The expression for updating the loss function of the evaluation network is:
[0056]
[0057] where s represents the state information in the experience group; a represents the action variable corresponding to the current state information s in the experience group; s′ represents the state information at the next moment in the experience group; r represents the feedback for taking action a when the current state information is s in the experience group; y represents the value function update target; Q π (s,a) represents the output of the evaluation network;
[0058] The expression for updating the network parameter φ of the evaluation network is:
[0059]
[0060] where ρ φ represents the learning step size of the evaluation network; represents the gradient with respect to φ.
[0061] Furthermore, the expression for the affine result V(s) of the value function is:
[0062]
[0063] where c i(s) represents the output of the affine coefficient module; s represents the state information; α represents the maximum entropy parameter of the action generation model; N represents the total number of discrete devices in the distribution network; T represents the gear position of the i-th discrete device; Q i (s, ·) represents the output of the second multi-head output module.
[0064] Furthermore, it also includes obtaining experience samples, including:
[0065] Obtain the state information at the next moment;
[0066] Based on the current state information and the state information at the next moment, determine the feedback variable;
[0067] Take the current state information, the action variable, the feedback variable, the state information at the next moment, and the current experience library as the experience sample;
[0068] Input the experience sample into the experience library to obtain an updated experience library.
[0069] The technical solution of the embodiment of the present invention has at least the following advantages and beneficial effects:
[0070] 1. A reactive power voltage control method for power grids proposed by the present invention, compared with the traditional reactive power voltage optimization method for accessing multiple discrete devices, has the characteristic of obtaining the optimal control strategy through online learning without the need for an accurate model of the distribution network. This characteristic can not only avoid the optimization error problem caused by model errors, but also eliminate the high cost of frequent model maintenance, improve the efficiency and safety of the distribution network operation, and is suitable for deployment in actual power systems.
[0071] 2. The present invention uniquely designs a strategy function network with output decoupling (i.e., the policy network). On the one hand, the action strategies of each device share the output of the feature processing module, and thus have extremely high efficiency in feature learning; on the other hand, the action strategies of each device are output decoupled and can be calculated independently, avoiding the problem of combinatorial explosion. Compared with the traditional reactive power voltage control method based on reinforcement learning, this method can ensure that when multiple discrete devices are connected, all devices can be controlled simultaneously, so that their adjustment capabilities can be utilized to the greatest extent, and then fully optimize the operation state of the distribution network and improve the consumption of renewable energy.
[0072] 3. The present invention uniquely designs an affine structure value function. On the one hand, it transforms the value solution with the original combinatorial explosion problem into an affine combination of independent outputs, and adaptively adjusts the affine coefficients through learning, ensuring the expression ability of the model and thus the learning effect. On the other hand, this structure can make full use of the output values of each discrete device, avoid the inefficient sampling link during the learning stage, and thus can greatly improve the sample efficiency, and can obtain a fully optimized control strategy after a small number of iterations, which is applicable to scenarios where samples are precious in the distribution network. BRIEF DESCRIPTION OF THE DRAWINGS
[0073] Figure 1 An exemplary flowchart of a power grid reactive voltage control method provided by some embodiments of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0074] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. Usually, the components of the embodiments of the present invention described and illustrated herein can be arranged and designed in various different configurations.
[0075] Figure 1 An exemplary flowchart of a power grid reactive voltage control method provided by some embodiments of the present invention. As Figure 1 shown, process 100 may include the following:
[0076] Step 110, based on the reactive voltage control objective of the distribution network, construct an optimization model for the reactive voltage control of the distribution network with multiple discrete devices; the control optimization model is used to adjust multiple discrete devices under constraint conditions so that the distribution network reaches the optimization objective. The specific steps are as follows:
[0077] Step 111, determine the variable form (i.e., the optimization objective) of the controlled multiple discrete devices:
[0078]
[0079] where T represents the number of control periods for controlling the multiple discrete devices. For example, the number of control periods of the discrete device within a day, and a typical value is 24; t represents the selected moment for counting the control period of the discrete device; T O represents the on-load tap changer (OLTC) tap position; T B represents the capacitor bank tap position; C O represents the OLTC regulation cost, C B represents the capacitor bank regulation cost, C P represents the active power network loss cost coefficient, It represents the change in the tap position of the OLTC at time t. It represents the change in the tap position of the capacitor station at time t. It represents the network loss of the active power at time t. They are calculated by the following formulas respectively:
[0080]
[0081]
[0082]
[0083] Among them, is the tap setting value of the i-th OLTC device at time t, and n OLTC is the total number of OLTC devices; is the tap setting value of the i-th capacitor station at time t, and n CB is the total number of capacitor station devices.
[0084] Step 112: Establish the reactive power voltage optimization constraints for the controlled distribution network with multiple discrete devices, including:
[0085] Voltage constraint:
[0086]
[0087] Among them, represents the set of all nodes in the distribution network; V i t represents the voltage amplitude of node i at time t; V and represent the lower limit and upper limit of the node voltage respectively.
[0088] Regulation constraint:
[0089]
[0090]
[0091] Among them, represents the tap setting value of the i-th OLTC device at time t; represents the number of taps of the i-th OLTC device; n OLTC represents the total number of OLTC devices; is the tap setting value of the j-th capacitor station at time t; represents the number of taps of the j-th capacitor station; n CB represents the total number of capacitor stations.
[0092] Step 120: Based on the control optimization model, construct an interactive training framework based on Markov decision-making; the interactive training framework is used to determine an intelligent control strategy based on the configuration of the distribution network.
[0093] Step 121: Corresponding to the measurement of the distribution network system, construct the state information s of the Markov decision process at time t as shown in the following formula:
[0094] s = (P, Q, V, T O , T B ) t
[0095] where P and Q respectively represent the vectors composed of the active power injection and reactive power injection of each node in the distribution network; V represents the vector composed of the voltages of each node in the distribution network; T O represents the vector composed of the tap positions of each OLTC device; T B represents the vector composed of the tap positions of each capacitor bank; t represents the time when the state information is obtained, and (·) t represents the measured value at time t;
[0096] Step 122: Corresponding to the reactive power and voltage optimization objective of multiple discrete devices, construct the feedback variable r of the Markov decision process as shown in the following formula:
[0097] r = -C O T O,loss (s, s′) - C B T B,loss (s, s′) - C P P loss (s′) - C V V loss (s′)
[0098] where s and s′ respectively represent the current state information at time t and the current state information at time t + 1; T O,loss (s, s′) represents the tap change amount of the OLTC adjustment from state s to state s′; T B,loss (s, s′) represents the tap change amount of the capacitor bank adjustment from state s to state s′; P loss (s′) represents the network loss at time t + 1; C V represents the voltage violation rate cost coefficient; V loss (s′) represents the voltage violation amount at time t + 1.
[0099] The calculation expressions for the network loss P loss (s′) at time t + 1 and the voltage violation amount V loss (s′) at time t + 1 are as follows:
[0100]
[0101]
[0102] Step 123, construct the multi-discrete device agent action variable a at time t, as shown in the following formula:
[0103] a = (T a O , T a B ) t
[0104] where T a O represents the vector composed of the tap positions of each OLTC device after executing action a; T a B represents the vector composed of the tap positions of each capacitor station after executing action a. For the i-th discrete device, its action variable at time t corresponds to a i .
[0105] Step 130, input the current state information into the action generation model, and the model outputs the action variable based on the intelligent control strategy.
[0106] In some embodiments, it further includes constructing the action generation model; initializing the action generation model based on the actual situation of the distribution network; and updating the action generation model.
[0107] In some embodiments, constructing the action generation model includes:
[0108] Construct a policy network π, and all its parameters are denoted as θ; the input of the policy network is the state information s, and the output is the probability of taking action a when the state information is s, a ~ π(·|s). The construction of the policy network includes:
[0109] Construct a first feature processing module; the first feature processing module is composed of a neural network with three hidden layers; the input of the first layer is the state information, and the inputs of the second and third layers are the outputs of the previous layer; the output size of each layer is 512, and each layer contains 512 neurons. The activation function is selected as the ReLU function; the module parameters are denoted as θ h . The expression of this module is denoted as
[0110] Construct a first multi-head output module; the first multi-head output module includes a plurality of first output modules, and the number is N = n CB + n OLTC pieces, and the plurality of first output modules respectively correspond to the plurality of discrete devices; for each first output module, it is composed of a neural network with two hidden layers, and the module parameters are denoted as The input of the first layer is the output of the feature processing module The output size is 512, containing 512 neurons, and the activation function is ReLU; the input of the second layer is the output of the first layer, and the output size is the total number of the multiple discrete devices, denoted as n i , that is, when i ≤ n OLTC when when i > n OLTC when The activation function of the second hidden layer is the Softmax function, and the expression is:
[0111]
[0112] The output of the i-th multi-head output module corresponds to the probability distribution of all discrete actions of the i-th discrete device, denoted as a i ~π i (·|s). Then the final output form of the multi-discrete device policy network π is a~π(·|s), and the probability is:
[0113]
[0114] All parameters are Specific implementation can use open-source neural network solvers such as PyTorch
[0115] Construct the evaluation network Q, and all parameters are denoted as φ; the input of the evaluation network is the state information s, and the output is the value Q(s,a) of taking the action a when the state information is s; the construction of the evaluation network includes:
[0116] Construct the second feature processing module, and the expression of the second feature processing module is denoted as The second feature processing module consists of a neural network with three hidden layers; the input of the first layer is the state information, and the input of the second and third layers is the output of the previous layer; the output size of each layer is 512, each layer contains 512 neurons, and the activation function is the ReLU function. The parameters of the second feature processing module are denoted as φ h .
[0117] Construct the affine coefficient module; the affine coefficient module consists of a neural network with two hidden layers, and the parameters of the affine coefficient module are denoted as φ a 、The expression is denoted as c(s), and the input of the first layer is the output of the second feature processing module The output size of the first layer is 512, containing 512 neurons, and the activation function is the ReLU function; the input of the second layer is the output of the first layer, the output size of the second layer is N + 1, the activation function is the linear function, and each dimension of the output is denoted as c0(s), c1(s),…, c N(s); where N represents the total number of discrete devices.
[0118] Construct N second multi-head output modules; each second multi-head output module includes a plurality of second output modules, and the plurality of second output modules respectively correspond to the plurality of discrete devices; for each second output module, the module parameters are denoted as It is composed of a neural network with two hidden layers. The input of the first layer is the output of the second feature processing module The output size is 512, including 512 neurons, and the ReLU function is selected as the activation function; the input of the second layer is the output of the first layer, and the output size is the total number of the plurality of discrete devices, denoted as n i , that is, when i ≤ n OLTC At this time, When i > n OLTC At this time, The linear function is selected as the activation function, and the module expression is denoted as Q i (s, ·), when the discrete action corresponding to the i-th discrete device is a i At this time, the output value of this module is Q i (s, a i ).
[0119] The result of the multi-discrete device evaluation network Q is combined by the outputs of the affine coefficient module and the multi-head output module. For a certain state-action combination s, a, the module output is:
[0120]
[0121] All parameters are Specific implementation can use open-source neural network solvers such as PyTorch.
[0122] In some embodiments, initializing the action generation model includes:
[0123] Randomly initialize the network parameters θ of the policy network and the network parameters φ of the evaluation network;
[0124] Determine the discount factor γ, the maximum entropy parameter α of the action generation model, the learning step ρ of the evaluation network φ , the learning step ρθ of the policy network, and the maximum learning period T of the action generation model max . For example, a typical value of the given discount factor γ is 0.99; a typical value of the given maximum entropy parameter α of the action generation model is 0.01; the given learning step ρ of the evaluation network φ and the learning step ρ of the policy network θ The typical value is 0.0001; the given typical value of the maximum learning period T max is 10 6。
[0125] Initialize the discrete-time variable t, the control period Δt, and the model update period k. For example, initialize the discrete-time variable t = 0; a typical value of the given control period Δt is 15 min; a typical value of the initialized model update period k is 8.
[0126] Initialize the experience library D such that
[0127] In some embodiments, updating the action generation model includes:
[0128] Determine whether the current time is within the update period of the action generation model. For example, it can be determined whether the value of t mod k is 0; if t mod k = 0, update the action generation model, and if t mod k ≠ 0, do not update the model.
[0129] When the current time is not within the update period of the action generation model, input the state information of the next moment into the action generation model, and the model outputs the action variable of the next moment. For example, the next moment can be t + Δt.
[0130] When the current time is within the update period of the action generation model, extract the experience group D B ∈ D, and the number of samples in this group of experiences is B;
[0131] Input the experience group into the action model, and construct a loss function based on the output of the model and the data of the experience group. In some embodiments, the expression of the loss function of the evaluation network is:
[0132]
[0133] Among them, s represents the state information in the experience group; a represents the action variable corresponding to the current state information s in the experience group; s′ represents the state information of the next moment in the experience group; r represents the feedback for making the action a when the current state information is s in the experience group; y represents the value function update target; Q π (s,a) represents the output of the evaluation network.
[0134] The expression of the value function update target y is:
[0135] y = r + γV(s′)
[0136] Using the affine structure property of the value function, the calculation method of V(s) is as follows:
[0137]
[0138] Among them, among them, c i(s) represents the output of the affine coefficient module; s represents the state information; α represents the maximum entropy parameter of the action generation model; N represents the total number of discrete devices in the distribution network; T represents the gear position of the i-th discrete device; Q i (s, ·) represents the output of the second multi-head output module.
[0139] In some embodiments, the expression of the loss function of the policy network is:
[0140]
[0141] where s represents the state information; D B represents the extracted experience group; represents the expectation of V(s); V(s) represents the affine result of the value function.
[0142] Based on the loss function, update the network parameters θ of the policy network and the network parameters φ of the evaluation network.
[0143] In some embodiments, the expression for updating the network parameters φ of the evaluation network is:
[0144]
[0145] where ρ φ represents the learning step size of the evaluation network; represents the gradient with respect to φ; can be directly obtained using open-source neural network solvers such as PyTorch.
[0146] In some embodiments, the expression for updating the network parameters θ of the policy network is:
[0147]
[0148] where ρ θ represents the learning step size of the policy network, represents the gradient with respect to θ; can be directly obtained using open-source neural network solvers such as PyTorch.
[0149] Determine whether the action generation model meets a preset condition. The preset condition can refer to that the action generation model reaches a convergence state and the iteration period is greater than the maximum learning period T max .
[0150] If so, stop updating the action generation model.
[0151] Step 140, send the action variable to the multiple discrete devices.
[0152] The discrete devices adjust the gear based on the action variables, enabling the distribution network to reach a new operating state. In some embodiments, new state information can be obtained from the distribution network measurement device at a future time Δt from the current time. ; Calculate r and add an empirical sample to the empirical library and update
[0153] Update s to , and t to t + Δt; According to the state information s, use the multi-discrete device policy network π of the agent to generate the multi-discrete device action a ∼ π(||s); Correspondingly send a to each discrete device, and cycle in this way to realize the control of each discrete device at time t.
[0154] The present invention establishes the reactive power voltage control problem of multi-discrete devices as a Markov decision process, sets up a reinforcement learning model and optimizes each discrete device at the same time. The intelligent model is trained and improved using the reinforcement learning algorithm. The method proposed by the present invention can make full use of the correlation information between multi-discrete devices, improve the training sample efficiency, alleviate the challenge brought by the explosion of the set value combination of multi-discrete devices, and can converge efficiently in the interactive control. The present invention uniquely designs an affine structure value function, makes full use of the correlation of each discrete device in the distribution network system, efficiently mines the system measurement and feedback information to fit the value function, and can greatly improve the learning efficiency of the agent; on the other hand, the present invention uniquely designs a policy function network with output decoupling, reduces the learning complexity, alleviates the problem of the explosion of multi-discrete device combinations, further improves the sample efficiency, quickly obtains the optimal policy of the agent, and ensures the optimal operation of the system in the scenario of incomplete models.
[0155] The above are only the preferred embodiments of the present invention and are not used to limit the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A method for controlling reactive power and voltage of a power grid, characterized in that Including: Based on the reactive power voltage control objective of the distribution network, a control optimization model is constructed; The control optimization model is used to adjust multiple discrete devices under constraint conditions so that the distribution network reaches the optimization objective; Based on the control optimization model, an interactive training framework based on Markov decision is constructed; The interactive training framework is used to determine an intelligent control strategy based on the configuration of the distribution network; Input the current state information into the action generation model, and the model outputs an action variable based on the intelligent control strategy; Send the action variable to the multiple discrete devices; The optimization objective has the following expression: Among them, represents the number of cycles for controlling the multiple discrete devices; represents the selected moment; represents the on-load tap changer (OLTC) tap position; represents the capacitor bank tap position; represents the OLTC regulation cost, represents the capacitor bank regulation cost; represents the network loss cost coefficient; represents the tap change amount of the OLTC regulation at the moment; represents the tap change amount of the capacitor bank regulation at the moment; represents the network loss of the active power at the moment; The expression of the constraint conditions includes: Voltage constraint: Among them, represents the set of all nodes in the distribution network; represents the node at the voltage amplitude at the moment; and respectively represent the lower and upper limits of the node voltage; Regulation constraint: Among them, represents the tap setting value of the th OLTC device at moment; represents the number of taps of the th OLTC device; represents the total number of OLTC devices; is the tap setting value of the th capacitor station at moment; represents the number of taps of the th capacitor station; represents the total number of capacitor stations.
2. The method for controlling the reactive power and voltage of the power grid according to claim 1, wherein The described intelligent control strategy includes state information , feedback variables and action variables ; The said status information has the following expression: Among them, and respectively represent the vectors composed of the active power injection and the reactive power injection of each node in the distribution network; represents the vector composed of the voltages of each node in the distribution network; represents the vector composed of the tap positions of each OLTC device; represents the vector composed of the tap positions of each capacitor station; represents the moment when the state information is obtained; The feedback variable has the following expression: Among them, and respectively represent the current state information at the moment of and represents the step change amount of OLTC regulation from state to state ; represents the step change amount of capacitor bank regulation from state to state ; represents the network loss at the moment of represents the cost coefficient of voltage violation rate; represents the voltage violation amount at the moment of The action variable has the following expression: Among them, represents the vector composed of the positions of each OLTC device after executing the action; represents the vector composed of the positions of each capacitor station after executing the action.
3. The grid reactive power voltage control method according to claim 1, characterized in that It also includes: Construct the action generation model; Initialize the action generation model based on the actual situation of the distribution network; Update the action generation model.
4. The method for controlling reactive power and voltage of a power grid according to claim 3, wherein The construction of the action generation model includes: Construct a policy network; the input of the policy network is state information , and the output is the probability of taking action when the state information is ; the constructing of the policy network includes: Construct a first feature processing module; the first feature processing module is composed of a neural network with three hidden layers; the input of the first layer is the state information, and the inputs of the second and third layers are the outputs of the previous layer; the output size of each layer is 512, each layer contains 512 neurons, and the activation function is selected as the ReLU function; Construct a first multi-head output module; the first multi-head output module includes multiple first output modules, and the multiple first output modules correspond to the multiple discrete devices respectively; for each first output module, it is composed of a neural network with two hidden layers, the input of the first layer is the output of the feature processing module, the output size is 512, it contains 512 neurons, and the activation function is selected as ReLU; the input of the second layer is the output of the first layer, the output size is the total number of the multiple discrete devices, and the activation function is selected as the Softmax function; Construct an evaluation network; the input of the evaluation network is the state information , and the output is the value of taking action when the state information is ; the constructing of the evaluation network includes: ; Construct a second feature processing module; the second feature processing module is composed of a neural network with three hidden layers; the input of the first layer is the state information, and the inputs of the second and third layers are the outputs of the previous layer; the output size of each layer is 512, each layer contains 512 neurons, and the activation function is selected as the ReLU function; Construct an affine coefficient module; the affine coefficient module is composed of a neural network including two hidden layers. The input of the first layer is the output of the second feature processing module. The output size of the first layer is 512, including 512 neurons, and the activation function selects the ReLU function. The input of the second layer is the output of the first layer, and the output size of the second layer is , and the activation function selects the linear function; where represents the total number of discrete devices; Construct a second multi-head output module; the second multi-head output module includes multiple second output modules, and the multiple second output modules correspond to the multiple discrete devices respectively; for each second output module, it is composed of a neural network with two hidden layers, the input of the first layer is the output of the second feature processing module, the output size is 512, it contains 512 neurons, and the activation function is selected as the ReLU function; the input of the second layer is the output of the first layer, the output size is the total number of the multiple discrete devices; the activation function is selected as the linear function.
5. The grid reactive power voltage control method according to claim 4, characterized in that The initialization of the action generation model includes: Randomly initialize the network parameters of the policy network and the network parameters of the evaluation network ; Determine the discount factor , the maximum entropy parameter of the action generation model , the learning step size of the evaluation network , the learning step size of the policy network and the maximum learning period of the action generation model ; Initialize the discrete-time variables , the control period and the model update period ; Initialize the experience library .
6. The grid reactive power voltage control method according to claim 4, characterized in that, The update of the action generation model includes: Judge whether the current time is in the update period of the action generation model; If not, input the state information of the next moment into the action generation model, and the model outputs the action variable of the next moment; If so, extract an experience group from the experience library; Input the experience group into the action generation model, and construct a loss function based on the output of the model and the data of the experience group; Update the network parameters of the policy network based on the loss function and the network parameters of the evaluation network ; Determine whether the action generation model meets a preset condition; If so, stop updating the action generation model.
7. The method for controlling reactive power and voltage of a power grid according to claim 6, characterized in that Update the loss function of the policy network The expression of which is as follows: Among them, represents status information; represents the extracted experience group; represents the expectation of; represents the affine result of the value function; Update the network parameters of the policy network The expression of Among them, represents the learning step size of the said policy network, represents with respect to the gradient of; The expression for updating the loss function of the evaluation network is: Among them, represents the status information in the experience group; represents the current status information in the experience group corresponding action variable; represents the status information at the next moment in the experience group; represents that in the experience group, when the current status information is under the condition of making an action feedback; represents the value function update target; represents the output of the evaluation network; Update the network parameters of the evaluation network The expression of Among them, represents the learning step of the evaluation network; represents with respect to the gradient of.
8. The method for controlling the reactive power and voltage of the power grid according to claim 7, characterized in that The affine result of the value function The expression is as follows: Among them, represents the output of the affine coefficient module; represents the state information; represents the maximum entropy parameter of the action generation model; represents the total number of discrete devices in the distribution network; represents the gear position of the th discrete device; represents the output of the second multi-head output module.
9. The grid reactive power voltage control method according to claim 1, wherein It also includes obtaining experience samples, including: Obtain the state information at the next moment; Based on the current state information and the state information at the next moment, determine the feedback variable; Use the current state information, the action variable, the feedback variable, the state information at the next moment, and the current experience library as the experience sample; Input the experience sample into the experience library to obtain an updated experience library.
Citation Information
Patent Citations
Power distribution network multi-time scale reactive voltage control method based on reinforcement learning
CN113489015A
Power distribution network reactive power optimization method and device based on data driving
CN113746111A