Voltage control strategy modeling and voltage online regulation and control method and system of power distribution network
By adopting the voltage control strategy modeling method of distribution location marginal pricing and maximum entropy reinforcement learning algorithm in the distribution network, the problem of collaborative optimization of virtual power plants in the distribution network is solved, and the safety control of voltage and the maximization of economic benefits is achieved.
Patent Information
- Application Number
- CN202510202029.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-24
- Publication Date
- 2025-06-27
AI Technical Summary
The optimality and efficiency of the coordinated optimization of multiple virtual power plants in the distribution network are affected by the complexity of the distribution network, and the prior art fails to fully consider the maximization of economic benefits when adjusting voltages.
The voltage control strategy modeling method based on marginal pricing of distribution locations is adopted, and the voltage control strategy model is formed to regulate the distribution network voltage online through regional division, regional electricity price determination, optimization model construction and maximum entropy reinforcement learning algorithm training.
It realizes the safe operation of the distribution network voltage, improves the economy of each area (VPP) and distribution network, and optimizes the voltage control of the distribution network and the operating profit of the virtual power plant.
Smart Images

Figure CN120222399A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of distribution network regulation, and particularly to a voltage control strategy modeling, an on-line voltage regulation method and system for a distribution network. Background Art
[0002] In recent years, distributed energy resources (DERs) with intermittent and fluctuating characteristics, such as photovoltaic (PV) power generation and wind turbine (WT) power generation, have been increasingly integrated into the distribution network (DN). Although the operation flexibility of the DN may be enhanced, its voltage violations become more frequent. Effectively controlling small and medium-capacity distributed storage systems and ensuring the safe operation of distributed storage systems have become the research focus. The emergence of the virtual power plant (VPP) has solved these problems.
[0003] The VPP consists of a large number of individual distributed generation devices or loads, forming a dispatchable and controllable power plant, effectively managing distributed generation, and supporting complementary and collaborative operation. The virtual power plant is divided into a technical virtual power plant (TVPP) and a commercial virtual power plant (CVPP). The TVPP mainly provides auxiliary services such as congestion management, peak shaving, and frequency modulation for the operation management of the DN. However, when controlling distributed power sources to regulate the voltage of distributed power sources, the maximization of economic benefits is not fully considered. The CVPP mainly focuses on formulating optimal planning / operation strategies to maximize the economic profit of DERs. Currently, the research on CVPP mainly considers time-of-use (ToU) electricity prices or simplified electricity prices, while distribution locational marginal pricing (DLMP) can be used as a price signal to reflect the grid operation demand, thereby motivating rational end-users to contribute to the optimal operation of the grid. Since the transactive energy (TE) mechanism can incentivize DERs through economic signals and enhance the secure and economic operation of the DN, it is necessary to explore the potential value of TE for VPP management.
[0004] The voltage regulation method can effectively alleviate the voltage violation in the distributed storage system, but it ignores the benefits of the distributed storage system. Some research has proposed a method for effectively managing and controlling multiple micro-network groups to reduce the operating cost of the DN and improve the operating performance of the DN. Some research has further established a stochastic objective function for maximizing profit, coordinating the control of voltages by devices such as PV and WT to achieve the maximization of DN profit. In recent research, a completely decentralized double-loop feedback peer-to-peer energy trading mechanism based on voltage regulation ability has been proposed. There is also research that designs a centralized energy trading mechanism based on coalition graph game to study the overvoltage problem. There is also research that proposes a two-layer constrained peer-to-peer TE framework between multiple microgrids to ensure network security. The lower layer uses peer-to-peer TE, and the upper layer reconfigures the DN according to the results of the lower layer. In some research, the distributed robust method is used to consider the VPP energy trading in the unbalanced DN to prevent communication failures. In the above literature, in order to simplify the expression of the optimization model, ToU is used in TE, and good economic benefits cannot be achieved. And in the process of solving the optimization strategy, the optimality and efficiency of the calculation results are affected by the complexity of the distribution network. Summary of the Invention
[0005] In order to overcome the problem that the complexity of the distribution network affects the optimality and efficiency of the collaborative optimization problem of multiple virtual power plants in the distribution network, the present invention provides a voltage control strategy modeling, voltage online regulation method and system for a distribution network.
[0006] On the one hand, the present invention provides a voltage control strategy modeling method based on a distribution network, including:
[0007] Dividing the distribution network into multiple regions based on the electrical distance between nodes of the distribution network, with one region corresponding to one virtual power plant;
[0008] Determining the regional electricity price of each region based on the allocated locational marginal pricing;
[0009] Regarding each region as an object for energy trading with the transmission network, and constructing an optimization model with the maximization of the profit of each region as the goal on the basis of considering the safety of circuit elements in each region;
[0010] Optimizing the strategy modeling of the optimization model based on the observable Markov decision process to obtain an optimized strategy model;
[0011] Training the optimized strategy model based on the regional electricity price of each region and the maximum entropy reinforcement learning algorithm, and taking the action network of the trained optimized strategy model as the voltage control strategy model.
[0012] Optionally, partitioning the distribution network into multiple regions based on the electrical distances between distribution network nodes, including:
[0013] Determining the electrical distance based on the active power-voltage sensitivity and reactive power-voltage sensitivity between distribution network nodes;
[0014] Performing node clustering with the electrical distance as an index, and determining multiple regions based on the clustering result, where each region includes distributed energy nodes and / or energy storage nodes.
[0015] Optionally, determining the regional electricity price for each region based on the zonal marginal pricing, including:
[0016] For each region, taking the minimum purchase cost of distribution network nodes as the node electricity price objective function;
[0017] Considering the power balance of the distribution network, the power boundary conditions of distributed energy, the voltage boundary conditions of network nodes, the energy storage capacity of energy storage nodes and the boundary conditions of energy storage charging and discharging, the boundary conditions of branch power, and the substation power boundary conditions, to determine the electricity price constraint conditions;
[0018] Solving the non-convex optimization problem composed of the node electricity price objective function and the electricity price constraint conditions to obtain the node electricity price;
[0019] Based on the node electricity price, determining the regional electricity price of the corresponding region.
[0020] Optionally, the objective function expression of the optimization model is:
[0021]
[0022]
[0023] In the formula, and are respectively the voltage violation penalty and the regional electricity purchase / sale revenue at time t; represents the reactive power output by the inverter-based distributed energy at node i at time t, represents the adjustable input / output active power of distributed energy and energy storage nodes at node i at time t; T represents the optimization period, and k1 and k2 are respectively the voltage penalty target coefficient and the profit target coefficient; is the penalty factor for voltage violation at time t; is the profit coefficient of energy trading at time t; ReLU is the activation function; v i,t is the voltage phasor on the bus at node i at time t, v i,t is the amplitude of; v max , v minare the maximum voltage and minimum voltage of the node respectively; k is the number of divided distribution network regions; represents the active power and reactive power of region k having sufficient power at time t and selling it to the power grid, and its value is less than 0; represents the active power and reactive power of region k having insufficient power at time t and purchasing electricity from the power grid, and its value is greater than 0; P buy,t,k and Q buy,t,k represent the regional electricity price of region k for purchasing electricity at time t; P sell,t,k and Q sell,t,k represent the selling price of active power and the selling price of reactive power of region k at time t respectively; Ω region is the set of the number of regions; Ω k is the set of nodes of region k;
[0024] The constraint conditions of the optimization model include branch power constraints, network node voltage constraints, power constraints of distributed energy nodes, capacity and charge-discharge constraints of energy storage nodes, and reactive power capacity constraints of inverters of distributed energy nodes;
[0025] The branch power constraints include branch active power constraints and branch reactive power constraints;
[0026] The branch active power constraint is:
[0027]
[0028] In the formula, is the active power of the inverter-based distributed energy installed on node i at time t, is the active power of the load installed on node i at time t, is the active power of the energy storage installed on node i at time t; e i,t is the real part of the bus voltage phasor v i,t at time t; G ij,t and B ij,t represent the real part and the imaginary part of the admittance of branch ij at time t respectively; e j,t and f j,t represent the real part and the imaginary part of the bus voltage phasor v j,t at time t respectively;
[0029] The branch reactive power constraint is:
[0030]
[0031] In the formula, is the reactive power of the load installed on node i at time t, is the reactive power of the energy storage installed on node i at time t, f i,tThe real part of the bus voltage phasor v at time t i,t ;
[0032] The voltage constraint of the network node is:
[0033]
[0034] The power constraint of the distributed energy node is:
[0035]
[0036] Wherein, represents the maximum active power of the inverter-based distributed energy installed on node i at time t, represents the minimum reactive power of the inverter-based distributed energy installed on node i at time t, represents the maximum reactive power of the inverter-based distributed energy installed on node i at time t;
[0037] The capacity and charge-discharge constraints of the energy storage node are:
[0038]
[0039] Wherein, is the state of charge of the energy storage system installed on node i at time t + 1, is the state of charge of the energy storage system installed on node i at time t; represents the output power of the energy storage system installed on node i at time t, represents the charging power of the energy storage system, represents the discharging power of the energy storage system; represents the maximum charging power of the energy storage system installed on node i at time t, represents the maximum discharging power of the energy storage system installed on node i at time t; η charge and η discharge are the charging rate and discharging rate respectively; Δt is the time interval; is the electrical energy capacity of the energy storage system installed on node i; respectively represent the upper and lower limits of the state of charge of the energy storage system installed on node i at time t;
[0040] The reactive power capacity constraint of the inverter of the distributed energy node is:
[0041]
[0042] In the formula, are the active power and reactive power of the photovoltaic power generation on node i at time t respectively; are the apparent powers of photovoltaic power generation and wind power generation at node i at time t, respectively; are the active power and reactive power of wind power generation at node i at time t, respectively.
[0043] Optionally, optimizing the policy modeling of the optimization model based on the observable Markov decision process includes:
[0044] Regarding one of the regions as an agent, and taking the node voltage, the power of the distributed energy node, and the state of charge of the energy storage node in each region as the observation values of the corresponding agent;
[0045] Taking the output power of the distributed energy node and the output power of the energy storage node in each region as the actions of the corresponding agent;
[0046] Determining the constraint conditions of the corresponding agent based on the power output boundary of the distributed energy node in each region;
[0047] Determining the reward of each agent based on the profit of each region and the penalty for the violation of the distribution network voltage limit, where the agents share the reward.
[0048] Optionally, training the optimization policy model based on the regional electricity price of each region and the maximum entropy reinforcement learning algorithm includes:
[0049] For each region, through power flow calculation, the interaction between the agent and the environment is carried out to obtain training samples, and the training samples are put into the experience replay pool; the training samples include the current state, the current action, the current reward, and the state at the next moment;
[0050] Sampling samples from the experience replay pool according to a predetermined priority to obtain target samples;
[0051] Training the optimization policy model using the target samples;
[0052] With the goal of maximizing the reward expectation and maximizing the entropy in each state, the gradient descent method is used to update the parameters of the optimization policy model until the model converges, and a trained optimization policy model is obtained.
[0053] Optionally, the optimization policy model includes an action network and an evaluation network. Training the optimization policy model using the target samples includes:
[0054] Based on the current state of the target sample and the action network, selecting the current optimal action of the agent;
[0055] Determine the current evaluation state value corresponding to the current optimal action based on the current optimal action, the current state, and the evaluation network;
[0056] Update the parameters of the action network and the evaluation network based on the current evaluation state value, the reward value, and the state reward value at the next moment.
[0057] Optionally, the updating of the parameters of the action network and the evaluation network based on the current evaluation state value, the reward value, and the state reward value at the next moment is achieved through the following formula:
[0058] The policy update formula of the optimization policy model is:
[0059]
[0060] In the formula, π new is the updated policy, Z π (o t ) represents the partition function of the policy with respect to the state o t , which is used to normalize the numerator into a probability distribution, o t represents the state at time t; D KL is the divergence that measures the matching degree of two probability distributions; Π is the set of policies; π′ represents any policy selected from Π; α is the temperature parameter, Q θ (o t ,·) is the evaluation state value at state o t ;
[0061] The parameter update formula of the evaluation network is:
[0062]
[0063]
[0064] In the formula, J Q (θ) is the parameter update function of the evaluation network Q, o t is the state at time t, a t is the action at time t, θ is the parameter of the evaluation network, represents the experience replay pool, is to take the expectation when the state and action belong to the experience replay pool, Q θ (o t ,a t ) is the evaluation state value when the state is o t and the action is a t , is Q θ (o t ,a t)'s target value, r(o t ,a t ) is the reward when the state is o t and the action is a t . γ is the future reward factor, and its value range is (0, 1); o t+1 is the state at time t + 1, is the expectation when the state at time t + 1 belongs to ρ π , ρ π represents the state-action or state marginal of the trajectory distribution induced by the policy , is the action taken according to the policy when the state is o t , t , are the parameters of the action network, is the soft state value function when the state is o t+1 and the evaluation network parameters are ; are the evaluation network parameters at the next moment after taking the action; is the action taken according to the policy when the state is o t+1 and the action is a t+1 ; is the updated gradient with respect to θ, is to find the gradient with respect to θ, is the evaluation state value when the state is o t+1 , the action is a t+1 and the evaluation network parameters are ;
[0065] Update the policy parameters by minimizing the expected KL divergence. The parameter update formula of the action network is:
[0066]
[0067] In the formula, is the parameter update function of the action network, represents finding the expectation when the state belongs to the experience replay pool and the action belongs to the policy π, are the action network parameters corresponding to the policy, Q θ is the evaluation state value corresponding to the evaluation network parameters θ, represents that the input noise vector at time t is ∈ t , the input state is o t and the action output by the action network, which is equivalent to a t , obtained from the neural network in the action network; ∈ t is the input noise vector at time t; is the updated with respect to The gradient of Find the gradient with respect to The gradient of To find the gradient with respect to a t The gradient;
[0068] Considering automatic entropy adjustment, find the gradient of α by minimizing the update function J(α) of the temperature parameter α, and update the temperature parameter α. The specific formula is:
[0069]
[0070] In the formula, J(α) represents the update function of the temperature parameter α, For action a t Belongs to the policy When finding the expectation, Is the updated gradient with respect to the temperature parameter α, Is the minimum entropy constant.
[0071] On the other hand, the present invention also provides an online voltage regulation method for a distribution network. The method includes:
[0072] Obtain the real-time observed state values of each area in the distribution network; the real-time observed state values include the real-time network node voltage, the active power and reactive power of inverter-based distributed energy, and the state of charge of the energy storage system;
[0073] Based on the real-time observed state values and the voltage control strategy model obtained in any one of the foregoing, determine the current optimal strategy for each area;
[0074] Use the current optimal strategy to online control and adjust the voltage of network nodes in each area;
[0075] Among them, the optimal strategy corresponds to the optimal action of the voltage control strategy model.
[0076] On the other hand, the present invention also provides a voltage control strategy modeling system for a distribution network, including:
[0077] An area division module, configured to divide the distribution network based on the electrical distance between distribution network nodes to obtain a plurality of areas, and one area corresponds to one virtual power plant;
[0078] A electricity price determination module, configured to determine the area electricity price of each area based on the allocated locational marginal pricing;
[0079] An optimization model construction module, configured to use each area as an object for energy trading with the transmission network, and construct an optimization model with the goal of maximizing the profit of each area on the basis of considering the safety of circuit elements in each area;
[0080] A policy model construction module, which is used to perform optimization policy modeling on the optimization model based on the observable Markov decision process to obtain an optimization policy model;
[0081] A training module, which is used to train the optimization policy model based on the regional electricity price of each region and the maximum entropy reinforcement learning algorithm, and use the action network of the trained optimization policy model as the voltage control policy model.
[0082] On the other hand, the present invention also provides a voltage on-line regulation system for a distribution network, and the system includes:
[0083] An observation value acquisition module, which is used to acquire the real-time observation state values of each region in the distribution network; the real-time observation state values include the real-time network node voltage, the active power and reactive power of the inverter-based distributed energy, and the state of charge of the energy storage system;
[0084] An optimal policy determination module, which is used to determine the current optimal policy of each region based on the real-time observation state value and the voltage control policy model obtained by any one of the foregoing methods;
[0085] A regulation module, which is used to use the current optimal policy to on-line control and adjust the voltage of the network nodes in each region;
[0086] Wherein, the optimal policy corresponds to the optimal action of the voltage control policy model.
[0087] On the other hand, the present invention also provides an electronic device, including: at least one processor and a memory; the memory and the processor are connected by a bus;
[0088] The memory is used to store one or more programs;
[0089] When the one or more programs are executed by the at least one processor, the methods described in any one of the above are implemented.
[0090] On the other hand, the present invention also provides a readable storage medium, on which an execution program is stored, and when the execution program is executed, the methods described in any one of the above are implemented.
[0091] Compared with the prior art, the beneficial effects of the present invention are:
[0092] The present invention provides a voltage control strategy modeling and voltage online regulation method and system for a distribution network. On the one hand, by introducing zonal marginal pricing and regional division of the distribution network, the optimal operating state of each region (VPP) in the distribution network is dynamically captured; for the collaborative optimization problem of multiple virtual power plants based on interactive energy, it is transformed into an optimization model with the maximization of the profit of each region as the goal on the basis of considering the safety of circuit elements in each region, which can optimize the voltage control of the distribution network and the operating profit of each region at the same time. On the other hand, by optimizing the strategy modeling of the optimization model based on the observable Markov decision process and training the model based on the maximum entropy reinforcement learning algorithm, the problems of large computational difficulty and low computational efficiency are avoided, and the economy of each region (VPP) and the distribution network is improved on the basis of ensuring that the distribution network voltage operates within a safe range. BRIEF DESCRIPTION OF THE DRAWINGS
[0093] Figure 1 is a schematic flowchart of a voltage control strategy modeling method for a distribution network according to the present invention;
[0094] Figure 2 is a schematic framework diagram of centralized training of an optimization model for a distribution network according to the present invention;
[0095] Figure 3 is a schematic framework diagram of a voltage regulation collaborative optimization strategy for a distribution network according to the present invention;
[0096] Figure 4 is a schematic structural diagram of an electronic device according to the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0097] The following further elaborates on the specific embodiments of the present invention with reference to the accompanying drawings.
[0098] Embodiment 1
[0099] A voltage control strategy modeling method for a distribution network provided by the present invention, as Figure 1 shown, includes the following steps:
[0100] Step S110: Divide the distribution network into multiple regions based on the electrical distance between nodes of the distribution network, and one region corresponds to one virtual power plant;
[0101] Step S120: Determine the zonal electricity price of each region based on the zonal marginal pricing;
[0102] Step S130: Take each region as an object for energy trading with the transmission network, and construct an optimization model with the maximization of the profit of each region as the goal on the basis of considering the safety of circuit elements in each region;
[0103] Step S140, perform optimization strategy modeling on the optimization model based on the observable Markov decision process to obtain an optimization strategy model;
[0104] Step S150, train the optimization strategy model based on the regional electricity price and the maximum entropy reinforcement learning algorithm for each region, and use the action network of the trained optimization strategy model as the voltage control strategy model.
[0105] In this exemplary embodiment, the distribution network is an active distribution network, which may include distributed energy resources (DERs) with inverters and energy storage nodes. The distributed energy resources may include intermittent and fluctuating energy sources such as photovoltaic power generation (PV) and wind power generation (WT). The electrical distance is defined by the voltage magnitude sensitivity, and one region corresponds to one virtual power plant. The zonal marginal pricing can be used as a price signal to reflect the grid operation demand and can be used to incentivize rational end-users to contribute to the optimal operation of the grid. Each region is regarded as an object for energy trading with the transmission grid, and an optimization model is constructed with the goal of maximizing the profit of each region on the basis of considering the safety of circuit elements in each region. By transforming the multi-VPP collaborative optimization problem based on TE into a nonlinear programming model, the purpose of simultaneously optimizing the voltage control of the distribution network and the operating profit of the VPP is achieved. By performing optimization strategy modeling on the optimization model based on the observable Markov decision process and training the model based on the maximum entropy reinforcement learning algorithm, the problems of high computational difficulty and low computational efficiency are avoided, and the economy of each region (VPP) and the distribution network is improved on the basis of ensuring that the distribution network voltage operates within a safe range.
[0106] In an exemplary embodiment, the process of dividing the distribution network into multiple regions based on the electrical distance between the nodes of the distribution network in step S110 includes the following:
[0107] Determine the electrical distance based on the active power voltage sensitivity and the reactive power voltage sensitivity between the nodes of the distribution network;
[0108] Perform node clustering with the electrical distance as an index, and determine multiple regions based on the clustering result. Each region includes distributed energy nodes and / or energy storage nodes.
[0109] In this exemplary embodiment, the regional division is also the zoning of the distribution network. Specifically, from the perspective of voltage control, the electrical distance is defined by the voltage magnitude sensitivity, and the electrical distance is determined based on the active / reactive voltage sensitivity, as shown in the following formulas (1)-(3).
[0110]
[0111] In the formula, e ijdenotes the electrical distance between nodes i and j, and respectively denote the electrical distance based on active power voltage sensitivity and the electrical distance based on reactive power voltage sensitivity between nodes i and j; the parameter τ can be calculated according to the most severe overvoltage condition, τ represents the voltage breakdown recovery ratio caused by active power reduction, and (1 - τ) represents the voltage breakdown recovery ratio caused by reactive power compensation. and respectively denote the sensitivities of the active voltage amplitude and reactive voltage amplitude of node i to the injected voltage at node j, and respectively denote the sensitivities of the active voltage amplitude and reactive voltage amplitude of node j to the injected voltage at node i, and respectively denote the sensitivities of the active voltage amplitude and reactive voltage amplitude of node i to the injected voltage, and respectively denote the sensitivities of the active voltage amplitude and reactive voltage amplitude of node j to the injected voltage. Clustering can be performed with e ij as the clustering index.
[0112] After obtaining multiple clustering regions, the performance of each region can also be calculated to understand the voltage regulation ability of each region. Specifically, the voltage regulation ability is defined as follows:
[0113]
[0114] In the formula, denotes the voltage regulation ability of region k, D Pk is the active voltage regulation ability of distributed energy and energy storage nodes in region k, D Qk is the reactive voltage regulation ability of the inverter in region k.
[0115] The clustering performance index is determined by the regional voltage regulation ability and the electrical distance.
[0116]
[0117] Among them, Q represents the clustering cluster performance index, A ij represents the influence performance between nodes i and j. m is the total weight of all nodes, k i = ∑ j A ij is the weight of node i, k j is the weight of node j; δ(i, j) represents the clustering indicator function of nodes i and j, δ(u, j) = 0 indicates that nodes i and j are in the same clustering cluster, otherwise δ(i, j) = 1. N represents the set of buses in the distribution network except for idle buses, e ghRepresents the electrical distance between nodes g and h, n c is the number of clustering clusters.
[0118] In some exemplary embodiments, determining the regional electricity price for each of the regions based on the allocated locational marginal pricing in step S120 includes the following steps:
[0119] For each region, the purchase cost of the distribution network nodes is minimized as the node electricity price objective function;
[0120] Considering the power balance of the distribution network, the power boundary conditions of distributed energy, the voltage boundary conditions of network nodes, the energy storage capacity of energy storage nodes and the boundary conditions of energy storage charging and discharging, the boundary conditions of branch power, and the substation power boundary conditions, determine the electricity price constraint conditions;
[0121] Solve the non-convex optimization problem composed of the node electricity price objective function and the electricity price constraint conditions to obtain the node electricity price;
[0122] Based on the node electricity price, determine the regional electricity price for the corresponding region.
[0123] In this exemplary embodiment, considering the operation of the distribution network, the allocated locational marginal pricing DLMP is used to calculate the settlement price of each bus in the distribution market transaction. As a price signal, DLMP effectively responds to the operation requirements of the distribution network, thereby motivating users (such as various end-users) to contribute to the optimal operation of the distribution network. The expression of the node electricity price objective function is as follows:
[0124]
[0125] where T is the optimization period. If it is optimized once every hour, then T = 24. c t and d t represent the price of purchasing active power and the price of purchasing reactive power from the transmission network; P i,t and Q i,t respectively represent the active power and reactive power purchased by node i at time t, t represents the optimization time, and n represents the number of network nodes.
[0126]
[0127]
[0128] The above formulas (7a)-(7j) are the electricity price constraint conditions, w ij,t represents the square of the branch current I at time t ij,t of u i,t is the square of the voltage V of node i at time t, Ω i,t k is the set of all nodes in area k; (7b) and (7c) are the active power balance equation and reactive power balance equation of the distribution network respectively, and pr(i) and cr(i) represent the set of parent nodes and the set of child nodes of node i respectively, respectively represent the active power of the distributed energy resource DER, load, and energy storage installed at node i at time t, respectively represent the reactive power of the distributed energy resource DER, load, and energy storage installed at node i at time t, r gi +jx gi is the impedance of line gi, r gi is the resistance of line gi, x gi is the reactance of line gi; are the active power and reactive power of branch gi at time t respectively, w gi,t is the square of the current of branch gi at time t; (7d) and (7e) are the active power constraint and reactive power constraint of renewable energy (photovoltaic and wind power) respectively, and respectively represent the maximum output active power and maximum output reactive power of the distributed energy resource DER at node i at time t, represents the minimum output reactive power of the distributed energy resource DER at node i at time t; (7f) is the voltage constraint condition of node i, and are the upper and lower limits of the square of the voltage of node i respectively, used to constrain u i,t ; (7g), (7h), and (7i) are the constraint conditions of branch power respectively, and are the active power, reactive power, and maximum capacity of branch ij at time t respectively, is the maximum capacity of the substation; (7j) is the active / reactive power of the substation, P t,k ,Q t,k respectively represent the injected active power and injected reactive power into area k at time t, respectively represent the injected active power and injected reactive power into all areas at time t, and K is the number of areas.
[0129] It can be seen that equation (7g) is a non-convex constraint. Equation (7g) is relaxed to equation (8), and equation (8) is transformed into the standard second-order cone form of equation (9).
[0130]
[0131] Finally, the YALMIP toolbox and Gurobi solver are used to solve the above problems. The active power price and reactive power price of node i are calculated by solving the dual factors of formulas (7b) and (7c). The regional electricity price is the average electricity purchase price of each region obtained by calculating the electricity prices of each node according to DLMP.
[0132] In some exemplary embodiments, in step S130, each of the regions is taken as an object for energy trading with the transmission grid, and an optimization model is constructed with the goal of maximizing the profit of each region on the basis of considering the safety of circuit elements in each region, which can be achieved through the following process:
[0133] The goal of the optimization model is to fully explore potential economic benefits on the basis of regulating the voltage of the distribution network. Therefore, in order to ensure the safe operation of the distribution network voltage within a safe range and improve economic benefits, a voltage regulation framework considering interactive energy is established. The optimization model mainly considers PV, WT with inverters, and energy storage nodes (ES), and schedules and controls the reactive power of the inverters of distributed energy and the active power of the energy storage node ES at a certain time interval (such as hourly). It can be considered from two aspects: First, according to the division of the distribution network regions, the control devices in each region prevent voltage over-limit. Second, this region is regarded as an object for energy trading with the transmission grid, and this energy can be charged / discharged or traded with the transmission grid to achieve the maximization of economic profit. The first part of the objective function is to prevent voltage violations, and the second part fully considers the potential benefits of voltage regulation under interactive energy. The optimization problem is expressed as follows:
[0134]
[0135] Equation (10) represents the optimization goal of the distribution network, and equations (10a) and (10b) are the voltage over-limit penalty and the regional electricity purchase / sale revenue respectively; in the formula, and are the voltage over-limit penalty and the regional electricity purchase / sale revenue at time t respectively; represents the reactive power output by the inverter-based distributed energy on node i at time t, represents the adjustable input / output active power of the distributed energy and the energy storage node on node i at time t; T represents the optimization period, and k1 and k2 are the voltage penalty target coefficient and the profit target coefficient respectively; is the penalty factor for voltage violation at time t; is the profit coefficient of energy trading at time t; ReLU is the activation function, and the function definition of ReLU is ReLU(x) = max(0, x), where x is the independent variable of the function; v i,t is the voltage phasor v i,t on the bus at node i at time t; max, v min They are the maximum voltage and minimum voltage of the node respectively; k is the number of divided distribution network regions; It represents the active power and reactive power of region k having sufficient power at time t and selling it to the power grid, and its value is less than 0; It represents the active power and reactive power of region k having insufficient power at time t and purchasing power from the power grid, and its value is greater than 0; P buy,t,k and Q buy,t,k It represents the regional electricity price for region k to purchase electricity at time t; P sell,t,k and Q sell,t,k They represent the active power selling price and reactive power selling price of region k at time t respectively, usually 1 / 10 of the power purchase price; Ω region is the set of the number of regions; Ω k is the set of nodes of region k;
[0136]
[0137] Among them, Equation (11) and Equation (12) are the active power constraint condition and reactive power constraint condition of the branch respectively; is the active power of the inverter-based distributed energy installed on node o at time t, is the active power of the load installed on node i at time t, is the active power of the energy storage installed on node i at time t; e i,t is the real part of the bus voltage phasor v i,t at time t; G ij,t and B ij,t represent the real part and imaginary part of the admittance of branch ij at time t respectively; e j,t and f j,t represent the real part and imaginary part of the bus voltage phasor v j,t at time t respectively; is the reactive power of the load installed on node i at time t, is the reactive power of the energy storage installed on node i at time t, f i,t is the real part of the bus voltage phasor v i,t at time t.
[0138]
[0139] Among them, Equation (13) is the network node voltage constraint, and Equation (14) and (15) are the active power constraint and reactive power constraint of distributed energy (such as photovoltaic and wind power). It represents the maximum active power of the inverter-based distributed energy installed on node i at time t, It represents the minimum reactive power of the inverter-based distributed energy installed on node i at time t, Denotes the maximum reactive power of the inverter-based distributed energy installed at node i at time t.
[0140] To improve energy efficiency, an energy storage system / node ESs is introduced into the model:
[0141]
[0142]
[0143] Equations (16)-(18) are the constraints for the energy storage capacity and the charging and discharging of the energy storage, where is the state of charge of the energy storage system installed at node i at time t + 1, is the state of charge of the energy storage system installed at node i at time t; Denotes the output power of the energy storage system installed at node i at time t, Denotes the charging power of the energy storage system, Denotes the discharging power of the energy storage system; Denotes the maximum charging power of the energy storage system installed at node i at time t, Denotes the maximum discharging power of the energy storage system installed at node i at time t; η charge and η discharge Are the charging rate and discharging rate respectively; Δt is the time interval; Is the electrical energy capacity of the energy storage system installed at node i; Denote the upper and lower limits of the state of charge of the energy storage system installed at node i respectively.
[0144] The reactive power capacity constraint of the inverter of the distributed energy node is:
[0145]
[0146] Equations (19)-(20) give the reactive power capacity of the inverter-based photovoltaic power generation PV and wind power generation WT; Are the active power and reactive power of the photovoltaic power generation at node i at time t respectively; Are the apparent powers of the photovoltaic power generation and wind power generation at node i at time t respectively; Are the active power and reactive power of the wind power generation at node i at time t respectively.
[0147] The collaborative optimization of multiple VPPs in a distribution network can be formulated as a non-convex non-linear programming model. Generally, the solution algorithms mainly include approximation algorithms or heuristic algorithms. In recent years, reinforcement learning, as a data-driven method, has been widely studied. Single-Agent Deep Reinforcement Learning (SADRL) and Multi-Agent Deep Reinforcement Learning (MADRL) are two mainstream reinforcement learning methods. MADRL has solved many complex problems due to its high scalability and low dependence on data and communication. Some studies have proposed an RL voltage control framework based on the Deep Deterministic Policy Gradient algorithm, which can achieve offline centralized training and online execution. Some studies have applied the algorithm based on Deep Q-Learning to unbalanced distribution systems to reduce power losses and avoid voltage violations. There are also studies that have proposed a multi-agent attention critic algorithm based on MADRL to solve the point-to-point energy trading problem of large-scale prosumers, demonstrating the performance of the MADRL algorithm. However, the existing algorithms cannot be directly used for the cooperative optimization of VPP and TE. Therefore, in some other examples, the optimization strategy modeling of the optimization model described in step S140 based on the partially observable Markov decision process includes the following process:
[0148] Regarding a region as an agent, the node voltage, the power of distributed energy nodes, and the state of charge of energy storage nodes within each region are used as the observations of the corresponding agent;
[0149] The output power of distributed energy nodes and the output power of energy storage nodes in each region are used as the actions of the corresponding agent;
[0150] Based on the power output boundaries of distributed energy nodes within each region, the constraint conditions of the corresponding agent are determined;
[0151] Based on the profit of each region and the penalty for voltage violation in the distribution network, the reward of each agent is determined, where each agent shares the reward.
[0152] In the embodiment of this example, the mathematical problem corresponding to the optimization model in step S130 is modeled using the Partially Observable Markov Decision Process (POMDP). The important elements of this model include the action space, the observation space, and the reward function, which are specifically described as follows.
[0153] 1) Observation space: One region corresponds to one agent, and the observation of agent k is expressed as:
[0154]
[0155] Among them, o k,t is the observation value of agent k at time t. The distributed energy RES includes photovoltaic power generation PV and wind power generation WT; i ∈ Ω k represents node i belonging to area k.
[0156] 2) Action space: The action of area agent k at time step t is a k,t , which corresponds to the output power of the distributed energy RES and energy storage ES that area k can control.
[0157]
[0158] 3) Constraints: The action space and observation space must satisfy the output constraints of distributed energy and energy storage constraints.
[0159]
[0160] Equation (23) represents the active output boundary of WT and PV, is the maximum active output of PV at node i at time t, is the maximum active output of WT at node i at time t. Equation (24) represents the reactive output boundary of WT and PV; δ1 and δ2 are the maximum curtailment factors of PV and WT respectively.
[0161]
[0162]
[0163] Equation (25) represents the constraint conditions of the energy storage node ESS. is the partial action value directly output by the agent, represents the maximum active power of the energy storage system installed on node i at time t, is the state of charge of the energy storage system installed on node i at time t-1.
[0164] 4) Reward function: The above optimization problem constructs a POMDP for multiple agents, and each agent can share the reward. The reward of each agent is expressed as:
[0165]
[0166] Equation (27) includes two parts: the profit of the distribution network economy and the penalty for the distribution network voltage exceeding the limit.
[0167] In some exemplary embodiments, training the optimization policy model based on the regional electricity price and the maximum entropy reinforcement learning algorithm for each of the regions in step S150 includes the following processes:
[0168] For each of the regions, the interaction between the agent and the environment is performed through power flow calculation to obtain training samples, and the training samples are placed in an experience replay pool; the training samples include the current state, the current action, the current reward, and the state at the next moment;
[0169] Sample samples from the experience replay pool according to a predetermined priority to obtain target samples;
[0170] Train the optimization policy model using the target samples;
[0171] With the goal of maximizing the reward expectation and maximizing the entropy in each state, the parameter update of the optimization policy model is performed using the gradient descent method until the model converges to obtain a trained optimization policy model.
[0172] In this exemplary embodiment, the voltage control policy model is the action network in the trained optimization policy model. With the goal of maximizing the reward expectation and maximizing the entropy in each state, it can be implemented based on the Soft Actor-Critic (SAC) algorithm of reinforcement learning. The reinforcement learning algorithm in this example searches for the optimal policy to maximize the reward expectation, and promotes the standard goal to achieve the maximum entropy goal by adding an entropy term to the formula, so that another goal of the optimal policy is to maximize its entropy in each state. Specifically, the expression of the objective function for searching for the optimal policy is:
[0173]
[0174] In the formula, is the entropy function, π(·|o t ) is the policy under the state o t ; π is an arbitrary policy, and π * is the optimal policy; is the expectation when the policy of the state o t , action a t belongs to ρ π ; o t represents the state at time t, a t represents the action at time t, and the randomness of the optimal policy is controlled by the temperature parameter α. The temperature parameter α is different at different time periods of the same task training, and the influence of this hyperparameter on the performance is obvious; ρ π represents the state-action or state margin of the trajectory distribution induced by the policy , is the expectation at the state o tMake action a according to the policy at time t t , are the parameters of the action network.
[0175] Express the entropy-regularized Bellman equation as:[[]]
[0176]
[0177] where denotes the entropy-regularized Bellman operation on Q θ (s t , a t ); Q θ (s t , a t ) represents the evaluated state value when the state is s t and the action is a t ; θ are the parameters of the evaluation network; r(s t+1 , a t+1 ) is the reward when the state is s t+1 and the action is a t+1 ; a t+1 is the action at time t + 1, s t+1 is the state at time t + 1; V θ (s t+1 ) is the soft state value function for the state s t+1 ; γ is the factor of future rewards in the range (0, 1); denotes that the state s t+1 belongs to ρ π ; denotes the expectation that the action a t+1 belongs to the policy π; is to make action a t+1 according to the policy at the state s t+1 . Q θ (s t+1 , a t+1 ) represents the evaluated state value when the state is s t+1 and the action is a t+1 .
[0178] Exemplarily, as Figure 2 shown, is the centralized training framework of the algorithm. During the centralized training process in Figure 2 , the agent collects samples by interacting with the environment. For example, the samples include the state (s k,t ), the action (a k,t ), the reward (r t ) and the next state (s k,t+1) The reward is obtained through energy trading and AC power flow calculation, and then the samples are put into the experience replay pool. The action network selects the optimal action of the agent based on the current state of the policy network and the trained parameters. The evaluation networks (Q-value network and target Q-value network) output the evaluation state value according to the action and the current state, that is, evaluate the quality of the action. The update process of the neural network weights is introduced during the training process. First, a batch of samples, that is, target samples, are drawn from the experience replay pool according to the sample priority. The sample priority can be determined according to the sample usage frequency, and then the target samples are used to train the model. Finally, the weights of the action network are updated using the policy gradient method (such as the gradient descent method) and the soft Bellman residual until the training sample set converges to a certain number.
[0179] During the above training process, two parts of the algorithm are improved to parameter sharing and prioritized experience replay. Parameter sharing includes parameter sharing of the evaluation network and parameter sharing of the action network. Parameter sharing shares the parameters for all agents through the centralized training method during the training phase. Prioritized experience replay samples the samples in the experience replay pool according to a certain priority, which is more efficient than randomly and uniformly sampling from the experience replay pool. This example can further improve the reward value and training speed of the model.
[0180] Exemplarily, the optimization policy model includes an action network and an evaluation network. Training the optimization policy model using the target samples includes the following process:
[0181] Based on the current state of the target sample and the action network, select the current optimal action of the agent;
[0182] Based on the current optimal action, the current state, and the evaluation network, determine the current evaluation state value corresponding to the current optimal action;
[0183] Based on the current evaluation state value, the reward value, and the state reward value at the next moment, update the parameters of the action network and the evaluation network.
[0184] In the embodiment of this example, the policy update formula of the optimization policy model is expressed as:
[0185]
[0186] In the formula, π new is the updated policy, Z π (o t ) represents the partition function of the policy with respect to the state o t , which is used to normalize the numerator into a probability distribution, o t represents the state at time t; D KLis the divergence that measures the matching degree of two probability distributions; Π is the set of policies; π′ represents any policy selected from Π; α is the temperature parameter, Q θ (o t ,·) is the evaluation state value at state o t ;
[0187] The parameter update formula of the evaluation network is as follows:
[0188]
[0189] In the formula, J Q (θ) is the parameter update function of the evaluation network Q, o t is the state at time t, a t is the action at time t, θ is the parameter of the evaluation network, represents the experience replay pool, is to take the expectation when the state and action belong to the experience replay pool, Q θ (o t ,a t ) is the evaluation state value when the state is o t and the action is a t ; is the target value of Q θ (o t ,a t ), r(o t ,a t ) is the reward when the state is o t and the action is a t ; γ is the future reward factor, and its value range is (0,1); o t+1 is the state at time t+1, is the expectation when the state at time t+1 belongs to ρ π , ρ π represents the state-action or state marginal of the trajectory distribution induced by the policy ; is to make the action o t at the state o t according to the policy, is the parameter of the action network, is the soft state value function when the state is o t+1 and the parameter of the evaluation network is ; is the parameter of the evaluation network at the next moment after making the action; is the expectation when the action at time t+1 belongs to the policy π, Q θ (o t+1 ,a t+1 ) is the evaluation state value when the state is o t+1 and the action is a t+1The evaluation function value in the case is to take action a according to the policy at state o t+1 ; t+1 ; is the updated gradient with respect to θ is to find the gradient with respect to θ is the evaluation state value when the state is o t+1 , the action is a t+1 , and the evaluation network parameters are ;
[0190] Update the policy parameters by minimizing the expected KL divergence. The parameter update formula of the action network is as follows:
[0191]
[0192] In the formula, is the parameter update function of the action network represents taking the expectation when the state belongs to the experience replay pool and the action belongs to the policy π represents that the input noise vector at time t is ∈ t , the input state is o t , and the action output by the action network in this case is equivalent to a t , which is obtained by the neural network in the action network; ∈ t is the input noise vector at time t; is the updated gradient with respect to ; Find the gradient with respect to ; is to find the gradient with respect to a t ;
[0193] Considering automatic entropy adjustment, find the gradient of α by minimizing the update function J(α) of the temperature parameter α, and update the temperature parameter α. The specific formula is as follows:
[0194]
[0195] In the formula, J(α) represents the update function of the temperature parameter α is to take the expectation when the action a t belongs to the policy ; is the updated gradient with respect to the temperature parameter α is the minimum entropy constant.
[0196] The integration of a large number of distributed energy sources brings new challenges to the operation and control of distribution networks. To effectively control distributed power sources, VPPs are introduced into distributed power source systems. Traditionally, the control strategy of VPPs is an optimization model that does not consider voltage control and TE. The present invention is based on the multi-VPP collaborative optimization problem of TE and transforms it into a non-linear programming model. Its purpose is to optimize the voltage control of the DN and the operating profit of VPPs. To dynamically capture the optimal operating state of VPPs in the distributed network, zonal marginal pricing and distributed network partitioning are introduced into the model. Generally speaking, the calculation of such models is difficult, and approximate algorithms are mainly used for solving, with low efficiency. Recently, MADRL has emerged as a scalable and promising data-driven method. We propose an enhanced MADRL solution to solve this problem. Through the above process, centralized training of the optimization strategy model is completed. After the model training is completed, the evaluation network (Q-value network) can be discarded for online execution.
[0197] Embodiment 2
[0198] The present invention also provides a method for on-line voltage regulation of a distribution network, the method comprising:
[0199] Obtaining real-time observed state values of each area in the distribution network; the real-time observed state values include real-time network node voltages, active and reactive powers of inverter-based distributed energy sources, and the state of charge of energy storage systems;
[0200] Determining the current optimal strategy for each area based on the real-time observed state values and the voltage control strategy model obtained in any one of the foregoing;
[0201] Using the current optimal strategy to on-line control and adjust the voltages of network nodes in each area;
[0202] Wherein, the optimal strategy corresponds to the optimal action of the voltage control strategy model.
[0203] In the present exemplary embodiment, when making on-line decisions, each agent only needs to observe the real-time observed values of the local area according to the trained action network (i.e., the voltage control strategy model), that is, the real-time network node voltages, the active and reactive powers of inverter-based distributed energy sources, and the state of charge of energy storage systems, obtain the optimal strategy, that is, the best action, based on the voltage control strategy model, and on-line regulate the node voltage based on the best action. Exemplarily, the optimal strategy includes the active power of charge and discharge of energy storage nodes and the reactive power of inverters of distributed energy nodes in each area. For each area, based on the optimal strategy, control and adjust the active power of charge and discharge of energy storage nodes and the reactive power of inverters of distributed energy nodes in the area to adjust the voltages of each network node in the area.
[0204] Exemplarily, as Figure 3 shown, it is the complete framework of the distribution network optimization strategy based on interactive energy proposed by the present invention. This strategy includes two parts: offline centralized training and online distributed execution. First, in the offline centralized training stage, the distribution network area is divided and the area electricity price is calculated. Each agent (agent 1, agent 2,... agent K) controls the charging and discharging of energy storage and the actions of wind power and photovoltaic inverters within one area (area 1, area 2,... area K). Among them, each agent can conduct energy transactions with the transmission network (main grid), determine the execution actions (action 1, action 2,... action K of agent 1, agent 2,... agent K) based on the current state of each agent (state 1, state 2, state of agent 1, agent 2,... agent K), and the actions act on the environment to return the reward state to continue the learning process. Through the above process, each area can fully improve the economic benefits on the basis of regulating the voltage. In the online distributed execution stage, by observing the states of each agent in the actual distribution network, the optimal actions of each agent are obtained, and the actions of energy storage, wind power, and photovoltaic inverters are quickly controlled.
[0205] Facing the new challenges brought by the integration of a large number of distributed energy sources to the operation and control of the distribution network, in order to effectively control distributed power sources, VPP is introduced into the distributed power source system. Traditionally, the control strategy of VPP is an optimization model that does not consider voltage control and interactive energy. The present invention is based on the collaborative optimization problem of multiple VPPs based on interactive energy and transforms it into a non-linear programming model, aiming to optimize the voltage control of the distribution network and the operating profit of VPP. In order to dynamically capture the optimal operating state of VPPs in the distributed network, zonal marginal pricing and distributed network partitioning are introduced into the model. Generally speaking, the calculation difficulty of such models is relatively large, and approximate algorithms are mainly used for solving, with low efficiency. To address this problem, an enhanced MADRL solution is introduced to avoid problems such as large calculation difficulty and low efficiency. This method coordinates PV, WT, and energy storage in each VPP to ensure that the distribution network voltage operates within a safe range and improve the economy of each VPP and the distribution network.
[0206] Embodiment 3
[0207] Based on the same inventive concept, the present invention also provides a voltage control strategy modeling system for a distribution network, including:
[0208] A regional division module for dividing the distribution network based on the electrical distance between distribution network nodes to obtain multiple regions, and one region corresponds to one virtual power plant;
[0209] An electricity price determination module for determining the regional electricity price of each region based on zonal marginal pricing;
[0210] The optimization model construction module is used to take each of the regions as an object for energy trading with the power grid, and construct an optimization model with the goal of maximizing the profit of each region on the basis of considering the safety of circuit elements in each region;
[0211] The policy model construction module is used to perform optimization policy modeling on the optimization model based on the observable Markov decision process to obtain an optimization policy model;
[0212] The training module is used to train the optimization policy model based on the regional electricity price of each region and the maximum entropy reinforcement learning algorithm, and use the action network of the trained optimization policy model as the voltage control policy model.
[0213] In a possible implementation manner, the region division module includes:
[0214] The distance determination sub-module is used to determine the electrical distance based on the active power voltage sensitivity and reactive power voltage sensitivity between the distribution network nodes;
[0215] The clustering sub-module is used to perform node clustering with the electrical distance as an index, and determine multiple regions based on the clustering result. Each region includes distributed energy nodes and / or energy storage nodes.
[0216] In a possible implementation manner, the electricity price determination module includes:
[0217] The electricity price target determination sub-module is used to, for each region, take the minimum purchase cost of the distribution network nodes as the node electricity price objective function;
[0218] The electricity price constraint condition determination sub-module is used to determine the electricity price constraint conditions by considering the power balance of the distribution network, the power boundary conditions of distributed energy, the voltage boundary conditions of network nodes, the energy storage capacity of energy storage nodes and the boundary conditions of energy storage charging and discharging, the boundary conditions of branch power, and the substation power boundary conditions;
[0219] The solution sub-module is used to solve the non-convex optimization problem composed of the node electricity price objective function and the electricity price constraint conditions to obtain the node electricity price;
[0220] The regional electricity price determination sub-module is used to determine the regional electricity price corresponding to the region based on the node electricity price.
[0221] In a possible implementation manner, the objective function expression of the optimization model is:
[0222]
[0223] In the formula, and The voltage violation penalty and the regional power purchase / sale revenue at time t, respectively; represents the reactive power output of the inverter-based distributed energy at node i at time t; represents the adjustable active power input / output of the distributed energy and the energy storage node at node i at time t; T represents the optimization period, and k1 and k2 are the voltage penalty target coefficient and the profit target coefficient, respectively; is the penalty factor for voltage violation at time t; is the profit coefficient for energy trading at time t; ReLU is the activation function; v i,t is the voltage phasor on the bus at node i at time t, v i,t is the amplitude of; v max , v min are the maximum voltage and the minimum voltage of the node, respectively; k is the region number; represents the active power and reactive power of region k selling sufficient power to the power grid at time t, and its value is less than 0; represents the active power and reactive power of region k purchasing power from the power grid due to insufficient power at time t, and its value is greater than 0; P buy,t,k and Q buy,t,k represent the regional electricity price for region k purchasing power at time t; P sell,t,k and Q sell,t,k represent the active power selling price and the reactive power selling price of region k at time t, respectively; Ω region is the set of the number of regions; Ω k is the set of nodes in region k;
[0224] The constraint conditions of the optimization model include branch power constraints, network node voltage constraints, power constraints of distributed energy nodes, capacity and charge / discharge constraints of energy storage nodes, and reactive power capacity constraints of inverters of distributed energy nodes;
[0225] The branch power constraints include branch active power constraints and branch reactive power constraints;
[0226] The branch active power constraint is:
[0227]
[0228] In the formula, is the active power of the inverter-based distributed energy installed at node i at time t, is the active power of the load installed at node i at time t, is the active power of the energy storage installed at node i at time t; e i,t is the real part of the voltage phasor v on the bus at time t i,t ; G ij,t and B ij,trespectively represent the real and imaginary parts of the admittance of branch ij at time t; e j,t and f j,t respectively represent the real and imaginary parts of the bus voltage phasor v j,t at time t;
[0229] The reactive power constraint of the branch is:
[0230]
[0231] In the formula, is the reactive power of the load installed at node i at time t, is the reactive power of the energy storage installed at node i at time t, f ,t is the real part of the bus voltage phasor v i,t at time t;
[0232] The network node voltage constraint is:
[0233]
[0234] The power constraint of the distributed energy node is:
[0235]
[0236] Among them, represents the maximum active power of the inverter-based distributed energy installed at node i at time t, represents the minimum reactive power of the inverter-based distributed energy installed at node i at time t, represents the maximum reactive power of the inverter-based distributed energy installed at node i at time t;
[0237] The capacity and charge-discharge constraints of the energy storage node are:
[0238]
[0239]
[0240] Among them, is the state of charge of the energy storage system installed at node i at time t + 1, is the state of charge of the energy storage system installed at node i at time t; represents the output power of the energy storage system installed at node i at time t, represents the charging power of the energy storage system, represents the discharging power of the energy storage system; represents the maximum charging power of the energy storage system installed at node i at time t, represents the maximum discharging power of the energy storage system installed at node i at time t; ηcharge and η discharge are the charging rate and the discharging rate respectively; Δt is the time interval; is the electrical energy capacity of the energy storage system installed at node i; respectively represent the upper limit and the lower limit of the state of charge of the energy storage system installed at node i at time t;
[0241] The reactive power capacity constraint of the inverter of the distributed energy node is:
[0242]
[0243] In the formula, are the active power and the reactive power of the photovoltaic power generation at node i at time t respectively; are the apparent powers of the photovoltaic power generation and the wind power generation at node i at time t respectively; are the active power and the reactive power of the wind power generation at node i at time t respectively.
[0244] In a possible implementation manner, the policy model construction module is specifically configured to:
[0245] Take one of the regions as an agent, and take the node voltage, the power of the distributed energy node, and the state of charge of the energy storage node in each region as the observation values of the corresponding agent;
[0246] Take the output power of the distributed energy node and the output power of the energy storage node in each region as the actions of the corresponding agent;
[0247] Determine the constraint conditions of the corresponding agent based on the power output boundary of the distributed energy node in each region;
[0248] Determine the reward of each agent based on the profit of each region and the penalty for the violation of the distribution network voltage limit, wherein the agents share the reward.
[0249] In a possible implementation manner, the training module includes:
[0250] A sample generation sub-module, which is used for each region to perform the interaction between the agent and the environment through power flow calculation to obtain training samples, and put the training samples into an experience replay pool; the training samples include the current state, the current action, the current reward, and the state at the next moment;
[0251] A sampling sub-module, which is used to sample samples from the experience replay pool according to a predetermined priority to obtain target samples;
[0252] A training sub-module, configured to train the optimization policy model by using the target samples;
[0253] A parameter update sub-module, configured to update the parameters of the optimization policy model by using the gradient descent method with the objectives of maximizing the expected reward and maximizing the entropy under each state until the model converges, thereby obtaining the voltage control policy model.
[0254] In a possible implementation manner, the optimization policy model includes an action network and an evaluation network, and the training sub-module includes:
[0255] An action selection sub-unit, configured to select the current optimal action of the agent based on the current state of the target sample and the action network;
[0256] An evaluation sub-unit, configured to determine the current evaluation state value corresponding to the current optimal action based on the current optimal action, the current state, and the evaluation network;
[0257] A parameter update sub-unit, configured to update the parameters of the action network and the evaluation network based on the current evaluation state value, the reward value, and the state reward value at the next moment.
[0258] In a possible implementation manner, the parameter update sub-unit is specifically configured to update the parameters of the action network and the evaluation network by using the following formula:
[0259] The policy update formula of the optimization policy model is:
[0260]
[0261] In the formula, π new is the updated policy, Z π (o t ) represents the partition function of the policy with respect to the state o t , which is used to normalize the numerator into a probability distribution, o t represents the state at time t; D KL is the divergence that measures the matching degree between two probability distributions; Π is the set of policies; π′ represents any policy selected from Π; α is the temperature parameter, and Q θ (o t , ·) is the evaluation state value at the state o t ;
[0262] The parameter update formula of the evaluation network is:
[0263]
[0264] In the formula, J Q (θ) is the parameter update function of the evaluation network Q, ot is the state at time t, a t is the action at time t, θ is the parameter of the evaluation network, represents the experience replay pool, is to take the expectation when the state and action belong to the experience replay pool, Q θ (o t , a t ) is the evaluation state value when the state is o t and the action is a t . is Q θ (o t , a t )'s target value, r(o t , a t ) is the reward when the state is o t and the action is a t . γ is the future reward factor, and its value range is (0, 1); o t+1 is the state at time t+1, is the expectation when the state at time t+1 belongs to ρ π , ρ π represents the state-action or state margin of the trajectory distribution induced by the policy . is to make the action o t when the state is o t , is the parameter of the action network, is the soft state value function when the state is o t+1 and the parameter of the evaluation network is . is the parameter of the evaluation network at the next moment after making the action; is to make the action a t+1 when the state is o t+1 ; is the updated gradient with respect to θ, is to calculate the gradient with respect to θ, is the evaluation state value when the state is o t+1 and the action is a t+1 and the parameter of the evaluation network is ;
[0265] The policy parameters are updated by minimizing the expected KL divergence. The parameter update formula of the action network is as follows:
[0266]
[0267] In the formula, is the parameter update function of the action network, Denote the expectation when the state belongs to the experience replay pool and the action belongs to the policy π. Denote that the input noise vector at time t is ∈ t and the input state is o t The action output by the action network in this case, which is equivalent to a t , obtained by the neural network in the action network; ∈ t is the input noise vector at time t; is the updated gradient with respect to ; Find the gradient with respect to ; is to find the gradient with respect to a t ;
[0268] Considering automatic entropy adjustment, find the gradient of α by minimizing the update function J(α) of the temperature parameter α, and update the temperature parameter α. The specific formula is:
[0269]
[0270] In the formula, J(α) represents the update function of the temperature parameter α, is the action a t belonging to the policy when finding the expectation, is the updated gradient with respect to the temperature parameter α, is the minimum entropy constant.
[0271] Embodiment 4
[0272] Based on the same inventive concept, the present invention also provides an on-line voltage regulation system for a distribution network. The system includes:
[0273] An observation value acquisition module for acquiring real-time observation state values of each area in the distribution network; the real-time observation state values include real-time network node voltages, active and reactive powers of inverter-based distributed energy, and the state of charge of the energy storage system;
[0274] An optimal strategy determination module for determining the current optimal strategy for each area based on the real-time observation state values and the voltage control strategy model obtained by any one of claims 1-8;
[0275] A regulation module for using the current optimal strategy to on-line control and adjust the voltages of network nodes in each area;
[0276] Among them, the optimal strategy corresponds to the optimal action of the voltage control strategy model.
[0277] In a possible implementation manner, the optimal strategy includes the active power of charge and discharge of energy storage nodes in each of the regions and the reactive power of the inverters of distributed energy nodes. Specifically, the regulation module is configured to:
[0278] For each of the regions, based on the optimal strategy, control and adjust the active power of charge and discharge of the energy storage nodes in the region and the reactive power of the inverters of the distributed energy nodes, so as to adjust the voltages of the network nodes in the region.
[0279] Embodiment 5
[0280] As Figure 4 shown, the present invention further provides an electronic device, which may be a computer device, a single-chip microcomputer device, a smart mobile device, etc. The electronic device in this embodiment may include a processor, a memory, a transceiver component, etc. The memory, the processor, and the transceiver component are connected through a bus; the memory may be used to store an execution program, and an exemplary execution program may include instructions; the processor is used to execute the instructions stored in the memory. The memory may also be used to store data, and the data may be called and / or modified when the instructions are executed.
[0281] The processor may be a Central Processing Unit (CPU), or may also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. It is the computing core and control core of the terminal, and is suitable for implementing one or more instructions. Specifically, it is suitable for loading and executing one or more instructions in the storage medium to implement the corresponding method flow or corresponding function, so as to implement the steps of a method for modeling a voltage control strategy of a distribution network and / or a method for online voltage regulation of a distribution network in the above embodiments.
[0282] Embodiment 6
[0283] Based on the same inventive concept, the present invention also provides a readable storage medium, specifically an electronic device-readable storage medium (Memory). The electronic device-readable storage medium is a memory device in the electronic device, used to store programs and data. It can be understood that the storage medium here can include both the built-in storage medium in the electronic device and, of course, the extended storage medium supported by the electronic device. The storage medium provides a storage space, and this storage space stores the operating system of the terminal. Moreover, in this storage space, there is also stored one or more instructions suitable for being loaded and executed by the processor. These instructions can be one or more executable programs (including program codes). It should be noted that the storage medium here can be a high-speed RAM memory or a non-volatile memory, such as at least one disk memory. By the processor loading and executing one or more instructions stored in the storage medium, the steps of a voltage control strategy modeling method for a distribution network and / or a voltage on-line regulation method for a distribution network in the above embodiments can be realized.
[0284] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) containing computer-usable program codes.
[0285] The present invention is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present invention. It should be understood that each flow and / or block in the flowchart and / or block diagram, and the combination of flows and / or blocks in the flowchart and / or block diagram, can be realized by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for realizing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0286] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device, and the instruction device realizes the functions in Figure 1 one flow or multiple flows and / or blocks Figure 1The functions specified in one or more boxes.
[0287] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process. Thus, the instructions executed on the computer or other programmable device provide for implementing the steps of the functions specified in Figure 1 one process or more processes and / or boxes Figure 1 the functions specified in one box or more boxes.
[0288] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention rather than to limit the scope of its protection. Although the present invention has been described in detail with reference to the above embodiments, those of ordinary skill in the art should understand that after reading the present invention, various changes, modifications or equivalent replacements can still be made to the specific implementation manners of the application. However, these changes, modifications or equivalent replacements are all within the scope of protection of the claims pending for approval of the application.
Claims
1. A voltage control strategy modeling method for a distribution network, characterized in that: include: The distribution network is divided into regions based on the electrical distance between the nodes of the distribution network, and multiple regions are obtained, and each region corresponds to a virtual power plant; determining a regional electricity price for each of said regions based on the allocated locational marginal pricing; Taking each of the regions as an object for energy trading with the transmission grid, building an optimization model with the goal of maximizing the profit of each of the regions on the basis of considering the safety of circuit elements in each of the regions; Based on the observable Markov decision process, the optimization model is optimized and the optimization strategy model is obtained; The optimization strategy model is trained based on the regional electricity price of each of the regions and the maximum entropy reinforcement learning algorithm, and the action network of the trained optimization strategy model is used as the voltage control strategy model.
2. The method according to claim 1, characterized in that The distribution network is divided into regions based on the electrical distances between the nodes of the distribution network to obtain multiple regions, including: Determining the electrical distance based on active power voltage sensitivity and reactive power voltage sensitivity between nodes of the distribution network; Node clustering is performed using the electrical distance as an indicator, and a plurality of regions are determined based on the clustering result, each of the regions including a distributed energy node and / or an energy storage node.
3. The method according to claim 1, characterized in that The determining of the regional electricity price of each of the regions based on the allocation location marginal pricing comprises: For each region, the node electricity price objective function is to minimize the electricity purchase cost of the distribution network node; Consider the power balance of the distribution network, the power boundary conditions of distributed energy, the voltage boundary conditions of network nodes, the energy storage capacity of energy storage nodes and the boundary conditions of energy storage charging and discharging, the boundary conditions of branch power and the boundary conditions of substation power, and determine the constraints of electricity prices; Solving a non-convex optimization problem consisting of the node electricity price objective function and the electricity price constraint condition to obtain the node electricity price; Based on the node electricity price, a regional electricity price of the corresponding area is determined.
4. The method according to claim 3, characterized in that The objective function expression of the optimization model is: In the formula, and They are the voltage over-limit penalty and regional power purchase / sale revenue at time t respectively; represents the reactive power output of the inverter-based distributed energy on node i at time t, represents the adjustable input / output active power of distributed energy and energy storage nodes on node i at time t; T represents the optimization period, k1 and k2 are the voltage penalty target coefficient and profit target coefficient respectively; is the penalty factor for voltage violation at time t; is the profit coefficient of energy trading at time t; ReLU is the activation function; v i,t is the voltage phasor v on the bus at node i at time t i,t The amplitude of v max ,v min are the maximum voltage and minimum voltage of the node respectively; k is the region number; It indicates that region k has sufficient electricity at time t and sells the active power and reactive power to the grid, and its value is less than 0; represents the active power and reactive power of area k when it is short of power and purchases electricity from the power grid at time t, and its value is greater than 0; P buy,t,k and Q buy,t,k represents the regional electricity price of region k at time t; P sell,t,k and Q sell,t,k They represent the active power selling price and reactive power selling price of region k at time t respectively; Ω region is the number of regions; Ω k is the node set of region k; The constraints of the optimization model include branch power constraints, network node voltage constraints, distributed energy node power constraints, energy storage node capacity and charging and discharging constraints, and distributed energy node inverter reactive capacity constraints; The branch power constraint includes branch active power constraint and branch reactive power constraint; The branch active power constraint is: In the formula, is the active power of the inverter-based distributed energy installed on node i at time t, is the active power of the load installed on node i at time t, is the active power of the energy storage installed on node i at time t; e i,t is the voltage phasor v on the bus at time t i,t The real part of G ij,t and B ij,t They represent the real and imaginary parts of the admittance of branch ij at time t respectively; e j,t and f j,t They represent the voltage phasor v on the bus at time t respectively. j,t The real and imaginary parts of The branch reactive power constraint is: In the formula, is the reactive power of the load installed on node i at time t, is the reactive power of the energy storage installed on node i at time t, f i,t is the voltage phasor v on the bus at time t i,t The real part of The network node voltage constraint is: The power constraint of the distributed energy node is: in, represents the maximum active power of the inverter-based distributed energy installed on node i at time t, represents the minimum reactive power of the inverter-based distributed energy installed on node i at time t, represents the maximum reactive power of the inverter-based distributed energy installed on node i at time t; The capacity and charging and discharging constraints of the energy storage node are: in, is the state of charge of the energy storage system installed on node i at time t+1, is the state of charge of the energy storage system installed on node i at time t; represents the output power of the energy storage system installed on node i at time t, represents the charging power of the energy storage system, Indicates the discharge power of the energy storage system; represents the maximum charging power of the energy storage system installed on node i at time t, represents the maximum discharge power of the energy storage system installed on node i at time t; η charge and η discharge are charging rate and discharging rate respectively; Δt is the time interval; is the electric energy capacity of the energy storage system installed on the i-node; They represent the upper and lower limits of the state of charge of the energy storage system installed on node i at time t; The reactive capacity constraint of the inverter of the distributed energy node is: In the formula, are respectively the active power and reactive power of photovoltaic power generation at node i at time t; are the apparent power of photovoltaic power generation and wind power generation at node i at time t respectively; are the active power and reactive power of wind power generation at node i at time t respectively.
5. The method according to claim 1, characterized in that The optimization strategy modeling of the optimization model based on the observable Markov decision process includes: Taking one of the regions as an intelligent agent, taking the node voltage, the power of the distributed energy node and the charge state of the energy storage node in each of the regions as the observation values of the corresponding intelligent agent; The output power of the distributed energy nodes and the output power of the energy storage nodes in each of the regions are used as actions of the corresponding intelligent agents; Determine the constraint condition of the corresponding intelligent agent based on the power output boundary of each distributed energy node in the area; The reward of each of the intelligent agents is determined based on the profit of each of the areas and the penalty for exceeding the voltage limit of the power distribution network, wherein the reward is shared by the intelligent agents.
6. The method according to claim 5, characterized in that The method of training the optimization strategy model based on the regional electricity price of each region and the maximum entropy reinforcement learning algorithm includes: For each of the regions, the interaction between the agent and the environment is performed through power flow calculation to obtain training samples, and the training samples are placed in an experience replay pool; the training samples include a current state, a current action, a current reward, and a state at the next moment; Sampling samples from the experience replay pool according to a predetermined priority to obtain target samples; Using the target sample to train the optimization strategy model; With the goal of maximizing the reward expectation and maximizing the entropy in each state, the gradient descent method is used to update the parameters of the optimization strategy model until the model converges to obtain a trained optimization strategy model.
7. The method according to claim 6, characterized in that The optimization strategy model includes an action network and an evaluation network. The optimization strategy model is trained by using the target sample, including: Based on the current state of the target sample and the action network, selecting the current optimal action of the agent; Determine a current evaluation state value corresponding to the current optimal action based on the current optimal action, the current state and the evaluation network; The parameters of the action network and the evaluation network are updated based on the current evaluation state value, the reward value and the state reward value at the next moment.
8. The method according to claim 7, characterized in that The updating of the parameters of the action network and the evaluation network based on the current evaluation state value, the reward value and the state reward value at the next moment is implemented by the following formula: The strategy update formula of the optimization strategy model is: In the formula, π new is the updated strategy, Z π (o t ) represents the state o t The partition function of the strategy is used to normalize the numerator into a probability distribution, o t represents the state at time t; D KL is the divergence that measures the degree of match between two probability distributions; Π is the set of strategies; π′ represents any strategy selected in Π; α is the temperature parameter, Q θ (o t ,·) is state o t The evaluation state value at time θ is the parameter of the evaluation network; The parameter update formula of the evaluation network is: In the formula, J Q (θ) is the parameter update function of the evaluation network Q, o t is the state at time t, a t is the action at time t, Represents the experience replay pool, Find the expectation when the state and action belong to the experience replay pool, Q θ (o t ,a t ) is state o t 、Action is a t The evaluation status value of the case, Q θ (o t ,a t ), r(o t ,a t ) is state o t 、Action is a t The reward in this case, γ is the future reward factor, and its value range is (0,1); t+1 is the state at time t+1, The state at time t+1 belongs to ρ π The expectation under the condition, ρ π Indicated by strategy The state-action or state-margin of the induced trajectory distribution, For the state o t Take action according to the strategy t , are the parameters of the action network, The state is o t+1 , the evaluation network parameters are The soft state value function in the case of The evaluation network parameters for the next moment after the action is taken; For the state o t+1 When taking action according to the strategy t+1 ; is the updated gradient about θ, To find the gradient with respect to θ, The state is o t+1 、Action is a t+1 , the evaluation network parameters are Evaluation status value of the situation; By minimizing the expected KL divergence to update the policy parameters, the parameter update formula of the action network is: In the formula, is the parameter update function of the action network, It means that the state belongs to the experience replay pool and the action belongs to the strategy π. The action network parameters are The corresponding strategy, Q θ is the evaluation state value corresponding to the evaluation network parameter θ, It means that the input noise vector at time t is ∈ t , input state is o t In this case, the action output by the action network is equivalent to a t , obtained by the neural network in the action network; ∈ t is the input noise vector at time t; For the updated The gradient of Ask about The gradient of In order to find out about a t The gradient of Considering automatic entropy adjustment, the gradient of α is calculated by minimizing the update function J(α) of the temperature parameter α, and the temperature parameter α is updated. The specific formula is: Where J(α) represents the update function of the temperature parameter α, For action a t Belong to strategy When you seek hope, is the updated gradient of the temperature parameter α, is the minimum entropy constant.
9. A method for online voltage control of a distribution network, characterized in that: The method comprises: Obtaining real-time observation status values of each area in the distribution network; the real-time observation status values include real-time network node voltage, active power and reactive power of inverter-based distributed energy, and charge state of the energy storage system; Determine the current optimal strategy for each of the regions based on the real-time observed state value and the voltage control strategy model obtained according to any one of claims 1 to 8; Using the current optimal strategy to online control and adjust the voltage of the network nodes in each of the areas; The optimal strategy corresponds to the optimal action of the voltage control strategy model.
10. The method according to claim 9, characterized in that The optimal strategy includes the active power of energy storage nodes charged and discharged in each of the regions and the reactive power of inverters of distributed energy nodes. The current optimal strategy is used to online control and adjust the voltage of network nodes in each of the regions, including: For each of the regions, the active power of charging and discharging of the energy storage nodes and the reactive power of the inverters of the distributed energy nodes in the region are adjusted based on the optimal strategy control to adjust the voltage of each network node in the region.
Citation Information
Cited By
Distributed energy storage aggregation scheduling method and system considering user fatigue effect
CN121390962A
A distributed energy storage aggregation scheduling method and system considering user fatigue effect
CN121390962B
Power distribution network adaptive voltage regulation method and system based on real-time data feedback
CN121440646A